← Back to project page

ECCV 2026 · Camera-ready results

Quantitative evaluation

240evaluation videos
82.50Dynamic Degree
61.23DINO-I
288 hA100 training

We use 30 reference images with 8 prompts per image. VBench measures motion, text alignment, and subject fidelity; OpenS2V-Eval v1.1 reports the Single-Domain benchmark.


Evaluation Metrics

Motion Smoothness

Measures temporal consistency across frames. Higher values indicate smoother, more coherent motion without jittering or flickering.

Dynamic Degree

Quantifies the amount of motion in generated videos. Higher values indicate more dynamic, non-static video content.

CLIP-T (Text Alignment)

Measures how well the generated video aligns with the input text prompt using CLIP embeddings.

CLIP-I / DINO-I (Identity)

Measures subject fidelity using CLIP and DINO embeddings. Higher DINO-I indicates better preservation of fine-grained subject details.


VBench Evaluation

We compare SDV-Gen methods with subject-driven image generation followed by CogVideoX-5B I2V. Phantom and VACE use their public Wan 2.1-1.3B versions.

Method Motion Smooth.↑ Dynamic Deg.↑ CLIP-T↑ CLIP-I↑ DINO-I↑
SDI-Gen → I2V methods
OminiControl 98.40 47.92 33.06 72.09 52.61
BLIP 97.81 48.10 28.83 79.27 57.21
IP-Adapter 97.74 51.90 28.94 76.60 51.41
SDV-Gen methods
VideoBooth 96.89 53.33 29.59 65.65 34.17
Phantom-1.3B 98.70 67.08 33.50 73.00 53.55
VACE-1.3B 98.74 43.75 33.70 73.54 54.47
MAGREF-480P 98.86 65.09 33.00 71.15 48.50
Ours (CogVideoX-5B) 98.23 76.25 33.65 76.19 61.23
Ours (Wan 2.1-1.3B) 98.97 53.75 32.51 74.65 58.77
Ours (Wan 2.2-5B) 98.09 82.50 31.29 76.97 57.27

Orange, yellow, and gray indicate first, second, and third place, respectively.

Key Findings

  • Wan 2.2-5B: highest Dynamic Degree, 82.50.
  • Wan 2.1-1.3B: highest Motion Smoothness, 98.97.
  • CogVideoX-5B: highest DINO-I, 61.23, with Dynamic Degree of 76.25.

OpenS2V-Eval v1.1

Single-Domain results without per-subject tuning or domain-specific adaptation. Total is the normalized weighted sum of the remaining scores.

Method Training Cost Total↑ Aes.↑ Smooth.↑ Amp.↑ FaceSim↑ Gme↑ Nexus↑ Natural↑
Closed-source models
Vidu 2.0 - 52.90 43.32 91.88 17.52 36.19 66.96 44.84 66.11
Pika 2.1 - 53.12 47.43 86.07 26.32 32.33 69.84 47.35 64.68
Kling 1.6 - 56.67 45.97 85.76 47.17 39.27 65.36 49.30 73.63
Zero-shot SDV-Gen baselines
VACE-P1.3B ≈70K h 49.20 48.93 95.68 11.91 18.04 70.78 36.24 66.85
VACE-1.3B ≈70K h 51.13 49.41 95.42 22.51 22.37 70.87 38.34 68.33
VACE-14B ≈210K h 61.75 48.94 93.16 19.69 64.65 65.86 50.82 70.56
Phantom-1.3B ≈10K h 54.50 49.00 93.70 16.38 44.03 69.54 37.72 66.76
Phantom-14B ≈30K h 57.02 47.46 94.86 41.55 51.82 70.07 35.30 71.11
SkyReels-A2-P14B - 55.06 40.85 85.54 26.41 54.42 61.81 48.60 61.85
MAGREF-480P - 53.44 46.31 92.63 27.43 33.77 69.02 42.45 68.33
Ours · 288 A100 hours
Ours (CogVideoX-5B) 288 h 50.05 45.40 93.90 19.38 18.05 70.53 41.23 68.52
Ours (Wan 2.1-1.3B) 288 h 51.45 43.80 93.68 17.55 13.90 71.12 45.21 75.65
Ours (Wan 2.2-5B) 288 h 56.84 44.70 84.97 14.38 48.78 61.98 49.97 71.30

Single-Domain benchmark. Orange, yellow, and gray indicate first, second, and third Total score.

Efficiency: 288 A100-hours is roughly 35–104× less compute than Phantom-1.3B/14B and 240–730× less than VACE-1.3B/14B.

Analysis

  • Wan 2.2-5B: Total score 56.84, close to Phantom-14B at 57.02.
  • Wan 2.1-1.3B: Total score 51.45, above VACE-1.3B at 51.13.
  • CogVideoX-5B: Gme score 70.53 with the same 288-hour training budget.

Training cost

Data and compute requirements reported in the camera-ready paper. Hours are total A100 GPU hours.

Method Zero-shot Required inputs Base Model A100 h
Per-subject tuning
CustomCrafter No 200 regularizing images / subject VideoCrafter2 (1.4B) ~200 / subj.
Still-Moving No Few reference images + 40 videos Lumiere (1.2B) -
Zero-shot methods
VideoBooth 48,724 subject-video pairs SD-based VDM (1.08B) 775–1,938*
VACE Yes 53M source videos LTX & Wan (1.3–14B) 70K–210K*
Phantom Yes 1M subject-video pairs Wan (1.3–14B) + Seed 10K–30K*
Ours-tiny/mini Yes 2K/4K subject-image pairs + 4K videos CogVideoX (5B) 288
Ours Yes 200K subject-image pairs + 4K videos CogVideoX (5B) & Wan (1.3B/5B) 288

* Estimated from reported batch sizes and wall-clock training time. △ follows the partial zero-shot designation in the paper.

Each of our backbones is trained for 288 A100-hours, using subject-image pairs for identity injection and 4,000 arbitrary videos for motion-awareness preservation.