ECCV 2026 · Camera-ready results
Quantitative evaluation
We use 30 reference images with 8 prompts per image. VBench measures motion, text alignment, and subject fidelity; OpenS2V-Eval v1.1 reports the Single-Domain benchmark.
Evaluation Metrics
Motion Smoothness
Measures temporal consistency across frames. Higher values indicate smoother, more coherent motion without jittering or flickering.
Dynamic Degree
Quantifies the amount of motion in generated videos. Higher values indicate more dynamic, non-static video content.
CLIP-T (Text Alignment)
Measures how well the generated video aligns with the input text prompt using CLIP embeddings.
CLIP-I / DINO-I (Identity)
Measures subject fidelity using CLIP and DINO embeddings. Higher DINO-I indicates better preservation of fine-grained subject details.
VBench Evaluation
We compare SDV-Gen methods with subject-driven image generation followed by CogVideoX-5B I2V. Phantom and VACE use their public Wan 2.1-1.3B versions.
| Method | Motion Smooth.↑ | Dynamic Deg.↑ | CLIP-T↑ | CLIP-I↑ | DINO-I↑ |
|---|---|---|---|---|---|
| SDI-Gen → I2V methods | |||||
| OminiControl | 98.40 | 47.92 | 33.06 | 72.09 | 52.61 |
| BLIP | 97.81 | 48.10 | 28.83 | 79.27 | 57.21 |
| IP-Adapter | 97.74 | 51.90 | 28.94 | 76.60 | 51.41 |
| SDV-Gen methods | |||||
| VideoBooth | 96.89 | 53.33 | 29.59 | 65.65 | 34.17 |
| Phantom-1.3B | 98.70 | 67.08 | 33.50 | 73.00 | 53.55 |
| VACE-1.3B | 98.74 | 43.75 | 33.70 | 73.54 | 54.47 |
| MAGREF-480P | 98.86 | 65.09 | 33.00 | 71.15 | 48.50 |
| Ours (CogVideoX-5B) | 98.23 | 76.25 | 33.65 | 76.19 | 61.23 |
| Ours (Wan 2.1-1.3B) | 98.97 | 53.75 | 32.51 | 74.65 | 58.77 |
| Ours (Wan 2.2-5B) | 98.09 | 82.50 | 31.29 | 76.97 | 57.27 |
Orange, yellow, and gray indicate first, second, and third place, respectively.
Key Findings
- Wan 2.2-5B: highest Dynamic Degree, 82.50.
- Wan 2.1-1.3B: highest Motion Smoothness, 98.97.
- CogVideoX-5B: highest DINO-I, 61.23, with Dynamic Degree of 76.25.
OpenS2V-Eval v1.1
Single-Domain results without per-subject tuning or domain-specific adaptation. Total is the normalized weighted sum of the remaining scores.
| Method | Training Cost | Total↑ | Aes.↑ | Smooth.↑ | Amp.↑ | FaceSim↑ | Gme↑ | Nexus↑ | Natural↑ |
|---|---|---|---|---|---|---|---|---|---|
| Closed-source models | |||||||||
| Vidu 2.0 | - | 52.90 | 43.32 | 91.88 | 17.52 | 36.19 | 66.96 | 44.84 | 66.11 |
| Pika 2.1 | - | 53.12 | 47.43 | 86.07 | 26.32 | 32.33 | 69.84 | 47.35 | 64.68 |
| Kling 1.6 | - | 56.67 | 45.97 | 85.76 | 47.17 | 39.27 | 65.36 | 49.30 | 73.63 |
| Zero-shot SDV-Gen baselines | |||||||||
| VACE-P1.3B | ≈70K h | 49.20 | 48.93 | 95.68 | 11.91 | 18.04 | 70.78 | 36.24 | 66.85 |
| VACE-1.3B | ≈70K h | 51.13 | 49.41 | 95.42 | 22.51 | 22.37 | 70.87 | 38.34 | 68.33 |
| VACE-14B | ≈210K h | 61.75 | 48.94 | 93.16 | 19.69 | 64.65 | 65.86 | 50.82 | 70.56 |
| Phantom-1.3B | ≈10K h | 54.50 | 49.00 | 93.70 | 16.38 | 44.03 | 69.54 | 37.72 | 66.76 |
| Phantom-14B | ≈30K h | 57.02 | 47.46 | 94.86 | 41.55 | 51.82 | 70.07 | 35.30 | 71.11 |
| SkyReels-A2-P14B | - | 55.06 | 40.85 | 85.54 | 26.41 | 54.42 | 61.81 | 48.60 | 61.85 |
| MAGREF-480P | - | 53.44 | 46.31 | 92.63 | 27.43 | 33.77 | 69.02 | 42.45 | 68.33 |
| Ours · 288 A100 hours | |||||||||
| Ours (CogVideoX-5B) | 288 h | 50.05 | 45.40 | 93.90 | 19.38 | 18.05 | 70.53 | 41.23 | 68.52 |
| Ours (Wan 2.1-1.3B) | 288 h | 51.45 | 43.80 | 93.68 | 17.55 | 13.90 | 71.12 | 45.21 | 75.65 |
| Ours (Wan 2.2-5B) | 288 h | 56.84 | 44.70 | 84.97 | 14.38 | 48.78 | 61.98 | 49.97 | 71.30 |
Single-Domain benchmark. Orange, yellow, and gray indicate first, second, and third Total score.
Efficiency: 288 A100-hours is roughly 35–104× less compute than Phantom-1.3B/14B and 240–730× less than VACE-1.3B/14B.
Analysis
- Wan 2.2-5B: Total score 56.84, close to Phantom-14B at 57.02.
- Wan 2.1-1.3B: Total score 51.45, above VACE-1.3B at 51.13.
- CogVideoX-5B: Gme score 70.53 with the same 288-hour training budget.
Training cost
Data and compute requirements reported in the camera-ready paper. Hours are total A100 GPU hours.
| Method | Zero-shot | Required inputs | Base Model | A100 h |
|---|---|---|---|---|
| Per-subject tuning | ||||
| CustomCrafter | No | 200 regularizing images / subject | VideoCrafter2 (1.4B) | ~200 / subj. |
| Still-Moving | No | Few reference images + 40 videos | Lumiere (1.2B) | - |
| Zero-shot methods | ||||
| VideoBooth | △ | 48,724 subject-video pairs | SD-based VDM (1.08B) | 775–1,938* |
| VACE | Yes | 53M source videos | LTX & Wan (1.3–14B) | 70K–210K* |
| Phantom | Yes | 1M subject-video pairs | Wan (1.3–14B) + Seed | 10K–30K* |
| Ours-tiny/mini | Yes | 2K/4K subject-image pairs + 4K videos | CogVideoX (5B) | 288 |
| Ours | Yes | 200K subject-image pairs + 4K videos | CogVideoX (5B) & Wan (1.3B/5B) | 288 |
* Estimated from reported batch sizes and wall-clock training time. △ follows the partial zero-shot designation in the paper.
Each of our backbones is trained for 288 A100-hours, using subject-image pairs for identity injection and 4,000 arbitrary videos for motion-awareness preservation.