Instruction-based video editing benchmark

OmniEdit-Bench

A comprehensive benchmark for evaluating whether video editing models can follow explicit and implicit instructions across spatial, temporal, audio, reference-based, and reasoning scenarios.

790Total Instances
5Editing tracks
4Evaluation dimensions
Taxonomy overview of OmniEdit-Bench tracks

Why this benchmark

Video editing evaluation needs more than frame-level checks.

Existing IVE benchmarks are often narrow and image-editing centric. OmniEdit-Bench expands evaluation to video-specific dimensions such as temporal dynamics, audio alignment, reference conditioning, and implicit reasoning.

The benchmark also treats instruction fidelity as central. Its accuracy-aware scoring mechanism reduces preservation, realism, and consistency scores when the edit itself is incorrect.

Benchmark taxonomy

Five complementary tracks for diagnosing editing capability.

Spatial240

Attribute, subject, and global scene editing

Fine-grained visual changes for color, texture, material, object-level operations, relighting, weather, and style.

Temporal200

Camera, motion, action, and composition

Video-native edits that require stable trajectories, coherent object dynamics, semantic motion changes, and action-level composition.

Reference200

Spatial and temporal reference guidance

Reference-conditioned edits that test whether models can align generated outputs with external visual and motion signals.

Audio100

Speech, object sound, and ambience

Multimodal editing tasks for speech content and tone, object sounds, and environmental audio aligned with visual events.

Reasoning50

Implicit intent and multi-step inference

Physical, spatial, temporal, causal, and hypothetical reasoning edits where the desired outcome must be inferred.

Evaluation framework

Accuracy-aware scoring makes plausible but wrong edits count less.

OmniEdit-Bench decomposes editing quality into four dimensions and evaluates them with track-specific VLM prompts. Accuracy acts as a gating signal for the remaining dimensions, emphasizing instruction fidelity over surface-level visual appeal.

AccuracyWhether the edited video faithfully follows the instruction.
PreservationWhether unrelated visual or audio content remains intact.
RealismWhether edited results stay plausible, natural, and artifact-free.
ConsistencyWhether frames, motion, modality, and reference cues remain coherent.
Evaluation pipeline for OmniEdit-Bench
Accuracy-aware evaluation pipeline for OmniEdit-Bench.

Model comparison

Track-specific leaderboards avoid mixing models with different task support.

Bar charts of accuracy, preservation, realism, and consistency across tracks
Spatial8 models
  1. 1Wan2.7-Edit
    69.8
  2. 2KlingV3-Omni
    59.6
  3. 3Seedance2.0
    56.9
  4. 4Runway Aleph
    54.2
  5. 5Grok Imagine
    51.7
  6. 6UniVideo
    39.9
  7. 7VIVA
    33.5
  8. 8Ditto
    22.9
Temporal8 models
  1. 1Wan2.7-Edit
    29.0
  2. 2Seedance2.0
    25.2
  3. 3KlingV3-Omni
    18.5
  4. 4Runway Aleph
    16.5
  5. 5Grok Imagine
    15.0
  6. 6Ditto
    13.5
  7. 7UniVideo
    11.3
  8. 8VIVA
    9.6
Audio2 models
  1. 1Wan2.7-Edit
    13.6
  2. 2Grok Imagine
    6.7
Reference6 models
  1. 1KlingV3-Omni
    49.9
  2. 2Wan2.7-Edit
    46.4
  3. 3Runway Aleph
    15.1
  4. 4Grok Imagine
    13.1
  5. 5VIVA
    9.8
  6. 6UniVideo
    4.9
Reasoning8 models
  1. 1Seedance2.0
    29.8
  2. 2Wan2.7-Edit
    24.6
  3. 3KlingV3-Omni
    22.9
  4. 4Runway Aleph
    20.3
  5. 5Grok Imagine
    17.9
  6. 6Ditto
    8.6
  7. 7UniVideo
    8.4
  8. 8VIVA
    7.5

Citation

Use OmniEdit-Bench to stress-test instruction fidelity.

@article{miao2026omnieditbench,
  title   = {OmniEdit-Bench: A Comprehensive Benchmark for Instruction-based Video Editing},
  author  = {Miao, Chenxuan and Feng, Yutong and Lu, Yi and Yan, Yunfeng and Qi, Donglian and Zhang, Shiwei and Liu, Yu and Chen, Xi and Zhao, Hengshuang},
  journal = {arXiv preprint arXiv:2608.05049},
  year    = {2026}
}