MotionInsight

Diagnosing Object Motion Deficiencies in Generated Videos

Jiahao Zhan1,2, Yongrui Ma1,2, Qunliang Xing2, Xuanyu Zhang4, Jingqi Tong3, Junlin Li2, Li Zhang2, Shijie Zhao2,†,✉, Tianfan Xue1,5,✉

1 MMLab, CUHK 2 ByteDance Inc. 3 Fudan University 4 Peking University 5 CPII under InnoHK

† Project Lead   ✉ Corresponding Authors

EMNLP 2026 Findings

Abstract

MotionInsight evaluates the fidelity of designated objects in generated videos through object consistency, motion continuity, and physical plausibility. We introduce VidMotion, a dataset with dimension-specific quality scores and annotations of motion failures. Our evaluator combines sampled RGB frames with explicit representations of object trajectories and camera motion to identify temporal deficiencies that can be difficult to detect from appearance alone. Motion description alignment connects these representations to a vision-language model, while motion-specific GRPO rewards improve scoring and diagnostic explanations. Experiments show strong agreement with human assessments. We further use MotionInsight scores to construct preference pairs for DPO post-training of a video generator, improving human preference for motion fidelity and overall video quality.

Overview

Overview of MotionInsight. Object tracking features and camera poses are fused into motion representations, then interleaved with RGB visual tokens. This explicit motion information supports object-level diagnosis along three dimensions: object consistency, motion continuity, and physical plausibility. Click the figure to enlarge it.

MotionInsight for Video Generation

We use MotionInsight as a reward model to construct preference pairs for DPO post-training of Wan2.1-T2V-1.3B. For each prompt, eight candidate videos are sampled. Each salient object receives scores across the three motion dimensions; the normalized scores are multiplied, and the lowest object reward is used as the video reward. The highest- and lowest-reward videos form a preference pair.

In a two-alternative forced-choice (2AFC) study with 20 volunteers, MotionInsight-DPO is preferred over both the original Wan 2.1 model and VideoPhy2-DPO.

Human preference for MotionInsight-DPO (%)
Compared againstMotion fidelityOverall quality
Wan 2.192.8686.67
VideoPhy2-DPO87.5687.50

Video Generation Comparisons

Explore examples of generated videos. Each clip compares three outputs for the same prompt on a shared timeline. Select a thumbnail to view another example.

Horse Jumping an Obstacle

6 / 13
Wan
DPO baseline
MotionInsight-DPO (Ours)

Prompt: A rider in a black helmet and white shirt, seated on a brown horse, leaning forward during a jump over an obstacle with decorative green and pink plants at its base.

BibTeX

Download
@misc{zhan2026motioninsightdiagnosingobjectmotion,
  title         = {MotionInsight: Diagnosing Object Motion Deficiencies in Generated Videos},
  author        = {Jiahao Zhan and Yongrui Ma and Qunliang Xing and Xuanyu Zhang and Jingqi Tong and Junlin Li and Li zhang and Shijie Zhao and Tianfan Xue},
  year          = {2026},
  eprint        = {2609.37030},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CV},
  url           = {https://arxiv.org/abs/2609.37030},
}

MotionInsight Framework

Enlarged MotionInsight framework diagram