Preprint

DeVI: Physics-based Dexterous Human-Object Interaction via Synthetic Video Imitation

Seoul National University1, RLWRLD2

TL;DR: Physics-based character control with dexterous HOI via Synthetic Video Imitation

Abstract

Recent advances in video generative models enable the synthesis of realistic Human-Object Interaction (HOI) videos across diverse scenarios and object categories, including complex dexterous manipulations that are difficult to capture using motion capture systems. While these synthetic videos offer rich knowledge of how people grasp and use objects, inaccurate object pose estimates and severe hand-object misalignment make them difficult to use for physics-based character control. We present DeVI (Dexterous Video Imitation), a novel framework that learns physics-based character control for dexterous HOI by imitating text-conditioned synthetic videos. To train a control policy from these videos, we introduce hybrid imitation targets that combine reconstructed 3D human motion with tracked 2D object trajectories. This formulation avoids explicit 3D object motion reconstruction that remains unreliable for dexterous interactions in synthetic videos. Additionally, our Visual HOI Alignment corrects misalignment of the human reference with the generated video and the initial object geometry, making the reference suitable for physical interaction with the object. Extensive experiments on synthetic video imitation show that DeVI's hybrid representation is more effective than those used by the baselines. Even when 3D HOI motion-capture data are provided as references, DeVI outperforms the baselines in imitating dexterous interactions. We further highlight the advantages of synthetic video imitation by showing text-controlled functional manipulation, target-aware interactions, and articulated object manipulation without requiring motion-capture demonstrations.

Method Overview

DeVI pipeline: HOI video generation, hybrid imitation target extraction, and humanoid control policy learning

Our method consists of three parts; (1) 2D HOI Video Generation, (2) Extracting Hybrid Imitation Target from the Video, and (3) Learning Humanoid Control Policy. First, we generate 2D HOI Video from the rendered 3D scene using the pre-trained image-to-video diffusion model. Then, the Hybrid Imitation Target which includes 3D human reference and 2D object reference is extracted from the video. Using the hybrid imitation target, we learn humanoid control policy imitating the dexterous HOI video via our hybrid tracking reward.

Key Takeaways

Limited Generalizability of Universal Controllers

Universal controller failure cases across trophy, camera, coke, apple, pot, garbage, wok, and cactus interactions
Universal controllers have limited generalizability to novel HOI motions. Due to the scarcity of 4D HOI data, existing universal controllers often fail to imitate novel dexterous HOI motions, including those reconstructed from videos.

Hybrid Imitation Target

Comparison of policies trained with RigVid, DAViD, Dex4D, and DeVI imitation targets across five dexterous human-object interactions
The hybrid imitation target enables more effective dexterous HOI policy learning. The hybrid imitation target is a novel representation that combines reconstructed 3D human motion with 2D object tracks from a generated HOI video. This formulation bypasses explicit 3D object trajectory reconstruction, which is used by baselines (e.g., RIGVid, DAViD, Dex4D) but remains unreliable (especially spatio-temporal hand-object misalignement) for dexterous manipulation.

Visual HOI Alignment

Visual HOI alignment produces human motion suitable for physical interaction. Visual HOI alignment refines human motion to align with both the generated video and the initial 3D object, allowing the human motion to be suitable for interaction.

Results

Various Simulated HOIs

Trophy
Camera
Coke
Apple
Wok
Garbage
Pot
Potted Plant

Target Awareness & Text Controllability

"Lifts a pot lid with right hand and places it onto the pot to close it"
"Lifts a frying pan with right hand and places it onto the induction"
"Picks up an apple with left hand and places into a brown basket"
"Picks up a tomato with right hand and places into a brown basket"

Visualization of Hybrid Imitation Target

Zero-Shot Cross-Object Policy Transfer

Egocentric Video w/ Articulated Object Manipulation

Generated Egocentric Video
Dexterous HOI Simulation
Generated Egocentric Video
Dexterous HOI Simulation

DeVI on GRAB dataset

Additional Results

Beyond the Table-top Scenarios

Reconstructed 3D scene used for the interaction simulation

DeVI on Reconstructed Scene

BibTeX

@misc{devi,
    title={DeVI: Physics-based Dexterous Human-Object Interaction via Synthetic Video Imitation}, 
    author={Hyeonwoo Kim and Jeonghwan Kim and Kyungwon Cho and Hanbyul Joo},
    year={2026},
    eprint={2604.20841},
    archivePrefix={arXiv},
    primaryClass={cs.CV},
    url={https://arxiv.org/abs/2604.20841}, 
}