OmniHOI

Dexterous Hand-Object Interaction from Monocular Human Video

  • 1Zhejiang University
  • 2The University of Hong Kong
  • 3Shanghai Jiao Tong University
  • 4Shanghai AI Laboratory

†Corresponding author

Abstract

Monocular videos of human manipulation provide abundant dexterous demonstrations, yet reconstructing hand-object interaction from a single view and transferring it to robot hands remain difficult, limiting their direct use for robot execution. Prior methods either require task-specific RL training, limiting scalability, or assume clean motion-capture trajectories and thus cannot operate directly on video. We present OmniHOI, a pipeline that turns an RGB video of hand-object interaction into an interaction-faithful trajectory on dexterous hands. The key idea is to enforce physical consistency using the evidence available at each stage: image evidence during reconstruction, contact geometry during retargeting, and dynamics during physics-in-the-loop refinement. Each stage optimizes the corresponding representation directly, correcting errors before they propagate downstream or must be absorbed by a learned policy. Its trajectories also execute on a real bimanual robot across diverse tasks.

Method

OmniHOI pipeline: adaptive keyframe sampling, keyframe reconstruction and contact-aware retargeting, then physics-in-the-loop refinement with CMA-ES in MuJoCo.
Overview. Adaptive sampling reduces the video clip to keyframes (left). The hand pose in each keyframe is reconstructed, refined and retargeted with contact information preserved (middle). Interpolating between keyframes gives an initial robot trajectory, which physics-in-the-loop refinement rolls out in simulation and corrects with CMA-ES (right).
Stage 1 · Keyframe reconstruction
Accumulated motion signal with keyframes at equal increments, compared with uniform sampling.

Adaptive keyframe sampling

Refining every frame is costly and mostly fits jitter. OmniHOI accumulates the second-order motion of the wrist and the finger joints, and places keyframes at equal increments of this signal: they crowd at reversals and transients and thin out on steady stretches. The frames in between are interpolated.

Hand-object reconstruction

SAM 3 segments hands, objects and table; MoGe-3 estimates intrinsics and per-frame point maps. WiLoR hands are placed by pixel-aligned hand–scene correspondences, and SAM 3D and FoundationPose give each object. Every keyframe is refined by alternating image alignment and contact–penetration optimization.

Stage 2 · Human2robot transfer

Contact-aware retargeting

A keypoint-vector solver keeps the finger arrangement, but morphology differences cause contact loss and penetration. Each contacting finger gets a contact region and surface anchors; the smallest pose change first removes link crossings, then penetration, while pulling the anchors to the demonstrated contact.

SDF gradients on a crossing finger, contact anchors and band, and MANO versus robot hand before and after repair.
Left: SDF gradients on a finger link that crosses a thin object cancel out, while the repair pushes the finger out along the bone chain. Middle: contact anchors on the robot surface and the contact band they must reach. Right: ground-truth MANO hand, kinematic retargeting (before) and the result after repair.

Physics-in-the-loop refinement

Keyframes are interpolated into a 60 Hz trajectory and replayed in MuJoCo. CMA-ES searches per-keyframe offsets in a rolling horizon, scoring each rollout by object tracking, contact, antagonistic grasp force, penetration and the state at the end of the window.

Real-world deployment

BibTeX

@article{mao2026omnihoi,
  title   = {OmniHOI: Dexterous Hand-Object Interaction from Monocular Human Video},
  author  = {Mao, Ting and Shao, Yanming and Wang, Ziheng and Liu, Haoyu and
             Wang, Yiqun and Wu, Xuanye and Mu, Yao},
  journal = {arXiv preprint arXiv:2610.10855},
  year    = {2026}
}