Learning robot manipulation policies from human video demonstrations constitutes a promising avenue for scalable robot learning. However, comparing different human-to-robot (H2R) transfer methods remains challenging, as existing approaches are evaluated under different settings, including differing task suites, scene layouts, object instances, and amounts of robot supervision.
To address this challenge, we present H2RBench, a Real2Sim benchmark for evaluating H2R transfer methods. H2RBench provides a standardized protocol built on real human video demonstrations and simulated robot demonstrations, and includes four manipulation tasks spanning diverse interaction requirements. We evaluate multiple representative H2R transfer methods, each adopting a different strategy for bridging the embodiment gap.
Using H2RBench, we systematically characterize how each method scales with the amount of human demonstrations, revealing that methods differ substantially in their ability to leverage additional human data. We further show that simulation performance is broadly predictive of real-world robot performance, with an overall Pearson correlation of r = 0.89, Spearman correlation of ρ = 0.85, and Mean Maximum Rank Violation (MMRV) of 0.06 across method–task configurations. These results establish H2RBench as a practical and scalable benchmark for comparative H2R evaluation prior to real-world deployment.
Figure 1: H2RBench Evaluation Pipeline. We reconstruct real-world workspaces into Isaac Lab simulation environments, collect task-specific human and robot demonstrations across four manipulation tasks, and train and evaluate all methods under a shared benchmark protocol.
H2RBench provides a unified observation-action interface supporting diverse H2R policy representations. Synchronized front, side, and wrist-mounted RGB-D camera views are provided together with robot proprioception and language observations. Human demonstrations are recorded using static front and side cameras, while robot demonstrations additionally include wrist-mounted observations. The robot demonstration budget is held fixed across methods; only the number of human demonstrations is varied, enabling controlled analysis of human-data scaling.
The four tasks span a broad range of manipulation challenges with varying precision requirements and task horizons, each reconstructed from a real-world scene.
Pick a mug from a randomized position and place it onto a fixed plate. 40 robot demos, up to 100 human demos.
Sim Robot Demos
Camera 0
Camera 1
Wrist
Real Human Demos
Front
Left
Stack one bowl onto another. Requires precise terminal pose alignment. 100 robot demos, up to 300 human demos.
Sim Robot Demos
Camera 0
Camera 1
Wrist
Real Human Demos
Front
Left
Insert a donut onto a fixed peg. Contact-rich with tight positional tolerances. 100 robot demos, up to 300 human demos.
Sim Robot Demos
Camera 0
Camera 1
Wrist
Real Human Demos
Front
Left
Sequentially place multiple blocks into a box. Tests long-horizon execution. 100 robot demos, up to 300 human demos.
Sim Robot Demos
Camera 0
Camera 1
Wrist
Real Human Demos
Front
Left
We benchmark four representative H2R transfer methods spanning the major embodiment-bridging strategies:
Inpaints the human arm and overlays a rendered robot arm on human demos before imitation learning.
Represents human and robot behavior via shared 3D keypoints from multi-view observations.
Decouples latent dynamics learning from action inference; learns dynamics from videos then predicts actions from robot demos.
Hierarchically decomposes transfer into a high-level subgoal planner (human+robot demos) and a low-level controller (robot demos only).
GHOST achieves the strongest and most consistent performance across all four tasks. Methods with hierarchical decomposition or explicit subgoal representations tend to achieve stronger and more robust performance, suggesting that structured intermediate abstractions are effective for bridging the embodiment gap.
| Method | Pick-Place | Stacking | Insertion | Sequential | ||||
|---|---|---|---|---|---|---|---|---|
| Prog. | SR | Prog. | SR | Prog. | SR | Prog. | SR | |
| Phantom | 77.0 | 56.7 | 75.6 | 25.6 | 39.2 | 7.8 | 37.8 | 13.3 |
| AMPLIFY | 74.4 | 58.9 | 85.3 | 61.1 | 35.0 | 2.2 | 8.1 | 0.0 |
| Point Policy | 82.6 | 62.2 | 79.7 | 61.1 | 37.8 | 6.7 | 17.4 | 0.0 |
| GHOST | 91.1 | 82.2 | 88.9 | 76.7 | 64.7 | 34.5 | 63.3 | 38.9 |
Best per column is bold; second-best is underlined. Prog. = average task progression; SR = final-stage success rate. Mean across random seeds, under robot + maximum human demonstration training.
Human demonstrations consistently improve performance on pick-and-place and stacking (high-level spatial tasks), while gains are limited for insertion and sequential manipulation, where fine-grained contact dynamics and error accumulation remain challenging.
Absolute performance: robot-only vs. robot+human.
Human transfer impact (log-odds gain). Positive = beneficial.
Figure 2: Capability and human transfer impact across H2R methods and tasks.
For simpler tasks most methods improve steadily with more human demonstrations. For precision-sensitive tasks (insertion) and long-horizon tasks (sequential manipulation), scaling trends are less consistent—some methods saturate or fluctuate, suggesting that compounding execution errors are difficult to address through additional human data alone.
Figure 3: Policy performance as a function of human demonstration count.
For each method–task pair we train (i) a simulation-pipeline policy (real human demos + simulated robot demos, evaluated in simulation) and (ii) a real-world-pipeline policy (real human demos + real robot demos, evaluated on the physical robot). We measure Sim-vs-Real correspondence using Pearson r, Spearman ρ, and MMRV.
| Condition | Pearson r ↑ | Spearman ρ ↑ | MMRV ↓ |
|---|---|---|---|
| Robot Only | 0.971 | 0.940 | 0.014 |
| Robot + Max Human | 0.894 | 0.851 | 0.060 |
N = 16 method–task pairs per condition.
Robot-Only
Robot + Max Human Demos
Figure 4: Sim-vs-Real correlation. Each point is one method–task pair. Higher simulated performance reliably predicts higher real-world performance.
Simulation performance is broadly predictive of real-world robot performance across both training conditions, supporting H2RBench as a practical proxy for comparative H2R evaluation prior to real-world deployment.
We evaluate all four methods across four manipulation tasks in simulation, showing representative success and failure rollouts for each method–task pair.
✓ Success
✗ Failure
✓ Success
✗ Failure
✓ Success
✗ Failure
✓ Success
✗ Failure
✓ Success
✗ Failure
✓ Success
✗ Failure
✓ Success
✗ Failure
✓ Success
✗ Failure
✓ Success
✗ Failure
✓ Success
✗ Failure
✓ Success
✗ Failure
✓ Success
✗ Failure
✓ Success
✗ Failure
✓ Success
✗ Failure
✓ Success
✗ Failure
✓ Success
✗ Failure
We evaluate all four methods across four manipulation tasks in real-world settings, showing representative success and failure rollouts for each method–task pair.
✓ Success
✗ Failure
Not grasp the mug
Not grasp the mug
✓ Success
✗ Failure
Final stacking step fail
Fail to grasp the green bowl
✓ Success
No successful trials recorded.
✗ Failure
Donut grasp contact point is not correct
Miss the donut grasp
Grasp but fail to lift to good height
✓ Success
No successful trials recorded.
✗ Failure
Grasp no block
Only grasp the first block
Only grasp the first two blocks
✓ Success
✗ Failure
✓ Success
✗ Failure
✓ Success
✗ Failure
✓ Success
No successful trials recorded.
✗ Failure
✓ Success
No successful trials recorded.
✗ Failure
Only grasp the mug
Fail to grasp the mug
✓ Success
✗ Failure
Fail to grasp the green bowl
Fail to grasp the green bowl
✓ Success
No successful trials recorded.
✗ Failure
Fail to grasp the donut
Grasp but fail to lift to good height
Fail to reach the donut
✓ Success
No successful trials recorded.
✗ Failure
Grasp no block
Grasp one block
Grasp two blocks
✓ Success
✗ Failure
✓ Success
✗ Failure
✓ Success
✗ Failure
✓ Success
✗ Failure