All research全部研究

ICRA · MANUSCRIPT UNDER REVIEW研究稿件 · 审稿中

Copy-Paste

Restore object positions to a visual reference.根据一张参考图像,恢复物体的位置。

Liu Xu · First author刘旭 · 第一作者

Drag right to reveal the real robot向右拖动,查看实机

Simulation仿真
Real robot实机

Simulation restoration · source footage 4×. Real-robot restoration · source footage 2×, motion smoothed for display. These are independent trials.仿真恢复:原片 4 倍速。实机恢复:原片 2 倍速,含展示用运动平滑。两段为独立试验。

Research title研究题名

Toward Whole-Body Intelligence through an Iterative World Action Model for Precise Scene Restoration面向全身智能:通过迭代世界动作模型实现精细场景恢复

A mobile manipulator restores object positions to match a visual reference. Learned docking connects base placement to manipulation outcomes, and an iterative World Action Model refines future predictions before generating arm actions.移动操作机器人根据视觉参考恢复物体位置。学习式停靠将底盘位置与操作结果联系起来,迭代世界动作模型则在生成机械臂动作之前细化未来预测。

More real-robot excerpts更多实机片段
Watch the complete research video观看完整研究视频 2:58

Complete research video完整研究视频

2:58 · Full video完整视频

The full demonstration includes the method, simulation, one real task and 18 additional real examples. Experimental footage uses accelerated playback. Motion capture is used only to measure outcomes.完整演示包含方法、仿真、一段真实任务及另 18 个真实实验示例。实验片段采用加速播放,动作捕捉仅用于结果测量。

Download full video下载完整视频

01

Learned docking & iterative WAM学习式停靠与迭代 WAM

Base placement changes both the arm's starting state and camera viewpoint. Docking utility is therefore learned from restoration outcomes of a frozen WAM.底盘位置改变机械臂起始状态和相机视点。因此,停靠效用以冻结 WAM 的恢复结果为监督。

Copy-Paste 框架:学习式底盘停靠与迭代世界动作模型
A coarse-to-fine execution pipeline: choose a docking pose, refine future-video predictions, generate actions, observe and replan.由粗到细的执行流程:选择停靠位置、细化未来视频预测、生成动作、重新观测并规划。 Full-size figure查看原图
  1. 1

    Dock for the task为任务选择停靠位置

    Docking takes the current head-camera XYZ-RGB point cloud; the goal image conditions the subsequent WAM manipulation stage. A point-cloud pose prior proposes base poses. After collision filtering, prior and utility scores are combined to select a feasible pose. The utility predictor is trained on restoration outcomes from a frozen manipulation model.停靠模块输入当前头部相机的彩色点云;目标图像用于后续 WAM 操作。点云位姿先验提出候选底盘位置,经过碰撞过滤后,融合先验与效用评分选择可行停靠位姿。效用预测器以冻结操作模型的恢复结果为监督进行训练。

  2. 2

    Refine future predictions细化未来预测

    The default model refines its future video twice. Each call starts from fresh noise and conditions on the preceding prediction, while current camera observations and the visual goal stay fixed. New observations after physical action start the next planning round.默认对未来视频进行两轮细化。每轮从新的噪声开始,以前一轮预测作为条件;当前相机观测和视觉目标保持不变。实际动作执行后的新观测用于下一轮规划。

  3. 3

    Execute, observe, replan执行、观测、重新规划

    Only the final world features guide action generation. The robot executes the first 24 actions in the predicted sequence, then obtains new observations and plans again.仅将最终世界特征用于动作生成。机器人执行预测动作序列的前 24 步,随后获取新观测并再次规划。

02

Evidence for the design设计依据

The component study separates docking utility from future refinement. Its controls use three training seeds and the same 150 simulation cases per seed, with a 10 mm success threshold plus release and safety criteria.组件实验分别考察停靠效用与未来预测细化。对照使用三个训练种子,每个种子评估相同的 150 个仿真案例;成功需满足 10 mm 位置阈值,以及释放和安全条件。

Docking learned from manipulation outcomes以操作结果监督停靠选择

The WAM manipulation policy is frozen before docking utility is learned. Its complete restoration outcomes label candidate poses, so ranking reflects the manipulation that a pose enables. In Table II, removing utility scoring while retaining the same twelve candidates reduces mean success from 71.1% to 47.3%. This supports the outcome-based ranking within the reported benchmark; it is a staged docking-and-manipulation framework.WAM 操作策略在学习停靠效用之前被冻结,其完整恢复结果用于标记候选位姿,使排序反映该位姿下后续操作的表现。表 II 中,保留相同的十二个候选、但去除效用评分后,平均成功率由 71.1% 降至 47.3%。这支持所报告基准中的结果监督排序;方法仍是停靠与操作分阶段执行的框架。

表 II 停靠组件消融:三个种子成功率均值及样本标准差
Table II, manuscript p. 6. Mean ± sample SD across three seeds; each seed uses 150 fixed cases. All variants retain filtering, base control and WAM manipulation.稿件第 6 页表 II。均值 ± 三个种子的样本标准差;每个种子使用 150 个固定案例。各变体均保留候选过滤、底盘控制和 WAM 操作。 Full-size figure查看原图

Recursive refinement under controlled denoising budgets去噪预算对照下的递归细化

Two recursive calls condition on the latest predicted future and reach 71.1% mean success. The repeated-initial control reaches 65.1% and matches calls and per-call evaluations. One twenty-evaluation call reaches 66.7% and matches total denoising evaluations only. These are inference controls using the same checkpoint and action branch within each seed. The comparison supports recursive conditioning under this protocol, rather than simply more denoising evaluations.两轮递归调用均以前一轮预测为条件,平均成功率为 71.1%。两次均以初始未来为条件的对照为 65.1%,匹配调用次数和每次求值数;一次调用、二十次去噪求值的对照为 66.7%,仅匹配总去噪求值数。每个种子内,这些推理对照使用相同 checkpoint 和动作分支。结果支持该协议中的递归条件作用,而不只是增加去噪求值数。

表 III 推理对照:递归、重复初始条件与一次调用、二十次去噪求值细化
Table III, manuscript p. 6. Error bars are sample SD across three seeds. Matched denoising budgets do not imply all computation is identical.稿件第 6 页表 III。误差线为三个种子的样本标准差。匹配去噪预算不等于全部计算完全相同。 Full-size figure查看原图

A third recursive call increases mean success by only 0.2 percentage points, from 71.1% to 71.3%, while planning-round latency rises from 0.654 to 0.838 s on RTX 6000 Ada. The real Thor platform is reported separately at 1.326 s per default planning round. More refinement therefore has an explicit latency trade-off in the tested settings.第三轮递归仅使平均成功率增加 0.2 个百分点(71.1% 至 71.3%),同时 RTX 6000 Ada 上的规划轮时延由 0.654 s 增至 0.838 s。实机 Thor 的默认规划轮时延另为 1.326 s。增加细化深度在所测试设置中有明确的时延代价。

03

Measured restoration恢复实验结果

Position accuracy with explicit success criteria.以明确成功条件评估位置恢复精度。

77.8%

Simulation · ID success仿真 · 域内成功率

90 cases · 10 mm tolerance90 个案例 · 10 mm 容差

61.7%

Simulation · OOD success仿真 · 域外成功率

60 cases · 10 mm tolerance60 个案例 · 10 mm 容差

19 / 30

Real restoration successes真实恢复成功次数

3 objects × 2 tables × 5 trials3 类物体 × 2 张桌子 × 5 次

30.7 ± 12.7 mm

Real final position error真实终态位置误差

Mean ± sample SD · all 30 trials均值 ± 样本标准差 · 全部 30 次

Success requires a safe release within the allowed budget, with the object resting on the task surface without gripper support and final position error ≤10 mm in simulation or ≤30 mm on the real robot. Reported errors include failed trials. Orientation is not evaluated.

成功要求在允许预算内安全释放,使物体无需夹爪支撑而停留在任务表面,且终态位置误差在仿真中 ≤10 mm、实机中 ≤30 mm。误差统计包含失败实验,不评估朝向。

Simulation comparison仿真对比

150 fixed cases · main evaluation, Table I150 个固定案例 · 主实验,表 I
Method方法ID success域内成功率OOD success域外成功率
DP-GC11.1%1.7%
π₀.₅-GC40.0%20.0%
Hyper-GoalNet8.9%1.7%
WAM + Task-Level IC64.4%28.3%
Copy-Paste77.8%61.7%

Simulation ID covers three seen object types in six training-seen scene configurations, with 90 cases. OOD covers two held-out object types in three held-out configurations, with 60 cases. Table I reports these groups separately; the 71.1% in the component studies is mean success across three training seeds, each evaluated on the same 150 cases. The 10 mm threshold also requires release and safety criteria. Simulation and real-robot rates use different tolerances and settings.仿真 ID 为三类已见物体、六个训练中见过的场景配置,共 90 个案例;OOD 为两类未见物体、三个保留场景配置,共 60 个案例。表 I 分别报告这两组结果;组件实验的 71.1% 是三个训练种子在每个种子相同 150 个案例上的成功率均值。10 mm 阈值还需结合释放与安全条件。仿真和实机成功率使用不同容差与测试设置。

Real robot真实机器人

The real platform uses an AgileX Ranger Mini 3.0 base, a PiPER arm and a Jetson AGX Thor. Tests include a seen cup, held-out can and bottle, and seen/held-out tables. With two refinement rounds, the mean planning-round latency on Thor is 1.326 seconds.真实平台使用 AgileX Ranger Mini 3.0 底盘、PiPER 机械臂和 Jetson AGX Thor。测试涵盖见过的杯子、未见过的罐子与瓶子,以及见过和未见过的桌子。在 Thor 上,两轮细化配置每规划一轮的平均时延为 1.326 秒。

杯子、罐子、瓶子在见过和未见过桌子上的代表性恢复实验,包括仅用于评估的动作捕捉轨迹
Twelve representative examples from 30 evaluated trials. Blue and green borders distinguish seen and held-out tables; motion capture is for evaluation only.从 30 次评估中选取的 12 个代表性示例。蓝色与绿色边框分别表示见过和未见过的桌子;动作捕捉仅用于评估。 Full-size figure查看原图
04

Scope & limitations范围与局限

Evaluation measures restoration position. Object and scene coverage is limited, and robot vibration, arm-control noise and base errors can reduce precision. Refinement adds latency; a third round brings diminishing gains.实验衡量恢复位置。物体与场景覆盖范围有限,机器人振动、机械臂控制噪声及底盘误差会影响精度。细化带来额外时延,第三轮的增益较小。

Training uses 1,769 simulation trajectories and 600 real demonstrations.训练使用 1,769 条仿真轨迹和 600 条真实示范。

05

Conclusion结论

The reported experiments support docking supervised by manipulation outcomes and recursive future refinement before action generation. The real system completes 19 of 30 restorations within a 30 mm position tolerance. Refinement has a planning-latency cost, and performance remains limited by the tested objects, scenes and physical execution noise.所报告实验支持以操作结果监督停靠选择,并在动作生成前递归细化未来预测。实机系统在 30 mm 位置容差下完成 30 次恢复中的 19 次。细化需要额外规划时间,表现仍受测试物体、场景范围及真实执行噪声限制。