Simulation evaluates ordered object-and-anchor localization, without grasping or placement; source footage 8×. The real clip shows online scene reconstruction at 4×. These are independent experiments.仿真评估先物体、后锚点的顺序定位,不含抓取或放置;原片 8 倍速。实机片段展示在线场景重建,原片 4 倍速。两段为独立实验。
Research title研究题名
A Flow-Matching Model over Online 3D Gaussian Splatting for Image-Goal Object Search and Restoration基于在线三维高斯泼溅的流匹配模型:图像目标驱动的物体搜索与复原
One past photo shows where an object used to belong. The robot must find the displaced object and its now-empty anchor. An online 3D Gaussian map and conditional flow matching generate a spatial Restoration Field that guides this process.一张旧照片记录物体原本的位置。机器人需要找到被挪走的物体和如今空置的原始锚点。在线三维高斯地图与条件流匹配共同生成空间恢复场,引导这一过程。
More real-robot excerpts更多实机片段
Real camera observations alongside the evolving online 3D Gaussian map.实机相机观测与持续更新的在线三维高斯地图。Source footage: 4× playback.原片为 4 倍速播放。
Parallel real-trial excerpts show grasping and returning toward the contextual anchor.并列实机实验片段展示抓取,以及返回上下文锚点的过程。Source footage: 4× playback; excerpts show different stages and outcomes.原片为 4 倍速播放;这些片段呈现不同阶段与结果。
Parallel excerpts from the ten real trials, including different outcomes and progress.十次实机实验的并列片段,包含不同结果与进度。Source footage: 4× playback. See the full video for all outcomes.原片为 4 倍速播放。完整结果见下方全片。
Watch the complete research video观看完整研究视频3:49
Complete research video完整研究视频
3:49 · Full video完整视频
The full video includes all ten real trials, successes and failures. Mapping footage plays at 4× and simulated rollouts at 8×. Simulation evaluates ordered localization of the object and anchor; physical grasping and placement are evaluated on the real robot.完整视频包含全部 10 次真实实验及成功与失败案例。建图片段以 4 倍速、仿真片段以 8 倍速播放。仿真评估按顺序定位物体与锚点;物理抓取和放置在真实机器人上评估。
After the object moves, its original support region is empty; object appearance alone cannot identify this return location.物体被移动后,原支撑区域已空置,仅靠物体外观无法确定归位位置。
The original support region is empty in the current scene. Finding the object and recognizing its former context are both necessary.当前场景中的原始支撑区域已经空置。任务需要同时找到物体,并识别它原来所处的环境上下文。Full-size figure查看原图The online map and goal image condition a latent flow model. Its decoder produces a Restoration Field that drives next-best-view exploration and persistent target readout.在线地图和目标图像为潜在流模型提供条件,解码器生成恢复场,用于下一最佳视角探索与稳定目标读出。Full-size figure查看原图
1
Map what is visible在线表示可见空间
Self-pose estimates are an input. RGB-D observations update an online 3D Gaussian map. Visibility-aware projection produces spatial triplanes; a frozen DINOv2 backbone represents the goal image.自位姿估计是系统输入。RGB-D 观测持续更新在线三维高斯地图。可见性感知投影生成空间三平面,冻结的 DINOv2 主干提取目标图像表示。
2
Generate a Restoration Field生成恢复场
Conditional flow matching jointly generates eight channels for scene state, search prior, support, object affinity and contextual-anchor affinity.条件流匹配联合生成八个通道,描述场景状态、搜索先验、支撑面、物体关联与上下文锚点关联。
3
Search, pick, return搜索、拾取、归位
The field guides exploration. Stable peaks in observed regions trigger object grasping, then anchor search and placement using manipulation primitives.恢复场引导探索,已观察区域中的稳定峰值触发物体抓取,随后进行锚点搜索,并通过操作基元完成放置。
02
Evidence for the design设计依据
A past photograph supplies appearance and context; the online map supplies the evolving spatial representation. The experiments test whether both inputs are used and where the ordered restoration task still fails.旧照片提供外观和上下文,在线地图提供持续更新的空间表示。实验考察模型是否使用这两种输入,以及有序恢复任务仍在哪个阶段失败。
The goal image and current map are both used目标照片与当前地图均参与定位
In Table 4, ordered localization success drops from 34.8% to 1.8% when goal images are shuffled and to 0.8% when the online triplane is set to zero. Removing contextual-anchor affinity gives 8.4%. These controls use the same held-out scenes and evaluation episodes. They support dependence on the goal and scene inputs in this setup; they do not compare 3DGS against every alternative map representation.表 4 中,打乱目标照片后,顺序定位成功率由 34.8% 降至 1.8%;将在线三平面置零后为 0.8%;移除上下文锚点关联后为 8.4%。这些对照使用相同的未见场景与评测任务,支持模型在该设置中对目标和场景输入的依赖,但并未比较 3DGS 与所有其他地图表示。
The empty anchor remains the observed bottleneck空置锚点仍是主要瓶颈
The same ten real trials yield eight object localizations, five anchor localizations and five physical restorations. The three counts are nested stages of one trial set. All five successful anchor localizations proceed to restoration, while repeated background structure and weak context can mislead anchor matching. The largest observed stage loss is therefore anchor search in this small sample; it does not establish that placement is generally solved.同一组十次实机试验得到八次物体定位、五次锚点定位和五次物理复原。这三个数是同一试验集中的嵌套阶段。五次成功锚点定位均进入并完成复原,而重复背景结构和弱上下文会误导锚点匹配。因此,在这个小样本中,最大的阶段损失发生在锚点搜索;这并不说明放置问题已被普遍解决。
Table 3, manuscript p. 7. Five objects, two trials each; these are sequential counts from the same ten trials, not independent samples.稿件第 7 页表 3。五件物体,每件两次;这些是同一十次试验的顺序阶段计数,不是独立样本。Full-size figure查看原图
Simulated success means locating the displaced object first and the original anchor second, within a 1 m arrival threshold; no grasping or placement is simulated. Real success additionally requires stable placement on the correct support surface with horizontal error <0.25 m. Final orientation is unconstrained.
15 held-out HM3D scenes · three categories and three seeds averaged, Table 215 个未见过的 HM3D 场景 · 三类物体与三随机种子均值,表 2
Method方法
Object success物体成功率
Anchor success锚点成功率
Ordered task顺序任务成功率
R-SPL
GaussNav
67.1%
31.1%
25.8%
0.118
IEVE
59.6%
25.9%
12.4%
0.061
Spot the Difference
62.4%
41.8%
34.8%
0.213
GaussNav and IEVE are sequentially adapted with object/anchor crops. Object and anchor rates are marginal stage scores, so anchor success may include visiting the anchor before the object. The full task requires the correct order. GaussNav retains higher object-only success.GaussNav 与 IEVE 使用物体及锚点裁图进行顺序适配。物体和锚点成功率为各阶段边际指标,因此锚点成功可能包含在找到物体前经过锚点。完整任务要求正确顺序,GaussNav 的仅物体成功率仍较高。
Qualitative bowl, cup and bottle examples. Time traces show field responses near policy targets, rather than experimental success percentages.碗、杯子和瓶子的定性示例。时间曲线表示策略目标附近的恢复场响应,不表示实验成功率。Full-size figure查看原图
Ten real-world trials十次真实实验
An AgileX Ranger MINI 3 with a six-DoF PiPER arm performs the complete task. Across ten trials, the system localizes the object eight times and the anchor five times; all five successful anchor localizations lead to physical restoration. Onboard RTX 4080 Super handles mapping, field generation and grasp proposals; Intel NUC runs Nav2. Mapping updates at approximately 8 Hz. The learned model supplies spatial information for search and position readout; hand-designed next-best-view scoring, Nav2 navigation and Contact-GraspNet grasp proposals connect it to execution.AgileX Ranger MINI 3 与六自由度 PiPER 机械臂执行完整任务。在十次实验中,系统八次成功定位物体、五次成功定位锚点;五次成功锚点定位均完成物理复原。机载 RTX 4080 Super 负责建图、恢复场生成和抓取候选,Intel NUC 运行 Nav2。建图更新频率约为 8 Hz。学习模型提供用于搜索和位置读出的空间信息;手工设计的下一最佳视角评分、Nav2 导航与 Contact-GraspNet 抓取候选将其接入动作执行。
All ten trials are represented. Black panels mark unachieved stages after earlier failure. Contextual anchor localization is the main observed bottleneck.图中呈现全部十次实验。黑色面板标记前序失败后未到达的阶段,上下文锚点定位是主要瓶颈。Full-size figure查看原图
04
Scope & limitations范围与局限
The method assumes one displaced rigid object and otherwise static scene geometry. Systematic ambiguity from identical objects or similar anchors is outside the evaluated scope. Large viewpoint changes and weak visual context can make anchor localization difficult.方法假设只有一个刚性物体被移动,其他场景几何保持不变。相同物体或相似锚点导致的系统性歧义不在已评估范围内。较大视角变化及较弱视觉上下文会增加锚点定位难度。
Restoration estimates position and relies on hand-designed next-best-view scoring. The real evaluation has ten trials. The approximately 8 Hz figure describes mapping updates, rather than an entire restoration or control cycle.恢复估计关注位置,并依赖手工设计的下一最佳视角评分。真实评估共十次。约 8 Hz 仅表示建图更新频率,不代表完整恢复或控制循环频率。
05
Conclusion结论
The work connects a past goal image to online 3DGS for ordered object-and-anchor localization. In ten real trials, anchor localization is the main observed bottleneck. Stronger contextual disambiguation and complete target-pose estimation remain open questions; current evaluation does not constrain final object orientation.这项工作将旧目标照片与在线 3DGS 联系起来,完成先物体、后原锚点的有序定位。在十次真实试验中,锚点定位是主要瓶颈。更可靠的上下文消歧与完整目标位姿估计仍是待解决的问题;当前评估不约束最终物体朝向。