All research全部研究

AAAI · MANUSCRIPT UNDER REVIEW研究稿件 · 审稿中

Spot the Difference

Find a displaced object and its original location from a past photo.从旧照片找到被移动物体与原始位置。

Liu Xu · First author刘旭 · 第一作者

Drag right to reveal the real robot向右拖动,查看实机

Simulation仿真
Real robot实机

Simulation evaluates ordered object-and-anchor localization, without grasping or placement; source footage 8×. The real clip shows online scene reconstruction at 4×. These are independent experiments.仿真评估先物体、后锚点的顺序定位,不含抓取或放置;原片 8 倍速。实机片段展示在线场景重建,原片 4 倍速。两段为独立实验。

Research title研究题名

A Flow-Matching Model over Online 3D Gaussian Splatting for Image-Goal Object Search and Restoration基于在线三维高斯泼溅的流匹配模型:图像目标驱动的物体搜索与复原

One past photo shows where an object used to belong. The robot must find the displaced object and its now-empty anchor. An online 3D Gaussian map and conditional flow matching generate a spatial Restoration Field that guides this process.一张旧照片记录物体原本的位置。机器人需要找到被挪走的物体和如今空置的原始锚点。在线三维高斯地图与条件流匹配共同生成空间恢复场,引导这一过程。

More real-robot excerpts更多实机片段
Watch the complete research video观看完整研究视频 3:49

Complete research video完整研究视频

3:49 · Full video完整视频

The full video includes all ten real trials, successes and failures. Mapping footage plays at 4× and simulated rollouts at 8×. Simulation evaluates ordered localization of the object and anchor; physical grasping and placement are evaluated on the real robot.完整视频包含全部 10 次真实实验及成功与失败案例。建图片段以 4 倍速、仿真片段以 8 倍速播放。仿真评估按顺序定位物体与锚点;物理抓取和放置在真实机器人上评估。

Download full video下载完整视频

01

Online 3DGS & conditional flow matching在线 3DGS 与条件流匹配

After the object moves, its original support region is empty; object appearance alone cannot identify this return location.物体被移动后,原支撑区域已空置,仅靠物体外观无法确定归位位置。

旧目标照片中的物体与原锚点,以及物体已被移动的当前场景
The original support region is empty in the current scene. Finding the object and recognizing its former context are both necessary.当前场景中的原始支撑区域已经空置。任务需要同时找到物体,并识别它原来所处的环境上下文。 Full-size figure查看原图
在线高斯输入、条件潜在流匹配模型,以及引导物体与锚点搜索的八通道恢复场
The online map and goal image condition a latent flow model. Its decoder produces a Restoration Field that drives next-best-view exploration and persistent target readout.在线地图和目标图像为潜在流模型提供条件,解码器生成恢复场,用于下一最佳视角探索与稳定目标读出。 Full-size figure查看原图
  1. 1

    Map what is visible在线表示可见空间

    Self-pose estimates are an input. RGB-D observations update an online 3D Gaussian map. Visibility-aware projection produces spatial triplanes; a frozen DINOv2 backbone represents the goal image.自位姿估计是系统输入。RGB-D 观测持续更新在线三维高斯地图。可见性感知投影生成空间三平面,冻结的 DINOv2 主干提取目标图像表示。

  2. 2

    Generate a Restoration Field生成恢复场

    Conditional flow matching jointly generates eight channels for scene state, search prior, support, object affinity and contextual-anchor affinity.条件流匹配联合生成八个通道,描述场景状态、搜索先验、支撑面、物体关联与上下文锚点关联。

  3. 3

    Search, pick, return搜索、拾取、归位

    The field guides exploration. Stable peaks in observed regions trigger object grasping, then anchor search and placement using manipulation primitives.恢复场引导探索,已观察区域中的稳定峰值触发物体抓取,随后进行锚点搜索,并通过操作基元完成放置。

02

Evidence for the design设计依据

A past photograph supplies appearance and context; the online map supplies the evolving spatial representation. The experiments test whether both inputs are used and where the ordered restoration task still fails.旧照片提供外观和上下文,在线地图提供持续更新的空间表示。实验考察模型是否使用这两种输入,以及有序恢复任务仍在哪个阶段失败。

The goal image and current map are both used目标照片与当前地图均参与定位

In Table 4, ordered localization success drops from 34.8% to 1.8% when goal images are shuffled and to 0.8% when the online triplane is set to zero. Removing contextual-anchor affinity gives 8.4%. These controls use the same held-out scenes and evaluation episodes. They support dependence on the goal and scene inputs in this setup; they do not compare 3DGS against every alternative map representation.表 4 中,打乱目标照片后,顺序定位成功率由 34.8% 降至 1.8%;将在线三平面置零后为 0.8%;移除上下文锚点关联后为 8.4%。这些对照使用相同的未见场景与评测任务,支持模型在该设置中对目标和场景输入的依赖,但并未比较 3DGS 与所有其他地图表示。

表 4 顺序定位的输入与上下文锚点消融
Table 4, manuscript p. 7; 15 held-out HM3D scenes. Ordered localization excludes grasping and placement.稿件第 7 页表 4;15 个未见过的 HM3D 场景。顺序定位不包含抓取和放置; Full-size figure查看原图

The empty anchor remains the observed bottleneck空置锚点仍是主要瓶颈

The same ten real trials yield eight object localizations, five anchor localizations and five physical restorations. The three counts are nested stages of one trial set. All five successful anchor localizations proceed to restoration, while repeated background structure and weak context can mislead anchor matching. The largest observed stage loss is therefore anchor search in this small sample; it does not establish that placement is generally solved.同一组十次实机试验得到八次物体定位、五次锚点定位和五次物理复原。这三个数是同一试验集中的嵌套阶段。五次成功锚点定位均进入并完成复原,而重复背景结构和弱上下文会误导锚点匹配。因此,在这个小样本中,最大的阶段损失发生在锚点搜索;这并不说明放置问题已被普遍解决。

同一十次实机试验的嵌套阶段计数:八次物体定位、五次锚点定位、五次复原
Table 3, manuscript p. 7. Five objects, two trials each; these are sequential counts from the same ten trials, not independent samples.稿件第 7 页表 3。五件物体,每件两次;这些是同一十次试验的顺序阶段计数,不是独立样本。 Full-size figure查看原图
03

Search and restoration results搜索与复原实验

34.8%

Ordered localization success顺序定位成功率

15 held-out HM3D scenes · simulation15 个未见过的 HM3D 场景 · 仿真

+9.0 pp

Gain over adapted GaussNav较适配 GaussNav 的提升

Full ordered task · percentage points完整顺序任务 · 百分点

5 / 10

Physical restorations物理复原成功次数

5 everyday objects · 2 trials each5 件日常物体 · 每件 2 次

Simulated success means locating the displaced object first and the original anchor second, within a 1 m arrival threshold; no grasping or placement is simulated. Real success additionally requires stable placement on the correct support surface with horizontal error <0.25 m. Final orientation is unconstrained.

仿真成功指先定位被移动物体,再定位原始锚点,到达阈值为 1 m;不模拟抓取和放置。真实成功还要求物体稳定放在正确支撑面上,水平误差 <0.25 m;不约束最终朝向。

Held-out scene comparison未见场景对比

15 held-out HM3D scenes · three categories and three seeds averaged, Table 215 个未见过的 HM3D 场景 · 三类物体与三随机种子均值,表 2
Method方法Object success物体成功率Anchor success锚点成功率Ordered task顺序任务成功率R-SPL
GaussNav67.1%31.1%25.8%0.118
IEVE59.6%25.9%12.4%0.061
Spot the Difference62.4%41.8%34.8%0.213

GaussNav and IEVE are sequentially adapted with object/anchor crops. Object and anchor rates are marginal stage scores, so anchor success may include visiting the anchor before the object. The full task requires the correct order. GaussNav retains higher object-only success.GaussNav 与 IEVE 使用物体及锚点裁图进行顺序适配。物体和锚点成功率为各阶段边际指标,因此锚点成功可能包含在找到物体前经过锚点。完整任务要求正确顺序,GaussNav 的仅物体成功率仍较高。

三个仿真实例,包括目标图像、搜索轨迹和恢复场随时间变化的数值
Qualitative bowl, cup and bottle examples. Time traces show field responses near policy targets, rather than experimental success percentages.碗、杯子和瓶子的定性示例。时间曲线表示策略目标附近的恢复场响应,不表示实验成功率。 Full-size figure查看原图

Ten real-world trials十次真实实验

An AgileX Ranger MINI 3 with a six-DoF PiPER arm performs the complete task. Across ten trials, the system localizes the object eight times and the anchor five times; all five successful anchor localizations lead to physical restoration. Onboard RTX 4080 Super handles mapping, field generation and grasp proposals; Intel NUC runs Nav2. Mapping updates at approximately 8 Hz. The learned model supplies spatial information for search and position readout; hand-designed next-best-view scoring, Nav2 navigation and Contact-GraspNet grasp proposals connect it to execution.AgileX Ranger MINI 3 与六自由度 PiPER 机械臂执行完整任务。在十次实验中,系统八次成功定位物体、五次成功定位锚点;五次成功锚点定位均完成物理复原。机载 RTX 4080 Super 负责建图、恢复场生成和抓取候选,Intel NUC 运行 Nav2。建图更新频率约为 8 Hz。学习模型提供用于搜索和位置读出的空间信息;手工设计的下一最佳视角评分、Nav2 导航与 Contact-GraspNet 抓取候选将其接入动作执行。

五件物体的全部十次真实实验,黑色面板表示失败后未到达的阶段
All ten trials are represented. Black panels mark unachieved stages after earlier failure. Contextual anchor localization is the main observed bottleneck.图中呈现全部十次实验。黑色面板标记前序失败后未到达的阶段,上下文锚点定位是主要瓶颈。 Full-size figure查看原图
04

Scope & limitations范围与局限

The method assumes one displaced rigid object and otherwise static scene geometry. Systematic ambiguity from identical objects or similar anchors is outside the evaluated scope. Large viewpoint changes and weak visual context can make anchor localization difficult.方法假设只有一个刚性物体被移动,其他场景几何保持不变。相同物体或相似锚点导致的系统性歧义不在已评估范围内。较大视角变化及较弱视觉上下文会增加锚点定位难度。

Restoration estimates position and relies on hand-designed next-best-view scoring. The real evaluation has ten trials. The approximately 8 Hz figure describes mapping updates, rather than an entire restoration or control cycle.恢复估计关注位置,并依赖手工设计的下一最佳视角评分。真实评估共十次。约 8 Hz 仅表示建图更新频率,不代表完整恢复或控制循环频率。

05

Conclusion结论

The work connects a past goal image to online 3DGS for ordered object-and-anchor localization. In ten real trials, anchor localization is the main observed bottleneck. Stronger contextual disambiguation and complete target-pose estimation remain open questions; current evaluation does not constrain final object orientation.这项工作将旧目标照片与在线 3DGS 联系起来,完成先物体、后原锚点的有序定位。在十次真实试验中,锚点定位是主要瓶颈。更可靠的上下文消歧与完整目标位姿估计仍是待解决的问题;当前评估不约束最终物体朝向。