Closed-Loop Visual Tracking and Robotic Grasping
An integrated Piper eye-in-hand RGB-D grasping system and a controlled study of when to commit a moving target pose.
Piper manipulator · ROS 2 · RGB-D · MoveIt 2
The system sees a target through a camera on its wrist, updates its 3D pose, and plans a grasp. The research question is when that changing pose is stable enough to commit to a motion plan.
- 150canonical simulation trials120 controlled + 30 simulated RGB-D
- 10/10gated RGB-D task success5/5 static + 5/5 move-stop
- 5/10 vs 0/10snapshot vs continuous tracking10 trials per policy · 5 static + 5 move-stop
When should the robot commit to a target pose?
A wrist-mounted RGB-D camera can update the target continuously, but the arm must eventually plan against one pose. Committing too early can leave the gripper at an old location; insisting on a fresh pose during every step can also make a valid plan expire. I studied that choice inside an end-to-end Piper pipeline, from perception to grasp verification.
The shared system segments the target in RGB, fuses the mask with registered depth, transforms the 3D estimate through the hand–eye calibration into base_link, and passes a target to MoveIt 2. The arm plans PREGRASP, GRASP, and LIFT; task success requires lift and hold evidence, not merely a valid trajectory.
One observation, from pixels to a planner target
These three images are from the same successful gated replay at ROS time 25.047 s. The frame identifiers and timestamps match exactly; the overlay is the actual segmentation output, rather than a redrawn illustration.



At the nearest telemetry sample, ROS time 25.000 s, depth fusion marked the observation valid and fresh. The estimated target in base_link was approximately (0.425, 0.000, 0.040) m. This position comes from the nearby state record; it is not claimed to have the exact image timestamp. The frozen segmentation model card reports 97.37% foreground mIoU on a 37-image held-out split from the recorded collection setup.
Three ways to choose the pose
Commit once
Take an observation and keep that target pose while planning and executing. It works when the target remains at the observed location, but can become stale when the cube moves.
Keep updating
Replan from fresh observations. In the controlled ground-truth track it handled move-stop in 20/20 trials. In the simulated RGB-D track, all ten task trials stopped at final task-level planning after the fresh, low-drift PREGRASP conditions were not met; that result is specific to this configuration.
Wait, then commit
Collect a short run of fresh target estimates, check their spatial spread and duration, then freeze a stable pose for execution. The gate is one policy in the same pipeline, not a separate perception or grasping system.
GRASP stage, with no verified lift. This 3×-accelerated configuration replay is illustrative and is not part of the 150 canonical trials.Two evaluation tracks, separate denominators
The frozen Phase 20 audit includes 120 controlled ground-truth trials (20 per policy–scenario cell) and 30 simulated semantic RGB-D trials (5 per cell). Both tracks compare static and move-stop targets under the same task-level success definition. The chart reports counts directly; the smaller RGB-D cells are not pooled with ground-truth trials.
In simulated RGB-D, gated succeeded in all five static and all five move-stop trials. Snapshot's five move-stop failures occurred at grasp. RGB-D tracking's ten task failures occurred at final task-level planning, despite initial plans being available in all ten cases. These observations identify where this frozen system failed; they do not show that continuous tracking cannot handle moving targets in general.
My contribution and the shared foundation
The Piper/RGB-D grasping foundation was collaborative work already in place. Yue Zhang proposed the continuous-tracking and stability-gated research direction and leads the manuscript; Wei Liu provided the Piper and RGB-D platform and collaborated on the physical system. I led ROS 2/Gazebo/MoveIt 2 simulation development and system integration, connecting wrist RGB-D perception, 3D target localization, collision-aware planning, and grasp-and-lift execution. I implemented the target policies and simulated RGB-D evaluation path, and led benchmark and audit tooling, quantitative analysis, and grasp-physics checks. I contributed to physical-robot deployment; the hardware work was collaborative. The research repository documents the implementation and provenance.
Physical demonstration and evidence boundary
This physical Piper video is a qualitative hardware demonstration. It is not included in the 150 canonical simulation trials and does not provide a repeated-trial hardware success estimate.
Replay and figure provenance
View replay logs, checksums, and audit sources
Both Gazebo videos reproduce the final Phase 20 simulated semantic RGB-D configuration for move-stop seed 42. The gated replay logged TRIAL_FINISHED with task success; the matched snapshot replay logged task failure at grasp. RGB, mask, and overlay carry the identical camera timestamp of 25.047 s in the gated replay. The videos are accelerated for viewing. Media provenance and checksums , gated replay events, snapshot replay events, and the research repository make the source and evaluation boundary inspectable.