Activation-Aligned Residual MPC for Robot Manipulators
Learned closed-loop dynamics and asynchronous residual CEM-MPC for position-controlled robot arms.
Research case study
I led the design, implementation, and end-to-end experimental development of a learned-dynamics residual MPC framework for position-controlled manipulators.
Overview
This project studies position-controlled manipulators whose actuator response and planner latency make a newly optimized command stale before it reaches the robot. The controller combines a learned closed-loop dynamics model with bounded residual MPC and an activation-aligned asynchronous execution protocol. The evidence spans an ABB IRB 2400 primary MuJoCo study, an independently trained UR5e replication, and a hardware-specific SO-ARM101 evaluation.
Key Results
- 61.2% joint RMSE reduction 0.910° → 0.353°
- 58.3% FK-derived TCP RMSE reduction 8.912 mm → 3.715 mm
- 0 control deadline misses across 63 hardware trials
These headline values summarize the SO-ARM101 protocol: the two error reductions compare NN-MPC with Direct IK across 54 held-out trials, while the timing count covers all 63 hardware trials.
Problem
A position actuator accepts an absolute joint reference, but the arm does not arrive there instantaneously. While CEM is optimizing a residual sequence, the physical or simulated plant continues moving under the previous packet and the IK nominal continues advancing. Replaying a plan in its launch-time frame therefore mixes old state, old reference, and current execution time.
The design challenge was to preserve the meaning of a residual command across planner delay without blocking the fast control loop.
Method
Learned closed-loop dynamics
A GRU predicts the next joint state from aligned histories of measured position, velocity, and the absolute references already sent to the position-controlled plant. It models the actuator-and-command loop, not free rigid-body dynamics.
Residual predictive control
CEM searches bounded corrections around a continuously generated damped-least-squares IK nominal. The command remains q_cmd = q_nom + r, so planning improves a task-derived reference instead of replacing it.
Activation-aligned async execution
The planner advances state, complete recurrent context, and the reference window to the expected activation time. A timestamped packet is then age-indexed and reanchored to the current IK nominal before projection and execution.
Learned Dynamics Model
The learned model represents the effective closed-loop response of the position-controlled robot: how measured joint position and velocity evolve after an absolute joint reference is issued. Inside CEM, its predictions are rolled forward autoregressively to score candidate residual sequences. It does not estimate free rigid-body or torque dynamics.
The ABB and UR5e simulation models use the same compact six-DOF architecture but independently collected datasets, checkpoints, and normalizers. This keeps the controller design comparable without treating the UR5e study as shared-model transfer.
- Context
- history length 16
- Input token
-
[q, dq, q_ref]· 18 values - Recurrent core
- single-layer GRU · hidden size 256
- Prediction head
-
256 → 256 → 6· SiLU - Target
- 6-D joint-velocity increment
Δdq - Model size
- 279,302 trainable parameters
- Data split
- 90/10 episode-disjoint train/validation
- Training objective
- one-step + 20-step rollout loss · weight 0.025
σu=0.8, rollout 019). Across this 200-step sequence, the six joint trajectories contain multiple reversals rather than a single smooth decay. Blue is MuJoCo ground truth and dashed orange is the independently trained GRU prediction. This visual is model-prediction evidence, not closed-loop tracking evidence; the aggregate values below use complete-history windows from every rollout in each group.| Robot | Excitation | 1-step q RMSE | 20-step q RMSE | 20-step dq RMSE | Divergent windows |
|---|---|---|---|---|---|
| ABB | σu=0.5 | 0.670 mrad | 4.852 mrad | 34.472 mrad/s | 0% |
| ABB | σu=0.8 | 1.069 mrad | 6.876 mrad | 51.398 mrad/s | 0% |
| UR5e | σu=0.5 | 0.132 mrad | 1.334 mrad | 2.745 mrad/s | 0% |
| UR5e | σu=0.8 | 0.160 mrad | 1.436 mrad | 3.023 mrad/s | 0% |
Each value is the mean over 10 held-out rollouts at the stated excitation level. The 20-step result pools prediction steps 1–20 within windows that already contain the full 16-token ground-truth history; it is not a terminal-error measurement or a closed-loop controller metric.
Simulation Evaluation
The ABB IRB 2400 is the primary MuJoCo platform for nominal comparisons, mechanism isolation, delay sweeps, and robustness tests. The UR5e is a nominal within-robot replication with its own collected data, normalizer, GRU checkpoint, IK references, limits, and delay calibration. Its model was trained independently; the experiment is not a shared-model or zero-shot transfer test.
Hardware Pipeline
The physical study uses a separately collected SO-ARM101 dataset to train a robot-specific GRU. Deployment retains joint, velocity, acceleration, braking, encoder-quantization, startup, and homing safeguards while the planner runs asynchronously beside the 30 Hz command loop. The planner has ±1° residual authority per joint.
SO-ARM101 data collection for the robot-specific learned dynamics model—not a controller-performance demo. Watch the data-collection video on YouTube
- Control rate
- 30 Hz
- Planner
- H=6 · CEM 128x2
- Model context
- history=16
- Tracking objective
- J = C_q · tracking-dominant
- Residual authority
- ±1° requested residual
- Controllers
- Direct / Preview6 / NN-MPC
- Protocol
- 54 held-out / 63 total
My Contributions
- Built the data, training, and validation pipeline for history-conditioned closed-loop manipulator dynamics.
- Designed and implemented bounded residual CEM-MPC around continuously generated IK references.
- Developed the activation-aligned packet protocol, including activation forecasting, age indexing, execution-time reanchoring, projection, feedback, and fallback behavior.
- Built the reproducible ABB and UR5e simulation workflows and the safety-gated SO-ARM101 experiment and evidence pipeline.