Activation-Aligned Residual MPC for Robot Manipulators

Learned closed-loop dynamics and asynchronous residual CEM-MPC for position-controlled robot arms.

Research case study

I led the design, implementation, and end-to-end experimental development of a learned-dynamics residual MPC framework for position-controlled manipulators.

ABB IRB 2400 complete figure-eight ThreadedASAP MuJoCo rollout with desired IK and actual TCP paths from initial hold through final hold
ABB IRB 2400 complete figure-eight ThreadedASAP MuJoCo rollout. Cyan shows the desired IK path and yellow shows the actual TCP path across all 3,607 execution steps, from initial hold through final hold.
UR5e complete figure-eight MuJoCo rollout with desired IK and actual TCP paths from initial hold through final hold
UR5e complete figure-eight MuJoCo rollout. Cyan shows the desired IK path and yellow shows the actual TCP path across all 1,405 steps of the independently trained and calibrated ThreadedAsync run, from initial hold through final hold.
Period
May 2026 – Aug 2026
Role
Lead Developer
Platforms
ABB IRB 2400 · UR5e · SO-ARM101

Overview

This project studies position-controlled manipulators whose actuator response and planner latency make a newly optimized command stale before it reaches the robot. The controller combines a learned closed-loop dynamics model with bounded residual MPC and an activation-aligned asynchronous execution protocol. The evidence spans an ABB IRB 2400 primary MuJoCo study, an independently trained UR5e replication, and a hardware-specific SO-ARM101 evaluation.

Key Results

  • 61.2% joint RMSE reduction 0.910° → 0.353°
  • 58.3% FK-derived TCP RMSE reduction 8.912 mm → 3.715 mm
  • 0 control deadline misses across 63 hardware trials

These headline values summarize the SO-ARM101 protocol: the two error reductions compare NN-MPC with Direct IK across 54 held-out trials, while the timing count covers all 63 hardware trials.

Problem

A position actuator accepts an absolute joint reference, but the arm does not arrive there instantaneously. While CEM is optimizing a residual sequence, the physical or simulated plant continues moving under the previous packet and the IK nominal continues advancing. Replaying a plan in its launch-time frame therefore mixes old state, old reference, and current execution time.

The design challenge was to preserve the meaning of a residual command across planner delay without blocking the fast control loop.

Method

01

Learned closed-loop dynamics

A GRU predicts the next joint state from aligned histories of measured position, velocity, and the absolute references already sent to the position-controlled plant. It models the actuator-and-command loop, not free rigid-body dynamics.

02

Residual predictive control

CEM searches bounded corrections around a continuously generated damped-least-squares IK nominal. The command remains q_cmd = q_nom + r, so planning improves a task-derived reference instead of replacing it.

03

Activation-aligned async execution

The planner advances state, complete recurrent context, and the reference window to the expected activation time. A timestamped packet is then age-indexed and reanchored to the current IK nominal before projection and execution.

Architecture of the learned residual CEM planner, timestamped packet, nominal-frame reanchoring, safety projection, and position-controlled plant
Method architecture. Planning forecasts the activation state and publishes residuals rather than cached absolute commands; execution retrieves the age-indexed residual and reanchors it to the current nominal.

Learned Dynamics Model

The learned model represents the effective closed-loop response of the position-controlled robot: how measured joint position and velocity evolve after an absolute joint reference is issued. Inside CEM, its predictions are rolled forward autoregressively to score candidate residual sequences. It does not estimate free rigid-body or torque dynamics.

The ABB and UR5e simulation models use the same compact six-DOF architecture but independently collected datasets, checkpoints, and normalizers. This keeps the controller design comparable without treating the UR5e study as shared-model transfer.

Context
history length 16
Input token
[q, dq, q_ref] · 18 values
Recurrent core
single-layer GRU · hidden size 256
Prediction head
256 → 256 → 6 · SiLU
Target
6-D joint-velocity increment Δdq
Model size
279,302 trainable parameters
Data split
90/10 episode-disjoint train/validation
Training objective
one-step + 20-step rollout loss · weight 0.025
UR5e high-variation held-out joint trajectories comparing MuJoCo ground truth with learned GRU predictions
High-variation held-out open-loop UR5e validation trace (σu=0.8, rollout 019). Across this 200-step sequence, the six joint trajectories contain multiple reversals rather than a single smooth decay. Blue is MuJoCo ground truth and dashed orange is the independently trained GRU prediction. This visual is model-prediction evidence, not closed-loop tracking evidence; the aggregate values below use complete-history windows from every rollout in each group.
Robot Excitation 1-step q RMSE 20-step q RMSE 20-step dq RMSE Divergent windows
ABB σu=0.5 0.670 mrad 4.852 mrad 34.472 mrad/s 0%
ABB σu=0.8 1.069 mrad 6.876 mrad 51.398 mrad/s 0%
UR5e σu=0.5 0.132 mrad 1.334 mrad 2.745 mrad/s 0%
UR5e σu=0.8 0.160 mrad 1.436 mrad 3.023 mrad/s 0%

Each value is the mean over 10 held-out rollouts at the stated excitation level. The 20-step result pools prediction steps 1–20 within windows that already contain the full 16-token ground-truth history; it is not a terminal-error measurement or a closed-loop controller metric.

Simulation Evaluation

The ABB IRB 2400 is the primary MuJoCo platform for nominal comparisons, mechanism isolation, delay sweeps, and robustness tests. The UR5e is a nominal within-robot replication with its own collected data, normalizer, GRU checkpoint, IK references, limits, and delay calibration. Its model was trained independently; the experiment is not a shared-model or zero-shot transfer test.

Representative ABB end-effector paths in which ThreadedASAP stays closer to the desired path while NaiveDelayed shows larger deviations
Representative ABB task-space tracking. ThreadedASAP stays closer to the desired path while NaiveDelayed shows larger deviations in this representative case; this single plot is not an aggregate comparison.
ABB planner-delay sweep for circle and fast-ellipse references comparing naive, alignment-only, reanchored, and full residual MPC variants
ABB fixed-delay component sweep. Alignment without nominal-frame reanchoring degrades sharply at longer delays, while the reanchored variants remain in the low-error region across the tested delays.

Hardware Pipeline

The physical study uses a separately collected SO-ARM101 dataset to train a robot-specific GRU. Deployment retains joint, velocity, acceleration, braking, encoder-quantization, startup, and homing safeguards while the planner runs asynchronously beside the 30 Hz command loop. The planner has ±1° residual authority per joint.

SO-ARM101 data collection for the robot-specific learned dynamics model—not a controller-performance demo. Watch the data-collection video on YouTube

Control rate
30 Hz
Planner
H=6 · CEM 128x2
Model context
history=16
Tracking objective
J = C_q · tracking-dominant
Residual authority
±1° requested residual
Controllers
Direct / Preview6 / NN-MPC
Protocol
54 held-out / 63 total
SO-ARM101 formal representative result with hardware setup, desired and FK-derived TCP paths, and encoder joint RMS error for three controllers
Representative formal-protocol SO-ARM101 block. The panels show the five-joint setup, desired and encoder-state FK TCP paths, and encoder joint RMS error for Direct IK, fixed Preview6, and NN-MPC.

My Contributions

  • Built the data, training, and validation pipeline for history-conditioned closed-loop manipulator dynamics.
  • Designed and implemented bounded residual CEM-MPC around continuously generated IK references.
  • Developed the activation-aligned packet protocol, including activation forecasting, age indexing, execution-time reanchoring, projection, feedback, and fallback behavior.
  • Built the reproducible ABB and UR5e simulation workflows and the safety-gated SO-ARM101 experiment and evidence pipeline.

Scope and Evidence