Preprint 2026

IronManInformation-Constrained Video-Action
Learning for Robot Manipulation

Less is more.

Yuanshuo Zhang1Wenzhe Zhao2Zixing Lei1Bin Chen2Siheng Chen1*

1 Shanghai Jiao Tong University2 Joy Future Academy, JD

* Corresponding author

99.0%LIBERO success rate
79.1%LIBERO-Plus success rate
79.4%RoboTwin clean2clean success rate
12.1%RoboTwin clean2random success rate

Abstract

Video Action Models (VAMs) couple visual dynamics modeling with action generation for robot manipulation. However, video representations are not naturally suited to action generation, as exposing the action policy to excessive visual detail can impair its generalization ability. Therefore, we introduce IronMan (Information-constRained videO-actioN learning for robot MANipulation), a robust video-action learning framework built on the information bottleneck principle. The core principle of this framework is to impose information constraints that suppress irrelevant visual information while preserving action-relevant dynamics cues. IronMan employs a dynamics-aware bottleneck that distills noisy, entangled one-step video features into compact world representations. Extensive simulation and real-world experiments demonstrate strong in-distribution (ID) performance and out-of-distribution (OOD) robustness while maintaining efficient inference. IronMan achieves success rates of 99.0% on LIBERO and 79.4% on RoboTwin clean2clean, outperforming all the evaluated baselines. Under OOD shifts, IronMan achieves a success rate of 79.1% on LIBERO-Plus, exceeding the strongest baseline by 10.4 percentage points.

01 / The method

A compact interface from video to action.

Existing video–action interfaces allow the action policy to freely attend to irrelevant visual signals, increasing its sensitivity to visual details. IronMan instead uses an information-constrained bottleneck designed to retain the predictive dynamics needed for action while suppressing redundancy.

Motivation of IronMan
Motivation. Unlike existing IDM-style Video Action Models that fully denoise future frames (high latency) and expose the policy to redundant features, IronMan extracts one-step video features and distills them into compact, efficient world tokens through a dynamics-aware bottleneck.
IronMan architecture overview
Overview of IronMan. A Video DiT (initialized from Wan2.2-5B) models visual dynamics and provides one-step video features. The dynamics-aware bottleneck aggregates these features with a spatio-temporal transformer (spatial cross-attention + temporal self-attention) and maps them through a Gaussian head to a latent distribution, yielding compact world tokens. The Action DiT attends to world tokens and robot state tokens to generate action chunks via flow matching. Video decoding is optional and inactive during action inference. Training combines video and action flow-matching losses with a variational information bottleneck (KL) regularization on the world-token distribution.

02 / The evidence

Strong ID performance and OOD robustness.

In-Domain: LIBERO & RoboTwin clean2clean

ModelEm. PT.SpatialObjectGoalLongLIBERO Avg.RoboTwin c2c
π0.5✓98.898.298.092.496.970.7
X-VLA✓98.298.697.897.698.168.0
Abot-M0✓98.899.899.096.698.657.4
starVLA✗98.799.798.694.297.846.5
FastWAM✗98.2100.097.095.297.677.8
LingBot-VA✓98.599.697.298.598.5–
Motus✓96.899.896.697.697.7–
IronMan (ours)✗99.0100.097.899.099.079.4

In-domain success rates (%). Bold = best, underline = second best. Em. PT. denotes embodied pretraining. Without embodied pretraining, IronMan achieves the highest LIBERO average (99.0%) — including 99.0% on long-horizon LIBERO-Long — and 79.4% on RoboTwin clean2clean.

Out-of-Distribution: LIBERO-Plus & RoboTwin clean2random

ModelCameraRobotLanguageLightBackgroundNoiseLayoutOverallRoboTwin c2rLatency (ms)
DiT4DiT59.761.188.490.833.760.083.668.7–180
FastWAM16.843.768.177.251.536.559.549.21.9190
FastWAM-Joint41.365.591.086.453.856.279.267.45.2580
FastWAM-IDM37.767.290.392.754.256.679.267.99.3810
OpenWAM-Joint1.253.471.177.343.818.452.243.713.2429
IronMan (ours)64.475.992.793.464.478.383.779.112.1200

OOD success rates (%) on LIBERO-Plus (seven perturbation categories) and RoboTwin clean2random, with inference latency measured on an NVIDIA RTX 5090. IronMan ranks first across all seven perturbation categories, surpassing the strongest baseline by 10.4 percentage points overall, with the largest gains on sensor noise (+18.3 pp) and background textures (+10.2 pp) — while keeping latency comparable to the fastest baselines.

Bottleneck analysis & attention visualizations

Bottleneck Mechanism Analysis

Bottleneck size analysis
Bottleneck size Q×D. Smaller bottlenecks limit representation capacity; larger ones expose the action head to excessive irrelevant details.
Constraint strength analysis
Constraint strength β. Intermediate β performs best: weak constraints let irrelevant information pass; overly strong constraints discard control-relevant cues.

Attention Visualization

Attention heatmap comparison — ketchup task Attention heatmap comparison — moka pot task
Attention heatmaps under a background shift from LIBERO (ID) to LIBERO-Plus (OOD). IronMan focuses selectively on the gripper–object interaction region, whereas the baseline without the dynamics-aware bottleneck also attends broadly to task-irrelevant backgrounds — and fails.

03 / In the real world

Robustness in the real world

We evaluate IronMan on three real-world tasks across the AGIBOT G2 and ROBOTERA M7 platforms — insert flowers, arrange fruits, and organize parcels — testing spatial perception, fine-grained manipulation, and long-horizon execution. We collect 100 demonstrations per task and evaluate 10 trials per task and setting. Each task is evaluated under three settings: ID (training distribution), Background (strong tabletop texture distractions), and Layout (shifted object positions and orientations).

Real-world performance
Real-world performance under ID, Background, and Layout settings. IronMan achieves the highest average normalized score under ID (91.11) and Layout (66.11), and scores 56.67 under Background shifts, ahead of FastWAM (6.39) and below π0.5 (68.61). These are normalized task scores, not success rates.

Real-world rollouts on AGIBOT G2 and ROBOTERA M7. Videos autoplay muted.

Insert flowers · example 1
Arrange fruits · example 1
Organize parcels · example 1
Insert flowers · example 2
Arrange fruits · example 2
Organize parcels · example 2
Insert flowers · example 1
Arrange fruits · example 1
Organize parcels · example 1
Insert flowers · example 2
Arrange fruits · example 2
Organize parcels · example 2
Insert flowers · example 1
Arrange fruits · example 1
Organize parcels · example 1
Insert flowers · example 2
Arrange fruits · example 2
Organize parcels · example 2

04 / In simulation

Attention heatmap.

Rollouts under background shifts on LIBERO-Plus. In each video, the left panel shows IronMan and the right panel shows the baseline without the dynamics-aware bottleneck. The attention heatmap video on the right visualizes where each policy attends. These are selected qualitative examples; aggregate results are reported in the section above. After IronMan succeeds, its final frame is held while the baseline continues to its step limit.

Rollout · left: IronMan — success, right: baseline w/o bottleneck — failure
Attention heatmap · left: IronMan, right: baseline

“Pick up the ketchup and place it in the basket” (LIBERO-Object, background shift).

Rollout · left: IronMan — success, right: baseline w/o bottleneck — failure
Attention heatmap · left: IronMan, right: baseline

“Pick up the alphabet soup and place it in the basket” (LIBERO-Object, background shift).

Rollout · left: IronMan — success, right: baseline w/o bottleneck — failure
Attention heatmap · left: IronMan, right: baseline

“Put the yellow and white mug in the microwave and close it” (LIBERO-10, background shift).

Rollout · left: IronMan — success, right: baseline w/o bottleneck — failure
Attention heatmap · left: IronMan, right: baseline

“Turn on the stove and put the moka pot on it” (LIBERO-10, background shift).

05 / Reference

Cite this work

If you find IronMan useful for your research, please consider citing our work.

Download .bib
@misc{zhang2026ironman,
  title         = {IronMan: Information-Constrained Video-Action Learning for Robot Manipulation},
  author        = {Zhang, Yuanshuo and Zhao, Wenzhe and Lei, Zixing and Chen, Bin and Chen, Siheng},
  year          = {2026},
  eprint        = {2610.07961},
  archivePrefix = {arXiv},
  primaryClass  = {cs.RO},
  doi           = {10.48550/arXiv.2610.07961},
  url           = {https://arxiv.org/abs/2610.07961}
}