World models · latent intent · direct control

INTACT

Isomorphic Intent-to-Action Learning for Search-Free World Models

World models already learn what actions do. INTACT teaches the same representation which action realizes a requested latent change.

Junhan Sun1,4, Hao Zhao2,4,†, Guofeng Zhang1,3,†

1State Key Laboratory of CAD&CG, Zhejiang University   2AIR, Tsinghua University   3InSpatio   4RoboParty Lab

Watch the project film
INTACT in one loop. Synchronized failures and recoveries lead into the intent-to-action interface, E1–E5 correspondence, controlled shared-encoder results, and the paper's headline operating point. Demo identities and quantitative frames come from audited artifacts. 2K4K
Interactive evidence

Unified-task JEPA representation geometry, from initialization to outcome.

Follow the same fixed observations from PushT, Cube, Reacher, and TwoRoom through 7,075 measured training states. Compare complete collapse, partial collapse, and a well-formed outcome in independently fitted 2D and 3D t-SNE.

Effective rank–
Mean cosine–
Measured frame–
2D t-SNE
3D t-SNE
Preparing measured checkpoints…
Training state E0.00 step 0
Well-formed geometry

Loading trajectory metadata.

Complete E0–E5 trajectory Fixed observations · real snapshots · no latent interpolation

Measurement note. Every colored point tracks the same fixed observation through training. Each projected frame is robustly normalized to expose morphology, so absolute projected radius is not compared. t-SNE is qualitative evidence about local neighborhoods; effective rank and mean pairwise cosine are computed in the original high-dimensional latent space.

The idea

A world model should not need broad search to express control knowledge it already learned.

Forward latent models learn the conditional “given this action, what changes?” Goal-conditioned control then numerically inverts that knowledge with thousands of candidate sequences. This creates a representation-control asymmetry: action shapes the latent during training, but deployment begins again from generic proposals.

Every action-labelled transition already contains the missing supervision. It identifies a state-conditioned motion intent and the action that realizes it. INTACT exposes that relation as a first-class interface, trained end to end with the forward JEPA.

INTACT v31 teaser showing matched Cube control, representation geometry, direct success rate, latency, and action-family alignment
INTACT turns action-aligned latent intent into search-free control. The matched Cube audit, one-epoch macro result, latency study, and representation correlations are distinct evaluations and are not pooled.
1full-data epoch
95.61%Direct macro SR
0Direct candidates
3.9–4.8 msplanner latency
23.44×fewer verification samples
Method

Two intent families. One conditional action law.

A real successor provides a physically grounded intent that is unavailable before acting. A future goal provides the condition available at deployment, but it is not a one-step successor. INTACT gives both families the same typed predictor graph while retaining their different causal roles.

Physical successor intent mtlocal = zt+1 − zt successor attached
Deployable goal intent mtgoal = sg(zg) − zt goal detached

The predictor receives the same four-slot grammar in both calls: current state, first-order intent, the matched state-intent interaction, and previous-action context. There is no pointwise loss forcing the local and goal endpoints to coincide.

INTACT inference and deployment pipeline
Shared-encoder INTACT. Four visual domains update one encoder, while task-specific head pairs preserve domain dynamics. Direct control emits an action chunk; world rollout is retained for recurrent replanning and optional verification.
Why the name matters

Isomorphic between calls. Intact across representation and control.

Graph isomorphism

Local and goal calls use the same parameter-labelled backbone-input graph.

Family isomorphism

Supported motion-intent families correspond through the action-law semantics induced by one predictor.

Representation intactness

Action-effective information survives the RGB-to-latent bottleneck without retaining every nuisance.

Conversion intactness

The relation between latent intent families and their action families survives direct readout.

Results

One epoch makes broad action search optional.

Task-specific INTACT is trained end to end for one full-data epoch. Direct evaluates no candidate sequence and makes no terminal latent-cost call. Guarded A keeps the coherent Direct plan as its center and verifies only a small local neighborhood.

TaskDirect · 0Guarded A · 384Published LeWM
PushT87.44 ± 1.2691.56 ± 0.6996.0 ± 4.0
Cube100.00 ± 0.0099.67 ± 0.3374.0 ± 3.0
Reacher97.33 ± 0.0097.56 ± 0.5186.0 ± 5.0
TwoRoom97.67 ± 1.2097.56 ± 1.2687.0 ± 2.5
Macro95.61 ± 0.5996.58 ± 0.4485.75

INTACT values average three independently trained models per task, each evaluated with three 100-episode seeds. Published LeWM retains its own 10-epoch CEM protocol and is landscape context rather than a paired significance control.

Search-free control

Controller first, search second.

One epoch turns a latent goal into a coherent action plan. Search becomes a small local verifier, not the controller itself.

01
Control interface comparisonTask-specific checkpoints · published LeWM uses CEM
Published LeWM · CEM INTACT Direct INTACT Guarded A
96.087.4+4.191.6
PushT
74.0100−0.399.7
Cube
86.097.3+0.297.6
Reacher
87.097.7−0.197.6
TwoRoom
02
Verify locallySame PushT goal-displacement checkpoints
Direct87.4% Guarded A91.6% Actor-on CEM93.9%
0 → 384 → 9,000candidate sequences per solve
Controller first, search second. Direct is the primary zero-search interface; Guarded A is a bounded local verifier, while actor-on CEM remains the broad-search reference.
Unified learning

One visual encoder, four control domains.

At E5, Goal-displacement INTACT reaches 91.22 ± 0.51% Direct macro SR with one encoder shared across PushT, Cube, Reacher, and TwoRoom. The matched shared LeWM baseline reaches 66.17 ± 2.67% using CEM 300×30.

Disabling every INTACT action head still raises actor-off pure-CEM macro to 70.08 ± 1.13%. Representation shaping is measurable before direct readout, while restoring the learned interface supplies most of the closed-loop gain.

Shared encoder · E5

One controlled matrix, rendered as data.

Success rate (%) · three training seeds

MethodPushTCubeReacherTwoRoomMacro
LeWMCEM 300×3074.667.383.139.766.17
Inverse onlyDirect36.167.690.678.668.19
Waypoint intent onlyDirect58.9100.065.972.174.23
Goal intent onlyDirect81.8100.088.769.284.92
Waypoint INTACTDirect71.299.058.277.276.40
Goal-displacement INTACTDirect · 0 candidates86.1100.097.281.691.22
Goal-displacement INTACTGuarded A · 128×385.499.297.080.490.53
Controlled E5 shared-encoder result. Goal-displacement INTACT improves all four tasks over matched jointly trained LeWM and exceeds its goal-intent-only control by 6.31 macro points.
Theory to measurement

Control follows action-family correspondence, not latent spread alone.

At a fixed current state, two endpoint conditions are equivalent when they induce the same expert action law. This conditional action quotient predicts that local and global agreement between predicted and expert action families should track control.

0.968predicted–expert kNN ↔ Direct SR across 15 eligible checkpoints
0.988predicted–expert linear CKA ↔ Direct SR
0.983pointwise action R² ↔ Direct SR

The v31 teaser shows the matched 15-checkpoint goal-displacement cohort used in the audited readout above. Across this cohort, predicted–expert kNN overlap, linear CKA, and pointwise action R² correlate with Direct SR at r = 0.968, 0.988, and 0.983, respectively; SRS reaches r = 0.940.

Checkpoint geometry

Correspondence becomes predictive during training.

Through epochE5
Macro SR91.22%
Mean CKA0.505
Pearson r+0.988
Cumulative n = 15 checkpoints
Native, inspectable measurement. Each epoch adds three independently trained Goal-INTACT checkpoints. Select E1–E5 to inspect Direct macro SR, mean linear CKA, and their cumulative Pearson relation; the three-point E1 estimate is marked provisional.
Measured E5 correspondence

Intent and action families evolve together.

E5 · standardized PCA-3D
Preparing measured E5 coordinates…
PushT action correspondence
Action R²—
Predicted–expert CKA—
kNN overlap—
Interactive E5 measurement. Shared intent uses one projection across all four tasks. Task views show predicted actions, expert actions, and their paired samples in a task-specific frozen basis. Drag to rotate and scroll to zoom; no raw latent or checkpoint is shipped to the browser.
Action-family alignment diagnostics correlated with direct success rate
Intent–action relation, rather than task clustering, tracks control. Higher rank or cleaner task clusters do not certify a useful control representation.
Open research exchange

Keep the world-model conversation moving.

Join an open community around JEPA, LeWM, INTACT, representation learning, and efficient robot control. Follow new results, compare reproductions, challenge assumptions, and bring early ideas.

Open the Community page
Permanent QR code for the INTACT World Model Community page Permanent entryCurrent WeChat invitation lives behind this page.
Reproducible release

From the paper to runnable training and evaluation.

PaperMethod + audited results

Definitions, controlled comparisons, and machine-readable result records.

CodeTraining + evaluation

Task-specific and shared-encoder recipes with explicit inference modes.

ModelsVerified checkpoints

Published separately with manifests and SHA-256 verification.

ReproduceLocked environment

Versioned dependencies, data checks, smoke runs, and checkpoint-compatible evaluation.