Four ways to condition a visuo-tactile world model

Same 30-episode training split, same 100 epochs, same optimizer, same seed — four different conditioning architectures, evaluated on ten identical held-out windows.

Diffusion-Forcing Mixture-of-Transformers over Wan2.2 latents (48 ch, 4 latent frames, 16×16). Every number below comes from last.ckpt at step 117 199 — the split warns that val aliases test, so selecting best-epoch=* would be selecting on the very episodes reported here.

The task

Each window is 16 fps6 frames → 4 latent frames. The first 2 latent frames are given as clean context; the model generates the other 2 from the action alone. Decoded through the Wan2.2 VAE that is 13 pixel frames, of which frames 0–4 come from the context and frames 5–12 are generated. The strips at the bottom outline the generated frames in red.

All four models are the same trunk: 12 blocks, hidden 768, patch (1,2,2), temporal-causal joint attention, cross-first block order, per-frame independent noise levels per modality. They differ only in how the action reaches the tokens.

The four configurations

mot_df_vt_pixflownew — the proposal

Both experts are conditioned by a frozen pixel-space Stage-1 flow concatenated into their latent, plus FiLM from the raw action vector. Action cross-attention, the analytic ActionImage, the actionflow encoder and the latent-flow probe loss are all gone.

video expert
pixview_D flow (2 ch, camera-crop px) concatenated into the 48-ch visual latent  •  FiLM from action[0:18], the camera-projected (u, v, z) of both sensors
tactile expert
pixc5_D flow (2 ch, gel px), one per sensor, concatenated into each 48-ch tactile latent  •  FiLM from action[18+9s : 27+9s], that sensor's own raw 9-D SE(3) step
contact force
Yes — via pixc5_D only. Its 5-channel input map is the gel frame: ch2 is press depth along the true gel normal and ch3-4 are the contact force on that same axis (log1p(N) + a contact flag). Force therefore reaches the network on exactly one path and can only influence the tactile flows; every video-branch input is force-free.
data streams
stage1_flow_k5.pt (cached frozen flows, already on the 16×16 latent grid)
parameters
182.61 M
action_ctx: false
pixflow:      {enable: true, video: true, tactile: true, share_tactile: true}
vector_film:  {enable: true, hidden: 256}
action_image: {enable: false}
action_motion:{enable: false}
video_motion: {enable: false}
pred_flow_loss:{enable: false}

mot_df_vt_vecfilmnew — the control

Byte-identical to pixflow except pixflow.enable: false. The dataloader still emits the flow stream, so the two runs see the same batches in the same order and only the model differs. It exists because pixflow removes four things and adds two at once — without this control, nothing in its result could be attributed to the flow.

video expert
FiLM from action[0:18]. No flow.
tactile expert
FiLM from each sensor's raw 9-D SE(3) step. No flow.
contact force
No. Force entered pixflow only through pixc5_D, and that encoder is not read here — so this run is the force-free twin as well as the flow-free one.
data streams
same file loaded, model ignores it
parameters
182.60 M
action_ctx: false
pixflow:      {enable: false}      # the only difference
vector_film:  {enable: true, hidden: 256}
action_image: {enable: false}
action_motion:{enable: false}
video_motion: {enable: false}
pred_flow_loss:{enable: false}

mot_df_vt_gelforcebaseline

The previous best recipe. Conditioning is spread across three mechanisms: per-modality action cross-attention, a dense analytic ActionImage on the video expert, and the frozen VAE-latent actionflow encoder injected into the tactile expert. A frozen latent-flow probe adds an auxiliary loss.

video expert
18-D cross-attention on action[0:18]  •  ActionImage: 6 Gaussian channels (2 sensors × pos/normal/up) rasterized on the 16×16 latent grid, σ = 0.8 cells, injected as a zero-init token bias
tactile expert
18-D cross-attention on action[18:36]  •  frozen amx_gel5_s2 motion feature (128 ch) projected zero-init onto the future tactile tokens
contact force
Yes — through amx_gel5_s2, the VAE-latent gel-frame encoder, fed the same 5-channel [dx, dy, dz, log1p(N), contact] map on the 16×16 grid.
data streams
tactile_action_map_gel_fps6.pt (5 ch), tactile_inverse_action_flow_gel_fps6.pt, tactile_flow6_lat16.pt
parameters
241.90 M (240.83 M trainable + 1.06 M frozen)
action_ctx: true                     # (default)
action_image:  {enable: true, sigma_cells: 0.8}
action_motion: {enable: true, checkpoint: amx_gel5_s2/best.pt,
                freeze: true, inject_to_dit: true}
pred_flow_loss:{enable: true, checkpoint: latent_flow_probe/best.pt,
                weight: 0.05}

mot_df_vt_gelforce_vmxbaseline + view motion

gelforce plus the frozen view motion encoder vmx_E — the video-side twin of actionflow. Its 128-ch feature is projected onto the video tokens and its predicted dense flow is appended to the ActionImage, making that raster learned and state-dependent rather than purely analytic.

video expert
everything gelforce has, plus vmx_E's motion feature as a second token bias and its 2-ch predicted flow appended to the ActionImage (6 → 8 channels); also consumes the sparse camera-frame motion raster action_map_view
tactile expert
identical to gelforce
contact force
Yes, identical to gelforce — vmx_E is video-side and force-free.
data streams
as gelforce, plus view_motion_map: true
parameters
243.11 M (241.23 M trainable + 1.87 M frozen)
action_ctx: true
action_image:  {enable: true, sigma_cells: 0.8}
action_motion: {enable: true, checkpoint: amx_gel5_s2/best.pt, freeze: true}
video_motion:  {enable: true, checkpoint: vmx_E/best.pt,
                freeze: true, inject_to_dit: true, flow_to_action_image: true}
pred_flow_loss:{enable: true, weight: 0.05}
Why the concatenation is a token bias. A convolution is linear in its input channels and the patch and stride are identical, so Conv[Wz|Wc](concat([z,c])) = ConvWz(z) + ConvWc(c). Adding a projected flow to the tokens is concatenating those channels into the patch embedding — exactly, not approximately (asserted numerically to 1.07e-6). Widening the expert’s own Conv3d(in_ch, …) would have been the wrong move: in_ch also sizes the output head, so it would silently change the diffusion variable from 48 to 50.

What conditions what

pathwaypixflowvecfilmgelforcegelforce+vmx
action cross-attention (per modality)——✓✓
analytic ActionImage → video tokens——✓✓ (+vmx flow)
frozen actionflow (VAE-latent) → tactile tokens——✓✓
frozen view-motion encoder → video tokens———✓
frozen pixel-space flow concatenated✓ both———
FiLM from the raw action vector✓ both✓ both——
latent-flow probe auxiliary loss——✓ 0.05✓ 0.05

The raw action vector is a real gap the middle two columns share: with action_ctx: false and no rasters, an earlier FiLM-only variant left action[18:36] — the 18-D tactile SE(3) step — reaching the network nowhere at all. vector_film restores it, standardized with the model’s own action statistics (the raw 9-D is metres ~1e-2 next to Rot6D components near 1, which a zero-init MLP head cannot cope with).

Training result — validation loss

configval_loss_videoval_loss_tactilebest epoch
pixflow0.06823 ± 0.000240.03452 ± 0.0001057
vecfilm0.06798 ± 0.000230.03459 ± 0.0001063
gelforce0.06974 ± 0.000200.04635 ± 0.0001395
gelforce+vmx0.07048 ± 0.000110.04583 ± 0.0001773

Mean ± sd over the last 10 epochs. Total val_loss is not comparable across the families: the two baselines carry ≈0.15 of frozen-probe loss inside it. The two components above are.

The tactile gap here is a confound, not a result. Every run that has the latent-flow probe enabled sits at 0.045–0.046 tactile val loss (actionfilm 0.04491, actionimage 0.04524, actionflow 0.04562, aimotion 0.04539) and both new runs have it off. That auxiliary term backpropagates into the tactile expert, so removing it is what moved the number — the conditioning change did not. The pixel PSNR below is the metric that is actually comparable.

Held-out result — decoded pixel PSNR

per-window PSNR for all four configs
Every config on every window. Tactile PSNR swings ±6 dB between windows — contact events dominate it — which is why the per-window paired difference below is the readable statistic, not the spread of the means.
windowpixflowvecfilmgelforcegelforce+vmx
camera view — pixel PSNR, dB
005:024.5624.2821.0124.94
005:3422.2722.4521.4923.42
005:6824.3024.6722.3023.50
005:10223.5523.5922.3723.67
005:13623.3223.0421.7123.64
006:024.0924.0424.2924.12
006:15024.4525.1823.2824.28
006:30024.9524.8524.8124.52
006:45023.7924.1122.5623.68
006:60023.7123.5022.5523.48
mean23.9023.9722.6423.93
tactile left — pixel PSNR, dB
005:039.7941.4240.7937.91
005:3439.3840.2938.9139.40
005:6842.3343.1842.1540.83
005:10239.6039.1737.8137.54
005:13650.5151.2247.9848.59
006:051.7452.0748.9449.30
006:15042.5443.2240.9739.96
006:30051.3651.4348.2348.54
006:45042.8042.8540.4240.57
006:60050.6750.9347.9948.23
mean45.0745.5843.4243.09
tactile right — pixel PSNR, dB
005:051.2248.2343.6246.11
005:3452.5552.5849.9850.47
005:6848.2548.4546.9446.95
005:10252.7552.6350.0450.35
005:13638.5938.7037.1436.68
006:051.2951.3049.1649.70
006:15041.6741.4340.2340.79
006:30039.3439.6436.5738.44
006:45045.2645.3044.7144.63
006:60040.2840.0838.0539.45
mean46.1245.8343.6444.36

Paired per-window differences

Because every config saw the same ten windows, the difference can be taken window by window instead of comparing two noisy means.

comparisonstreammean Δmedian Δsdwins
pixflow − vecfilmcamera view-0.07 dB+0.01 dB0.335/10
pixflow − vecfilmtactile left-0.51 dB-0.50 dB0.581/10
pixflow − vecfilmtactile right+0.28 dB-0.02 dB0.964/10
new pair − baseline paircamera view+0.65 dB+0.57 dB0.588/10
new pair − baseline pairtactile left+2.07 dB+2.37 dB0.7910/10
new pair − baseline pairtactile right+1.98 dB+1.80 dB1.1610/10

Reading

Why one might have expected the flow to help, and a guess at why it did not: the tactile encoder pixc5_D is weak — held-out direction cosine 0.645 and magnitude ratio 0.44, i.e. it recovers less than half the true flow length. A separate study also found that a perfect, dense, correctly-splatted flow explains only ~12.5% of the frame-to-frame change on a gel image, because a vision-based tactile sensor images surface normals: a contact dimple deepening changes pixels in place, and optical flow cannot express that. The video encoder pixview_D is much better (cos 0.812, ratio 0.885), but the view stream is where there was least headroom.

Ten held-out windows

Ground truth on top, then the four configs. Frames 0–4 are the clean context (identical for every row); frames 5–12 outlined in red are generated. Per-row PSNR is over the whole 13-frame clip, computed on the raw decoded tensors — the images here are h264, which costs a little fidelity equally in every row.

Each window also gets an error map: |prediction − ground truth| on one shared scale across all four configs. That panel is the one worth reading for the tactile streams — a gel image is nearly uniform, so the side-by-side strip cannot show where the configs actually differ.

005:0 — sample_000
camera view sample_000
camera view error sample_000
tactile left sample_000
tactile left error sample_000
tactile right sample_000
tactile right error sample_000
005:34 — sample_001
camera view sample_001
camera view error sample_001
tactile left sample_001
tactile left error sample_001
tactile right sample_001
tactile right error sample_001
005:68 — sample_002
camera view sample_002
camera view error sample_002
tactile left sample_002
tactile left error sample_002
tactile right sample_002
tactile right error sample_002
005:102 — sample_003
camera view sample_003
camera view error sample_003
tactile left sample_003
tactile left error sample_003
tactile right sample_003
tactile right error sample_003
005:136 — sample_004
camera view sample_004
camera view error sample_004
tactile left sample_004
tactile left error sample_004
tactile right sample_004
tactile right error sample_004
006:0 — sample_005
camera view sample_005
camera view error sample_005
tactile left sample_005
tactile left error sample_005
tactile right sample_005
tactile right error sample_005
006:150 — sample_006
camera view sample_006
camera view error sample_006
tactile left sample_006
tactile left error sample_006
tactile right sample_006
tactile right error sample_006
006:300 — sample_007
camera view sample_007
camera view error sample_007
tactile left sample_007
tactile left error sample_007
tactile right sample_007
tactile right error sample_007
006:450 — sample_008
camera view sample_008
camera view error sample_008
tactile left sample_008
tactile left error sample_008
tactile right sample_008
tactile right error sample_008
006:600 — sample_009
camera view sample_009
camera view error sample_009
tactile left sample_009
tactile left error sample_009
tactile right sample_009
tactile right error sample_009

Reproducing

# cache the frozen Stage-1 flows (32 episodes)
python data_preprocessing/build_stage1_pixel_flows.py

# gate the architecture before spending a GPU-day
python -m vm_diffusion.scripts.verify_pixflow

# train
sbatch vm_diffusion/scripts/train_mot.sh mot_df_vt_pixflow \
       vm_diffusion/config/train_mot_df_vt_pixflow.yaml
sbatch vm_diffusion/scripts/train_mot.sh mot_df_vt_vecfilm \
       vm_diffusion/config/train_mot_df_vt_vecfilm.yaml

# evaluate all four on the same ten windows and decode to pixels
sbatch vm_diffusion/run_eval_4way_test10.sh

Held-out episodes motherboard_0510_episode_005 and 006, split episode_holdout_0510_gel (30 train / 2 held out; the 4 pusht_0618 episodes are excluded — they have no force export and therefore no gel-frame action map). Rollouts are teacher-forced on the conditioning streams: the frozen encoders read ground-truth frames, so these numbers are an upper bound on what a closed-loop rollout would give.