Same 30-episode training split, same 100 epochs, same optimizer, same seed — four different conditioning architectures, evaluated on ten identical held-out windows.
Each window is 16 fps6 frames → 4 latent frames. The first 2 latent frames are given as clean context; the model generates the other 2 from the action alone. Decoded through the Wan2.2 VAE that is 13 pixel frames, of which frames 0–4 come from the context and frames 5–12 are generated. The strips at the bottom outline the generated frames in red.
All four models are the same trunk: 12 blocks, hidden 768, patch (1,2,2), temporal-causal joint attention, cross-first block order, per-frame independent noise levels per modality. They differ only in how the action reaches the tokens.
Both experts are conditioned by a frozen pixel-space Stage-1 flow concatenated into their latent, plus FiLM from the raw action vector. Action cross-attention, the analytic ActionImage, the actionflow encoder and the latent-flow probe loss are all gone.
pixview_D flow (2 ch, camera-crop px) concatenated into the 48-ch visual latent • FiLM from action[0:18], the camera-projected (u, v, z) of both sensorspixc5_D flow (2 ch, gel px), one per sensor, concatenated into each 48-ch tactile latent • FiLM from action[18+9s : 27+9s], that sensor's own raw 9-D SE(3) steppixc5_D only. Its 5-channel input map is the gel frame: ch2 is press depth along the true gel normal and ch3-4 are the contact force on that same axis (log1p(N) + a contact flag). Force therefore reaches the network on exactly one path and can only influence the tactile flows; every video-branch input is force-free.stage1_flow_k5.pt (cached frozen flows, already on the 16×16 latent grid)action_ctx: false
pixflow: {enable: true, video: true, tactile: true, share_tactile: true}
vector_film: {enable: true, hidden: 256}
action_image: {enable: false}
action_motion:{enable: false}
video_motion: {enable: false}
pred_flow_loss:{enable: false}
Byte-identical to pixflow except pixflow.enable: false. The dataloader still emits the flow stream, so the two runs see the same batches in the same order and only the model differs. It exists because pixflow removes four things and adds two at once — without this control, nothing in its result could be attributed to the flow.
action[0:18]. No flow.pixflow only through pixc5_D, and that encoder is not read here — so this run is the force-free twin as well as the flow-free one.action_ctx: false
pixflow: {enable: false} # the only difference
vector_film: {enable: true, hidden: 256}
action_image: {enable: false}
action_motion:{enable: false}
video_motion: {enable: false}
pred_flow_loss:{enable: false}
The previous best recipe. Conditioning is spread across three mechanisms: per-modality action cross-attention, a dense analytic ActionImage on the video expert, and the frozen VAE-latent actionflow encoder injected into the tactile expert. A frozen latent-flow probe adds an auxiliary loss.
action[0:18] • ActionImage: 6 Gaussian channels (2 sensors × pos/normal/up) rasterized on the 16×16 latent grid, σ = 0.8 cells, injected as a zero-init token biasaction[18:36] • frozen amx_gel5_s2 motion feature (128 ch) projected zero-init onto the future tactile tokensamx_gel5_s2, the VAE-latent gel-frame encoder, fed the same 5-channel [dx, dy, dz, log1p(N), contact] map on the 16×16 grid.tactile_action_map_gel_fps6.pt (5 ch), tactile_inverse_action_flow_gel_fps6.pt, tactile_flow6_lat16.ptaction_ctx: true # (default)
action_image: {enable: true, sigma_cells: 0.8}
action_motion: {enable: true, checkpoint: amx_gel5_s2/best.pt,
freeze: true, inject_to_dit: true}
pred_flow_loss:{enable: true, checkpoint: latent_flow_probe/best.pt,
weight: 0.05}
gelforce plus the frozen view motion encoder vmx_E — the video-side twin of actionflow. Its 128-ch feature is projected onto the video tokens and its predicted dense flow is appended to the ActionImage, making that raster learned and state-dependent rather than purely analytic.
gelforce has, plus vmx_E's motion feature as a second token bias and its 2-ch predicted flow appended to the ActionImage (6 → 8 channels); also consumes the sparse camera-frame motion raster action_map_viewgelforcegelforce — vmx_E is video-side and force-free.gelforce, plus view_motion_map: trueaction_ctx: true
action_image: {enable: true, sigma_cells: 0.8}
action_motion: {enable: true, checkpoint: amx_gel5_s2/best.pt, freeze: true}
video_motion: {enable: true, checkpoint: vmx_E/best.pt,
freeze: true, inject_to_dit: true, flow_to_action_image: true}
pred_flow_loss:{enable: true, weight: 0.05}
Conv[Wz|Wc](concat([z,c])) = ConvWz(z) + ConvWc(c). Adding a projected flow to the tokens is concatenating those channels into the patch embedding — exactly, not approximately (asserted numerically to 1.07e-6). Widening the expert’s own Conv3d(in_ch, …) would have been the wrong move: in_ch also sizes the output head, so it would silently change the diffusion variable from 48 to 50.| pathway | pixflow | vecfilm | gelforce | gelforce+vmx |
|---|---|---|---|---|
| action cross-attention (per modality) | — | — | ✓ | ✓ |
| analytic ActionImage → video tokens | — | — | ✓ | ✓ (+vmx flow) |
| frozen actionflow (VAE-latent) → tactile tokens | — | — | ✓ | ✓ |
| frozen view-motion encoder → video tokens | — | — | — | ✓ |
| frozen pixel-space flow concatenated | ✓ both | — | — | — |
| FiLM from the raw action vector | ✓ both | ✓ both | — | — |
| latent-flow probe auxiliary loss | — | — | ✓ 0.05 | ✓ 0.05 |
| config | val_loss_video | val_loss_tactile | best epoch |
|---|---|---|---|
| pixflow | 0.06823 ± 0.00024 | 0.03452 ± 0.00010 | 57 |
| vecfilm | 0.06798 ± 0.00023 | 0.03459 ± 0.00010 | 63 |
| gelforce | 0.06974 ± 0.00020 | 0.04635 ± 0.00013 | 95 |
| gelforce+vmx | 0.07048 ± 0.00011 | 0.04583 ± 0.00017 | 73 |

| window | pixflow | vecfilm | gelforce | gelforce+vmx |
|---|---|---|---|---|
| camera view — pixel PSNR, dB | ||||
| 005:0 | 24.56 | 24.28 | 21.01 | 24.94 |
| 005:34 | 22.27 | 22.45 | 21.49 | 23.42 |
| 005:68 | 24.30 | 24.67 | 22.30 | 23.50 |
| 005:102 | 23.55 | 23.59 | 22.37 | 23.67 |
| 005:136 | 23.32 | 23.04 | 21.71 | 23.64 |
| 006:0 | 24.09 | 24.04 | 24.29 | 24.12 |
| 006:150 | 24.45 | 25.18 | 23.28 | 24.28 |
| 006:300 | 24.95 | 24.85 | 24.81 | 24.52 |
| 006:450 | 23.79 | 24.11 | 22.56 | 23.68 |
| 006:600 | 23.71 | 23.50 | 22.55 | 23.48 |
| mean | 23.90 | 23.97 | 22.64 | 23.93 |
| tactile left — pixel PSNR, dB | ||||
| 005:0 | 39.79 | 41.42 | 40.79 | 37.91 |
| 005:34 | 39.38 | 40.29 | 38.91 | 39.40 |
| 005:68 | 42.33 | 43.18 | 42.15 | 40.83 |
| 005:102 | 39.60 | 39.17 | 37.81 | 37.54 |
| 005:136 | 50.51 | 51.22 | 47.98 | 48.59 |
| 006:0 | 51.74 | 52.07 | 48.94 | 49.30 |
| 006:150 | 42.54 | 43.22 | 40.97 | 39.96 |
| 006:300 | 51.36 | 51.43 | 48.23 | 48.54 |
| 006:450 | 42.80 | 42.85 | 40.42 | 40.57 |
| 006:600 | 50.67 | 50.93 | 47.99 | 48.23 |
| mean | 45.07 | 45.58 | 43.42 | 43.09 |
| tactile right — pixel PSNR, dB | ||||
| 005:0 | 51.22 | 48.23 | 43.62 | 46.11 |
| 005:34 | 52.55 | 52.58 | 49.98 | 50.47 |
| 005:68 | 48.25 | 48.45 | 46.94 | 46.95 |
| 005:102 | 52.75 | 52.63 | 50.04 | 50.35 |
| 005:136 | 38.59 | 38.70 | 37.14 | 36.68 |
| 006:0 | 51.29 | 51.30 | 49.16 | 49.70 |
| 006:150 | 41.67 | 41.43 | 40.23 | 40.79 |
| 006:300 | 39.34 | 39.64 | 36.57 | 38.44 |
| 006:450 | 45.26 | 45.30 | 44.71 | 44.63 |
| 006:600 | 40.28 | 40.08 | 38.05 | 39.45 |
| mean | 46.12 | 45.83 | 43.64 | 44.36 |
Because every config saw the same ten windows, the difference can be taken window by window instead of comparing two noisy means.
| comparison | stream | mean Δ | median Δ | sd | wins |
|---|---|---|---|---|---|
| pixflow − vecfilm | camera view | -0.07 dB | +0.01 dB | 0.33 | 5/10 |
| pixflow − vecfilm | tactile left | -0.51 dB | -0.50 dB | 0.58 | 1/10 |
| pixflow − vecfilm | tactile right | +0.28 dB | -0.02 dB | 0.96 | 4/10 |
| new pair − baseline pair | camera view | +0.65 dB | +0.57 dB | 0.58 | 8/10 |
| new pair − baseline pair | tactile left | +2.07 dB | +2.37 dB | 0.79 | 10/10 |
| new pair − baseline pair | tactile right | +1.98 dB | +1.80 dB | 1.16 | 10/10 |
pixflow would have looked like a clear win over the baselines.gelforce_vmx matches the new pair). Two clean context frames plus the camera-projected action apparently determine the next two latent frames about as well as any of these mechanisms can.gelforce’s minimum is at epoch 95. The lighter conditioning converges faster and then mildly overfits.pixc5_D is weak — held-out direction cosine 0.645 and magnitude ratio 0.44, i.e. it recovers less than half the true flow length. A separate study also found that a perfect, dense, correctly-splatted flow explains only ~12.5% of the frame-to-frame change on a gel image, because a vision-based tactile sensor images surface normals: a contact dimple deepening changes pixels in place, and optical flow cannot express that. The video encoder pixview_D is much better (cos 0.812, ratio 0.885), but the view stream is where there was least headroom.Ground truth on top, then the four configs. Frames 0–4 are the clean context (identical for every row); frames 5–12 outlined in red are generated. Per-row PSNR is over the whole 13-frame clip, computed on the raw decoded tensors — the images here are h264, which costs a little fidelity equally in every row.
Each window also gets an error map: |prediction − ground truth| on one shared scale across all four configs. That panel is the one worth reading for the tactile streams — a gel image is nearly uniform, so the side-by-side strip cannot show where the configs actually differ.




























































# cache the frozen Stage-1 flows (32 episodes)
python data_preprocessing/build_stage1_pixel_flows.py
# gate the architecture before spending a GPU-day
python -m vm_diffusion.scripts.verify_pixflow
# train
sbatch vm_diffusion/scripts/train_mot.sh mot_df_vt_pixflow \
vm_diffusion/config/train_mot_df_vt_pixflow.yaml
sbatch vm_diffusion/scripts/train_mot.sh mot_df_vt_vecfilm \
vm_diffusion/config/train_mot_df_vt_vecfilm.yaml
# evaluate all four on the same ten windows and decode to pixels
sbatch vm_diffusion/run_eval_4way_test10.sh