Source motion, extracted for the edit finetune

Ten held-out test clips. Each source video is turned into a dense motion field on the video-latent grid, so it can be channel-concatenated to the noisy latent as extra conditioning — and, separately, reduced back to the single line a person would have drawn to produce it.

video → AllTracker reference flow (81 frames @ 256×256) → forward splat into frame-t coordinates → causal pool to 21×32×32 → [3, 21, 32, 32] fp16 (dx, dy, validity)

Two things this measurement says, before any training.

Per clip

Only of 10 clips are expressible as one drawn line. wander is path length over net displacement. A hand-drawn drag is a single monotone stroke, so a region that oscillates — a strumming arm, water, a bird flapping — has a long path and almost no net travel, and no single line can represent it. The first version of this tool happily emitted a zig-zag for those; it now says so instead. Clips with wander < 2 and real travel are the ones where a drag is actually the right control surface.

AllTracker (window 16, 4 iters) on the exact 81 frames the cached VAE latents were built from, via JavisDiT's own load_video_at_target_fps. The splat is a pure-PyTorch reimplementation of softsplat's avg mode (cupy is unavailable here), validated against closed-form answers: exact on integer and sub-pixel translation, matching the inverse area Jacobian 1/s² to 0.3% on an affine zoom, and exactly zero round-trip error on locally-constant flow. Latent 0 is forced to exactly (0, 0, 1) — AllTracker leaves a ~0.2–0.4 px self-regression residual at frame 0 that would otherwise manufacture spurious sub-pixel holes.