Synth4D is an omnipercipient model: it recovers appearance and motion from synchronized multi-view capture, then lets you move a camera through that space and time.
Film and digital video made the same bargain: lock a lens to a path, flatten the world onto a plane, call that the record. It is a good bargain if you only ever need that path. It fails the moment you want a camera that was not there.
Synth4D is ViewSynth’s omnipercipient model for spatial intelligence. It takes synchronized multi-view capture from our rigs and estimates a 4D Gaussian splat: appearance and motion you can reframe. After the crew packs, the take is still present from every angle the array saw. That is the claim. Omniperception, bounded by the lenses.
The problem
Novel-view video is an old computer-vision problem. Dense studio volumes solve it with hundreds of locked cameras. Generative video models solve a different problem: they synthesize plausible frames, then ask you to trust the geometry.
We want a third path. Capture what actually happened, on location, with a rig you can carry. Then recover enough of the 4D field that a later cut can stand anywhere in that field. Not a second shoot. A second set of eyes on the same beat.
What omniperception requires
Being everywhere at once is not a prompt. It is a capture problem first. Cameras must share a clock. Baselines must be known. The subject has to sit inside a volume the array actually covered. Synth4D does not invent a city block the lenses never saw. It reconstructs the volume they did.
That constraint is the point. If a surface was never in any lens, the model should not invent a confident lie. Honesty to the take is more useful for visual media than a world that looks complete and never happened.