Introducing Synth4D: omnipercipient spatial intelligence Read the note →

Research note · 2026

Synth4D: An Omnipercipient Model for Spatial Intelligence

Synth4D estimates a 4D Gaussian splat from synchronized multi-view capture. Once that field exists, a new camera path is a render: spatial intelligence over a take the cameras actually saw.

Source take Synth4D reshoot
Synth4D is an omnipercipient model: it recovers appearance and motion from synchronized multi-view capture, then lets you move a camera through that space and time.

Film and digital video made the same bargain: lock a lens to a path, flatten the world onto a plane, call that the record. It is a good bargain if you only ever need that path. It fails the moment you want a camera that was not there.

Synth4D is ViewSynth’s omnipercipient model for spatial intelligence. It takes synchronized multi-view capture from our rigs and estimates a 4D Gaussian splat: appearance and motion you can reframe. After the crew packs, the take is still present from every angle the array saw. That is the claim. Omniperception, bounded by the lenses.

The problem

Novel-view video is an old computer-vision problem. Dense studio volumes solve it with hundreds of locked cameras. Generative video models solve a different problem: they synthesize plausible frames, then ask you to trust the geometry.

We want a third path. Capture what actually happened, on location, with a rig you can carry. Then recover enough of the 4D field that a later cut can stand anywhere in that field. Not a second shoot. A second set of eyes on the same beat.

What omniperception requires

Being everywhere at once is not a prompt. It is a capture problem first. Cameras must share a clock. Baselines must be known. The subject has to sit inside a volume the array actually covered. Synth4D does not invent a city block the lenses never saw. It reconstructs the volume they did.

That constraint is the point. If a surface was never in any lens, the model should not invent a confident lie. Honesty to the take is more useful for visual media than a world that looks complete and never happened.

ViewSynth multi-camera capture rig
ViewSynth capture rig: hardware-synchronized multi-view on location.

The stack

Rigs

ViewSynth designs handheld multi-camera arrays. Cameras are hardware-synchronized so every frame shares a timestamp. That is the input Synth4D expects: a short, dense burst of views around a subject, not a monocular clip, and not a text prompt.

Model

Synth4D fits a 4D Gaussian splat to that take. Gaussians are a compact way to store color and geometry that still renders in real time. The fourth dimension is time: the splat is allowed to change as the action moves, so you can freeze a beat and slide a virtual camera through it.

Reshoot

Once the splat exists, a new camera path is a render, not a recapture. That is the reshoot. Same take. Every viewpoint at once.

What Synth4D is not

It is not a world model in the Atlas or Odyssey sense. We do not generate persistent environments from text, or simulate physics for robots. We do not fill in what the cameras never saw and call it ground truth.

Atlas and Odyssey ask: can a model imagine a world? Synth4D asks: can a model be present to a world that already happened? Those are different research programs. Ours stays coupled to capture.

Why hardware and model stay coupled

Most 4DGS pipelines assume a capture partner and a processing partner. We keep both in-house so baseline, sync, and the model’s priors move together. A change in rig spacing is a change in what the model can resolve. That is slower than an API on other people’s footage. It is how you get a look that is ours.

Where it shows up

Marketing: campaign assets you can reframe after wrap. Entertainment: cameras that were not on the day. Education: a lecture or demonstration you can walk around after it ends.

Status

Synth4D is in active development on real takes from our rigs. This note is a position, not a benchmark card. If you are a researcher or a production team who wants to push on the capture-model loop, write to us.

Work with ViewSynth →

Cite as: ViewSynth, “Synth4D: An Omnipercipient Model for Spatial Intelligence,” 2026. viewsynth.ai/research.html