Multi-agent world models

OverSeer: Multi-Agent World Models Beyond What Agents See

Anonymous authors

Under review · project page for the anonymous submission

A multi-agent world model that renders each player's view together with a few fixed auxiliary cameras. Because the cameras never move, they anchor the shared world: places stay where they were when a player returns, and what one player does is seen by the others. A δ-VAE codes each auxiliary stream as one reference frame plus a compact residual, so an extra camera costs an 8 × smaller share of tokens than a player stream.

Illustration of the setting, rendered by the game engine — not model output. Two agents, six fixed cameras, one shared world. OverSeer predicts the agents' own views (bottom) jointly with the auxiliary views (right).

Idea

Egocentric streams are partial and entangled, so a multi-agent world model built on them alone drifts: revisited places change and the players disagree about each other. OverSeer adds fixed auxiliary cameras and predicts all streams together. A stationary view changes only when the world changes, so it is a persistent shared reference, and it compresses well: the δ-VAE keeps one reference latent per camera and a small residual for what moved (108 tokens per latent frame instead of 864).

How it works

One frame of every stream and the recorded actions go in; the model predicts all eight streams together, for thirty seconds. Everything from here on is what the model produces, at the resolution it produces it.

Method figure. Stage 1, training the δ-VAE: the encoder takes an auxiliary view's reference latent (t = 0) and later latents and keeps only a residual code; the decoder repeats the reference latent, adds a learnable positional embedding and reconstructs the later latents from it and the residual code. Stage 2, training OverSeer: the action text embedding, the agents' reference and noised latents, the auxiliary views' reference latent, their residual codes and the reference Plücker rays are concatenated into one sequence for a diffusion transformer (cross-attention to the action, self-attention, camera-ray modulation), which reconstructs the agent view latents and the auxiliary view latents together.

OverSeer rollouts

Every frame is generated from the first frame of each stream and the recorded actions. Agent 1 Agent 2 above, the numbered auxiliary cameras on the band below. The panel on each agent view is the action the model receives at that frame: the movement keys, the look direction (the dot in the ring) and the buttons, lit while active. In Doom the fire chip follows the shot the clip itself shows, read off the muzzle flash. Each domain shows a few episodes; More samples opens the rest, picked by the benchmark's own quality ranking.

Minecraft

Agent view

Auxiliary view

Agent view

Auxiliary view

Agent view

Auxiliary view

Agent view

Auxiliary view

Doom

Agent view

Auxiliary view

Agent view

Auxiliary view

Agent view

Auxiliary view

Agent view

Auxiliary view

Real multi-camera capture

Two head-mounted cameras and six fixed cameras in a studio. Here the panel is read off the head-mounted camera's poses: the walking direction in the wearer's own frame as the keys, the head turn as the look ring.

Agent view

Auxiliary view

Agent view

Auxiliary view

Agent view

Auxiliary view

Agent view

Auxiliary view

Comparison with baseline models

Scripted tasks that only a consistent world passes: walk apart and reunite, or leave an event and come back to its outcome. Same first frame and actions (the panel) for every system; the columns play in step.

Speed ×1.00

Ground truth

OverSeer

Solaris

γ-World

The two agents start side by side, walk apart and return. At the end each has to see the other again where it left it.

Consistency over time

The same tasks in both domains, and the same question: when a place, an event or the partner comes back into view, is it the one that was there before? Every block has the same three columns, and the middle one is the world without the auxiliary views.

Reunite at start · Doom

The marines walk out and come back to the pose they started in, so the last frames revisit the first.

Speed ×1.00

Ground truth

Ego only

OverSeer

E4M3. Both marines walk out and back to their starting pose, so the last frame revisits the first: the same room for OverSeer, a rebuilt one for the ego-only run.

Reunite far · Doom

The marines start out of sight of each other and walk to a meeting point, where the partner has to be.

Speed ×1.00

Ground truth

Ego only

OverSeer

E2M7. Only the run that keeps the auxiliary views reaches the partner; the ego-only world beyond the first frame is rebuilt and the meeting never happens.

Reunite at start · Minecraft

The two agents start side by side, walk apart and return; at the end each has to see the other where it was left.

Speed ×1.00

Ground truth

Ego only

OverSeer

The partner is back in view for OverSeer at the end of the walk; the ego-only run has lost it.

Reunite far · Minecraft

The agents start out of each other's sight and walk to the same place, which only a shared world brings them to.

Speed ×1.00

Ground truth

Ego only

OverSeer

The rendezvous at the end of the clip: the partner is there for OverSeer, absent without the auxiliary views.

Dynamic event · Minecraft

Something happens between the agents at the start, they walk away, and the outcome has to be there when they look back.

Speed ×1.00

Ground truth

Ego only

OverSeer

A creeper standing between the agents explodes early on and leaves a crater; both agents walk away and look back at the spot.

What the δ-VAE codes carry

The δ-VAE stores each auxiliary camera as its first frame (the reference latent) plus a small residual code per later frame: 108 tokens per latent frame instead of 864. The scene lives in the reference; the code carries what changes.

Reconstruction with the clip's own code

Decoding the reference with the clip's own code gives the clip back. The scene stays as it was, and everything that moves (monsters, players, animals, lava, fire) comes from the code.

Doom

Ground truth

Reconstruction

Ground truth

Reconstruction

Minecraft

Ground truth

Reconstruction

Ground truth

Reconstruction

Ground truth

Reconstruction

Ground truth

Reconstruction

Transfer: another clip's code

Decode a host camera's reference frame with a donor clip's residual code: the donor's mover appears in the host scene, the scene itself is untouched. The control decodes the same reference with the host's own code.

Doom

Donor · code source

Control · own code

Transfer · donor's code

Doom

Donor · code source

Control · own code

Transfer · donor's code

Minecraft

Donor · code source

Control · own code

Transfer · donor's code

Minecraft

Donor · code source

Control · own code

Transfer · donor's code