LEGO: A Lifting-Free Approach for
Exocentric-to-Egocentric Video Generation

Suhwan Cho* Yonwoo Choi* Soongjin Kim* Jicheol Park Taegyu Lim

GenGenAI

*Equal contribution

TL;DRGiven a single third-person video, LEGO generates what a person in it sees, without estimating depth or building a point cloud.

Egocentric videos generated by LEGO from stop-motion LEGO footage, far from the training data. The right side is the view from the head of one minifigure along a user-specified camera trajectory.

Abstract

Generating an egocentric video from a single exocentric recording is a challenging case of novel view synthesis, as the two cameras share little overlap and much of the target view is unobserved. Current state-of-the-art methods reconstruct the scene explicitly by estimating depth, lifting the video into a point cloud, and re-rendering it from the egocentric camera to condition a video diffusion model. This deterministic mapping assigns each pixel to a single reprojected location, which preserves texture but translates depth errors into misplaced content. We ask what a video diffusion model should receive as its condition and propose a lifting-free answer: a learned view synthesizer, an LVSM-style transformer fine-tuned to render the egocentric view directly without depth, point clouds, or reprojection, resolving cross-view correspondence internally. In contrast, its probabilistic mapping averages each region over candidate source locations according to a learned correspondence distribution, preserving structure while fine texture is averaged away. We argue that this trade-off suits a diffusion generator, whose denoising training excels at restoring detail, so an effective condition should prioritize structural alignment over sharpness. This distribution's concentration also yields a per-region confidence, used both to mask low-confidence regions and to guide the generator toward high-confidence areas during early layout-forming denoising steps. Our approach consistently outperforms the state-of-the-art explicit pipeline and generalizes to other datasets without retraining. The synthesizer thus supplies view structure, and the diffusion model its detail.

Motivation

Aligned before sharp

Reconstruction pipelines estimate depth, lift the exocentric video into a point cloud, and re-render it from the egocentric camera. Each pixel lands at one reprojected location, so the render stays sharp, but a depth error moves content to the wrong place, and much of the egocentric frame is left uncovered.

LEGO renders the egocentric view with a learned view synthesizer instead. It averages over the candidate source locations, so its errors are blur rather than misplaced content. Blur is the corruption that denoising training teaches a diffusion model to undo, while misplaced content has to be overridden. The render below is soft, but it puts the scene where it belongs.

Condition generation in EgoX and LEGO: depth estimation, lifting and re-rendering versus direct rendering with a frozen synthesizer
Condition generation. EgoX estimates depth, lifts the exocentric video into a point cloud, and re-renders it along the egocentric trajectory. LEGO renders the egocentric view directly with a frozen synthesizer.

Method

How LEGO works

LEGO keeps the video diffusion backbone and the canvas of EgoX, where the exocentric and egocentric views are placed side by side and the egocentric half of the conditioning latent carries the condition. It changes two things: what fills that condition, and how the early denoising steps are guided.

LEGO pipeline: a frozen view synthesizer renders the egocentric view and a confidence map; the gated render conditions the diffusion model, and ARC guides early denoising
Overview. The frozen synthesizer renders the egocentric view once per latent frame and exposes a confidence map. The render, gated to gray where confidence is low, fills the egocentric half of the conditioning latent, and the confidence also drives ARC during the early denoising steps.
  1. render

    Learned view synthesizer

    An LVSM-style transformer maps the exocentric frame and the ray maps of both cameras to the egocentric fisheye frame. It is fine-tuned for 10k steps, about 3.5 hours on four H200 GPUs, and then frozen.

  2. gate

    Confidence gating

    The synthesizer's attention forms a distribution over exocentric source locations. Its top-16 mass is a per-region confidence, and regions below the threshold are set to neutral gray, so the generator completes them from its own prior.

  3. guide

    Adaptive Rendering Consistency

    During the early steps, when the layout is decided, ARC moves the generator's clean estimate toward the render in proportion to the squared confidence. It needs no training and replaces a depth-derived attention bias that prior work trains into the generator.

Results

Comparison with EgoX

Both methods receive the same camera poses, text prompt, and random seed. EgoX is the authors' released checkpoint. EgoHumans and Nymeria are never used for training: LEGO runs with the weights trained on Ego-Exo4D.

Playback

Unseen scene, cooking

Exocentric input
EgoX
LEGO
Ground truth

Unseen scene, basketball

Exocentric input
EgoX
LEGO
Ground truth

Quantitative results

Comparison on Ego-Exo4D

Image criteriaVideo criteria
MethodPSNR ↑SSIM ↑LPIPS ↓CLIP-I ↑FVD ↓TF ↑MS ↑DD ↑
Seen scenes
Exo2Ego-V14.530.3840.5690.774622.470.9600.9660.985
TrajectoryCrafter13.050.3750.6060.780546.090.9600.9800.947
Wan Fun Control12.250.4630.6170.810595.070.9680.9800.901
Wan VACE12.950.4130.6260.829508.690.9890.9940.673
EgoX16.050.5560.4980.896184.470.9770.9900.974
EgoX†15.570.5560.4660.903197.950.9850.9930.990
LEGO20.280.6890.3100.925155.560.9890.9940.987
Unseen scenes
Exo2Ego-V12.700.4390.5970.6791283.500.9710.9760.978
TrajectoryCrafter12.240.2970.6190.778821.710.9660.9840.944
Wan Fun Control13.590.4390.6040.799968.780.9710.9850.944
Wan VACE12.170.3450.6380.8201045.450.9950.9960.427
EgoX14.380.4570.5520.877440.640.9810.9920.989
EgoX†13.530.4370.5420.886455.060.9860.9940.990
LEGO15.910.5050.5040.897404.770.9900.9951.000

Rows without a mark are quoted from the EgoX paper. †The released EgoX checkpoint rerun under our evaluation code. TF, MS, and DD are the VBench temporal flickering, motion smoothness, and dynamic degree scores.

Generalization to other datasets

Image criteriaVideo criteria
MethodPSNR ↑SSIM ↑LPIPS ↓CLIP-I ↑FVD ↓TF ↑MS ↑DD ↑
EgoHumans
EgoX†13.740.4600.5730.806421.830.9830.9911.000
LEGO15.520.5300.5450.819332.410.9850.9921.000
Nymeria
EgoX†11.640.4510.6370.804296.980.9840.9931.000
LEGO15.230.5340.5610.821210.270.9880.9941.000

Both methods use weights trained on Ego-Exo4D only, with no fine-tuning. Each dataset has 200 test clips.

Ablation

What each component contributes

All arms share one training recipe and differ only in the conditioning input and in how denoising is guided.

Seen scenesUnseen scenes
ArmPSNR ↑LPIPS ↓CLIP-I ↑FVD ↓PSNR ↑LPIPS ↓CLIP-I ↑FVD ↓
Conditioning input (exocentric tokens hidden)
No condition12.300.5910.862310.0011.380.6400.838695.44
Point-cloud render15.590.4780.895195.8013.420.5950.853549.77
Untuned synthesizer render15.710.5030.888212.9413.560.5960.848560.51
Synthesizer render18.240.3640.904183.4014.350.5670.860581.90
Completing the configuration (cumulative)
+ confidence gating18.270.3750.904166.8614.400.5660.868502.47
+ exocentric tokens18.910.3400.921159.0414.980.5150.897415.08
Denoising guidance
GGA (trained into the model)18.620.3560.915194.9815.200.5250.896429.42
ARC (point-cloud render)13.960.5860.801735.5013.720.6180.795812.88
ARC (synthesizer render) = LEGO20.280.3100.925155.5615.910.5040.897404.77
Qualitative ablation on one unseen clip following the rows of the ablation table
Qualitative ablation on one unseen clip, following the rows of the table. The second row shows enlarged crops of the corresponding regions in the first row.

Confidence gating

Without a gate, the synthesizer render fills every location, including those with no exocentric evidence. The gate sets locations whose confidence falls below the threshold to neutral gray.

The synthesizer render gated at thresholds from 0.1 to 0.5, next to the ungated render and the EgoX point-cloud render
The synthesizer render gated at τ = 0.1 to 0.5, shown between the ungated render and the point-cloud render of EgoX. LEGO uses τ = 0.3.

Limitations

LEGO inherits the limitations of its view synthesizer. The synthesizer renders each frame from a single exocentric frame, so it cannot recover content that the frame does not show, and it blurs where the correspondence is uncertain. On unfamiliar scenes its colors also fade. The pipeline needs no depth, but it still needs the camera poses of both views. Because the synthesizer is frozen and enters only through its render and confidence, a stronger feed-forward view synthesizer can replace it without changing the diffusion model or ARC.

Failure case where the exocentric frame does not cover the egocentric view and the condition is almost entirely gray
Where the exocentric frame does not cover the egocentric view, the condition is empty and the generator fills the region from its own prior.

BibTeX

@misc{cho2026lego,
  title  = {LEGO: A Lifting-Free Approach for Exocentric-to-Egocentric Video Generation},
  author = {Cho, Suhwan and Choi, Yonwoo and Kim, Soongjin and Park, Jicheol and Lim, Taegyu},
  year   = {2026}
}