Explicit Layer Modeling for Video Object Insertion and Video Layer DecompositionLayered diffusion framework that first jointly generates RGB scene content and RGBA foreground layers.
DBL-Diffusion runs an RGB branch and an RGBA branch side by side, exchanging information at every denoising step through joint cross-attention. The same backbone instantiates two tasks that were previously trained
without any direct layer supervision.
(VFrgb, α) = VF — the predicted RGBA foreground layer, split into its RGB and alpha channels. VB is the input background video and ⊙ is elementwise multiplication. The RGB branch's own composite prediction VC is only an auxiliary signal during denoising — the final result alpha-blends VF directly onto VB,
which is why DBL-Insert stays robust to minor inconsistencies in VC.
Green tiles denote the explicit RGBA foreground layer — pulled the way a compositor keys a matte against a green screen.
It carries both the object and the effects it induces (shadows, reflections).
Each row shows the model input on the left and
the two DBL-Diffusion outputs on the right.
The bottleneck for both tasks is the same: no existing dataset provides direct supervision for a foreground layer that includes object-induced effects.
TriLayer fills that gap.
"Happy dog running in sunny cambridge fields."
Stage-by-stage clip counts
Human verification rounds sit between stages 4, 5, 6, and 7, narrowing raw Pexels clips down to 3,908 high-quality triplets with a VLM-generated object-centric caption for each.
| Dataset | Composite | BG (V.E. removed) | 0/1 Mask | Alpha matte | Alpha (V.E. incl.) | Dynamic obj. | Real-world | # Frames |
|---|---|---|---|---|---|---|---|---|
| DVM | ✓ | ✗ | ✗ | ✓ | ✗ | ✓ | ✓ | 81< |
| VideoMatting108 | ✓ | ✗ | ✗ | ✓ | ✗ | ✓ | ✓ | 81< |
| VideoMatte240K | ✓ | ✗ | ✗ | ✓ | ✗ | ✓ | ✓ | 81< |
| Senorita-2M | ✓ | ✗ | ✓ | ✗ | ✗ | ✓ | ✓ | 33~64 |
| VPData | ✓ | ✗ | ✓ | ✗ | ✗ | ✓ | ✓ | 81< |
| ROSE++ | ✓ | ✓ | ✓ | ✗ | ✗ | ✗ | ✗ | 81< |
| TriLayer (Ours) | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | 81 |
0/1 Mask = SAM-based mask; V.E. = Visual Effect
Both tasks reuse the same VACE-based backbone, split into an RGB branch and an RGBA branch that talk to each other at every transformer block.
In-domain: predicts RGB video within the pretrained latent space (scene-level composite and clean background).
Out-of-domain: predicts the RGBA foreground layer using the WAN-alpha VAE (object's transparency, visual effects, and alpha mattes).
Each block enables bidirectional information exchange between the RGB and RGBA branches, allowing the RGB branch to incorporate foreground geometry while the RGBA branch captures scene-dependent context.
LoRA keeps the RGB branch close to its pretrained prior; DoRA's decoupled magnitude/direction update gives the RGBA branch room to learn transparency.
Each branch's noise level is controlled independently, which also enables conditional edits — e.g. freezing the RGBA branch at low noise to edit a foreground layer while keeping the background fixed.
On LayeredVid-Benchmark (65 videos) for insertion, and the Movie / Kubric benchmarks for decomposition. Pick an example below.
"Add Wooden Boat."
Row 1 is each method's composite output, row 2 its explicit foreground layer. VACE-Inp / VACE-C predict a composite directly and never expose a separate layer; VACE-F and DBL-Insert both learn a foreground layer, so their composite shown here is obtained by alpha-blending that layer onto the background.
VACE-Inp cannot synthesize object-induced effects beyond the masked region, while VACE-C and VACE-F show unstable effects and background appearance changes.
DBL-Insert jointly synthesizes explicit foreground layers and scene-consistent composites for more reliable integration.
| Model | ViCLIP-T↑ | BG Consist.↑ | Subj. Consist.↑ | Motion Smooth.↑ | Aesthetic↑ | Imaging Qual.↑ | CLIP-I↑ | DINO-I↑ |
|---|---|---|---|---|---|---|---|---|
| AnyV2V | 23.520 | 0.906 | 0.902 | 0.984 | 0.480 | 0.583 | 0.740 | 0.567 |
| ReVideo | 24.303 | 0.926 | 0.934 | 0.991 | 0.511 | 0.666 | 0.764 | 0.699 |
| VACE-Inp | 24.717 | 0.943 | 0.954 | 0.992 | 0.506 | 0.663 | 0.760 | 0.716 |
| VACE-C | 24.596 | 0.944 | 0.963 | 0.993 | 0.522 | 0.661 | 0.781 | 0.739 |
| VACE-F | 24.473 | 0.934 | 0.954 | 0.992 | 0.534 | 0.661 | 0.783 | 0.716 |
| DBL-Insert | 24.794 | 0.948 | 0.963 | 0.993 | 0.535 | 0.660 | 0.791 | 0.748 |
"Decompose Sheep."
Row 1 is each method's recovered background, row 2 its recovered foreground layer — all three methods produce both. Optimization-based (Gen-Omnimatte) and feed-forward (OmnimatteZero) methods often leave semi-transparent or incomplete foregrounds around fine boundaries and complex effects; DBL-Decompose keeps sharper boundaries and more complete effect separation, at a modest trade-off in pixel-wise reconstruction metrics.
| Method | Movie | Kubric | ||||
|---|---|---|---|---|---|---|
| PSNR↑ | LPIPS↓ | SSIM↑ | PSNR↑ | LPIPS↓ | SSIM↑ | |
| Reconstruction-based omnimatte methods | ||||||
| Omnimatte | 21.76 | 0.239 | 0.736 | 26.81 | 0.207 | 0.831 |
| D2NeRF | — | — | — | 34.99 | 0.113 | 0.887 |
| LNA | 23.10 | 0.129 | 0.847 | — | — | — |
| OmnimatteRF | 33.86 | 0.017 | 0.981 | 40.91 | 0.028 | 0.970 |
| Generative Omnimatte | 32.69 | 0.030 | 0.989 | 44.07 | 0.010 | 0.981 |
| OmnimatteZero | 35.11 | 0.014 | 0.992 | 44.97 | 0.010 | 0.988 |
| Generation-based diffusion methods | ||||||
| ObjectDrop | 28.05 | 0.124 | — | 34.22 | 0.083 | — |
| Lumiere Inp. | 26.62 | 0.148 | — | 31.46 | 0.157 | — |
| Propainter | 27.44 | 0.114 | — | 34.67 | 0.056 | — |
| DiffuEraser | 29.51 | 0.105 | — | 35.19 | 0.048 | — |
| DBL-Decompose | 33.42 | 0.017 | 0.971 | 38.78 | 0.021 | 0.977 |
While reconstruction-oriented methods such as Gen-Omnimatte and OmnimatteZero achieve higher PSNR/SSIM/LPIPS scores, this is expected since they are explicitly designed to reconstruct the input video by optimizing pixel-level fidelity or preserving latent representations through attention masking.
In contrast, DBL-Decompose leverages diffusion priors to generate perceptually plausible and visually coherent background layers rather than directly reproducing the ground-truth frames. As shown in above results, our method produces cleaner and more realistic backgrounds while maintaining high-quality foreground layers with accurate object boundaries and scene-dependent effects.
The RGB branch stays close to the pretrained latent space, so LoRA is enough; the RGBA branch has to learn an out-of-domain concept (transparency), where DoRA's decoupled magnitude/direction update adapts more stably.
| Method (RGB–RGBA) | ViCLIP-T↑ | BG Consist.↑ | Subj. Consist.↑ | Motion Smooth.↑ | Aesthetic↑ | Imaging Qual.↑ | CLIP-I↑ | DINO-I↑ |
|---|---|---|---|---|---|---|---|---|
| DoRA – DoRA | 24.520 | 0.927 | 0.935 | 0.993 | 0.525 | 0.641 | 0.774 | 0.672 |
| LoRA – LoRA | 24.681 | 0.938 | 0.954 | 0.994 | 0.539 | 0.657 | 0.775 | 0.742 |
| DoRA – LoRA | 24.604 | 0.935 | 0.944 | 0.993 | 0.539 | 0.664 | 0.788 | 0.701 |
| LoRA – DoRA (Ours) | 24.794 | 0.948 | 0.963 | 0.993 | 0.535 | 0.660 | 0.791 | 0.748 |
Object-removal models leave afterimage or ghosting artifacts. The proposed AGBI applies source-appearance-guided inpainting with a two-stage mask (bounding-box, then object-level) to remove these artifacts from the background videos used as training supervision.
Example 1
Example 2
Training on AGBI-refined background videos, rather than raw object-removal outputs, removes these artifacts from the supervision and prevents them from being propagated into the learned background reconstruction.
Example 1
Example 2
Unseen objects, scences, and VFX effects (e.g., fire, raindrop, and smoke).
DBL-Decompose achieves the superior perceptual fidelity for both foreground and background, as evidenced by consistently lower FID and more visually plausible results.
| Model | Kubric (BG) | Movie (BG) | LayeredVid-Benchmark (FG) |
|---|---|---|---|
| Gen-Omnimatte | 16.598 | 16.344 | 60.826 |
| OmnimatteZero | 96.077 | 36.428 | 46.882 |
| DBL-Decompose | 15.815 | 14.364 | 32.967 |
DBL-Insert is consistently preferred over all competing methods across every evaluation criterion, demonstrating superior insertion quality and better alignment with the intended editing objective.
| Model | Subject Consistency(↑) | Insertion Rationality(↑) | Prompt Alignment(↑) | Overall Quality(↑) |
|---|---|---|---|---|
| AnyV2V | 1.0% | 1.0% | 1.5% | 1.5% |
| ReVideo | 0.5% | 1.0% | 1.0% | 1.5% |
| VACE-Inp | 17.5% | 10.0% | 21.0% | 15.5% |
| VACE-C | 15.0% | 11.5% | 18.5% | 13.0% |
| VACE-F | 9.5% | 8.5% | 7.5% | 7.0% |
| DBL-Insert | 56.5% | 68.0% | 50.5% | 61.5% |
DBL-Decompose receives the highest preference for both background and foreground reconstruction by a large margin, indicating that our explicit layer modeling produces more visually plausible decomposed layers than existing methods.
| Model | Background Quality(↑) | Foreground Quality(↑) |
|---|---|---|
| Gen-Omnimatte | 8.0% | 19.0% |
| OmnimatteZero | 20.5% | 6.0% |
| DBL-Decompose | 71.5% | 75.0% |
The supplementary video showcases more qualitative examples and ablations, along with various applications including layer-based video editing, scene-aware video editing, and other applications enabled by our layered video representation.
@article{han2026explicitlayer,
title = {Explicit Layer Modeling for Video Object Insertion and Video Layer Decomposition},
author = {Han, kyujin and Shin, seungjoo and Cho, sunghyun},
journal = {arxiv},
year = {2026},
}