arXiv 2026

DBL-Diffusion

Explicit Layer Modeling for Video Object Insertion and Video Layer DecompositionLayered diffusion framework that first jointly generates RGB scene content and RGBA foreground layers.

Kyujin Han, Seungjoo Shin, Sunghyun Cho†
POSTECH
Paper PDF Code 🤗TriLayer Dataset 🤗Weights Supple. Video
🤗We will publicly release the code, dataset, and pretrained models upon paper acceptance.🤗
01 — Two tasks, one backbone

One dual-branch model,
two layered video tasks

DBL-Diffusion runs an RGB branch and an RGBA branch side by side, exchanging information at every denoising step through joint cross-attention. The same backbone instantiates two tasks that were previously trained
without any direct layer supervision.

DBL-Insert: Add Spyder-man and Iron Man Helmat.
BG video (input)
Edited first frame
FG layer (RGBA)
Composite
Vout = α ⊙ VFrgb + (1 − α) ⊙ VB

(VFrgb, α) = VF — the predicted RGBA foreground layer, split into its RGB and alpha channels. VB is the input background video and ⊙ is elementwise multiplication. The RGB branch's own composite prediction VC is only an auxiliary signal during denoising — the final result alpha-blends VF directly onto VB,
which is why DBL-Insert stays robust to minor inconsistencies in VC.

DBL-Decompose: Decompose Bird.
Composite (input)
BG video
FG layer (RGBA)

Green tiles denote the explicit RGBA foreground layer — pulled the way a compositor keys a matte against a green screen.
It carries both the object and the effects it induces (shadows, reflections).

02 — Qualitative results

General performance,
lavered video results

Each row shows the model input on the left and
the two DBL-Diffusion outputs on the right.

Video Layered Object Insertion — BG → FG + Composite

BG (input)
FG
Comp
"Add Man."
BG (input)
FG
Comp
"Add Dusty Car."
BG (input)
FG
Comp
"Add animated Golem."
BG (input)
FG
Comp
"Add White Swan."

Video Layer Decomposition — Composite → BG + FG

Comp (input)
BG
FG
"Decompose Cat."
Comp (input)
BG
FG
"Decompose Dog."
Comp (input)
BG
FG
"Decompose Sheep."
Comp (input)
BG
FG
"Decompose Surfing Board."
03 — TriLayer dataset

Aligned 3,908 video triplets
composite–background–foreground

The bottleneck for both tasks is the same: no existing dataset provides direct supervision for a foreground layer that includes object-induced effects.
TriLayer fills that gap.

"Happy dog running in sunny cambridge fields."

Composite
Background
Foreground (RGBA)

Construction pipeline

Stage-by-stage clip counts

1
Video database
18,000
2
Preprocess & shot filter
18,000
3
VLM filtering
9,178
4
Object mask gen.
9,178
5
BG layer by removal
5,492
6
FG layer + alpha
4,956
7
BG refinement (AGBI)
3,908

Human verification rounds sit between stages 4, 5, 6, and 7, narrowing raw Pexels clips down to 3,908 high-quality triplets with a VLM-generated object-centric caption for each.

TriLayer vs. existing datasets

DatasetCompositeBG (V.E. removed)0/1 MaskAlpha matteAlpha (V.E. incl.)Dynamic obj.Real-world# Frames
DVM✓✗✗✓✗✓✓81<
VideoMatting108✓✗✗✓✗✓✓81<
VideoMatte240K✓✗✗✓✗✓✓81<
Senorita-2M✓✗✓✗✗✓✓33~64
VPData✓✗✓✗✗✓✓81<
ROSE++✓✓✓✗✗✗✗81<
TriLayer (Ours)✓✓✓✓✓✓✓81

0/1 Mask = SAM-based mask; V.E. = Visual Effect

04 — Architecture

A dual-branch diffusion backbone

Both tasks reuse the same VACE-based backbone, split into an RGB branch and an RGBA branch that talk to each other at every transformer block.

RGB Branch

VACE + LoRA

In-domain: predicts RGB video within the pretrained latent space (scene-level composite and clean background).

⇄
RGBA Branch

VACE + DoRA

Out-of-domain: predicts the RGBA foreground layer using the WAN-alpha VAE (object's transparency, visual effects, and alpha mattes).

Joint cross-attention

Each block enables bidirectional information exchange between the RGB and RGBA branches, allowing the RGB branch to incorporate foreground geometry while the RGBA branch captures scene-dependent context.

LoRA–DoRA hybrid

LoRA keeps the RGB branch close to its pretrained prior; DoRA's decoupled magnitude/direction update gives the RGBA branch room to learn transparency.

Disentangled timestep sampling

Each branch's noise level is controlled independently, which also enables conditional edits — e.g. freezing the RGBA branch at low noise to edit a foreground layer while keeping the background fixed.

05 — Baseline comparison

How DBL-Diffusion stacks up against prior methods

On LayeredVid-Benchmark (65 videos) for insertion, and the Movie / Kubric benchmarks for decomposition. Pick an example below.

Video Object Insertion

"Add Wooden Boat."

Source video (input)
Edited first frame (input)
Source video & edited first frame
Composite
No FG output
VACE-Inp
Composite
No FG output
VACE-C
Composite
FG layer
VACE-F
Composite
FG layer
DBL-Insert (Ours)

Row 1 is each method's composite output, row 2 its explicit foreground layer. VACE-Inp / VACE-C predict a composite directly and never expose a separate layer; VACE-F and DBL-Insert both learn a foreground layer, so their composite shown here is obtained by alpha-blending that layer onto the background.
VACE-Inp cannot synthesize object-induced effects beyond the masked region, while VACE-C and VACE-F show unstable effects and background appearance changes.
DBL-Insert jointly synthesizes explicit foreground layers and scene-consistent composites for more reliable integration.

ModelViCLIP-T↑BG Consist.↑Subj. Consist.↑Motion Smooth.↑Aesthetic↑Imaging Qual.↑CLIP-I↑DINO-I↑
AnyV2V23.5200.9060.9020.9840.4800.5830.7400.567
ReVideo24.3030.9260.9340.9910.5110.6660.7640.699
VACE-Inp24.7170.9430.9540.9920.5060.6630.7600.716
VACE-C24.5960.9440.9630.9930.5220.6610.7810.739
VACE-F24.4730.9340.9540.9920.5340.6610.7830.716
DBL-Insert24.7940.9480.9630.9930.5350.6600.7910.748

Video Layer Decomposition

"Decompose Sheep."

Composite (input)
Source video
Background
FG layer
Gen-Omnimatte
Background
FG layer
OmnimatteZero
Background
FG layer
DBL-Decompose (Ours)

Row 1 is each method's recovered background, row 2 its recovered foreground layer — all three methods produce both. Optimization-based (Gen-Omnimatte) and feed-forward (OmnimatteZero) methods often leave semi-transparent or incomplete foregrounds around fine boundaries and complex effects; DBL-Decompose keeps sharper boundaries and more complete effect separation, at a modest trade-off in pixel-wise reconstruction metrics.

MethodMovieKubric
PSNR↑LPIPS↓SSIM↑PSNR↑LPIPS↓SSIM↑
Reconstruction-based omnimatte methods
Omnimatte21.760.2390.73626.810.2070.831
D2NeRF———34.990.1130.887
LNA23.100.1290.847———
OmnimatteRF33.860.0170.98140.910.0280.970
Generative Omnimatte32.690.0300.98944.070.0100.981
OmnimatteZero35.110.0140.99244.970.0100.988
Generation-based diffusion methods
ObjectDrop28.050.124—34.220.083—
Lumiere Inp.26.620.148—31.460.157—
Propainter27.440.114—34.670.056—
DiffuEraser29.510.105—35.190.048—
DBL-Decompose33.420.0170.97138.780.0210.977

While reconstruction-oriented methods such as Gen-Omnimatte and OmnimatteZero achieve higher PSNR/SSIM/LPIPS scores, this is expected since they are explicitly designed to reconstruct the input video by optimizing pixel-level fidelity or preserving latent representations through attention masking.
In contrast, DBL-Decompose leverages diffusion priors to generate perceptually plausible and visually coherent background layers rather than directly reproducing the ground-truth frames. As shown in above results, our method produces cleaner and more realistic backgrounds while maintaining high-quality foreground layers with accurate object boundaries and scene-dependent effects.

06 — Ablations

What each design choice buys you

LoRA vs. DoRA training strategy

The RGB branch stays close to the pretrained latent space, so LoRA is enough; the RGBA branch has to learn an out-of-domain concept (transparency), where DoRA's decoupled magnitude/direction update adapts more stably.

Method (RGB–RGBA)ViCLIP-T↑BG Consist.↑Subj. Consist.↑Motion Smooth.↑Aesthetic↑Imaging Qual.↑CLIP-I↑DINO-I↑
DoRA – DoRA24.5200.9270.9350.9930.5250.6410.7740.672
LoRA – LoRA24.6810.9380.9540.9940.5390.6570.7750.742
DoRA – LoRA24.6040.9350.9440.9930.5390.6640.7880.701
LoRA – DoRA (Ours)24.7940.9480.9630.9930.5350.6600.7910.748

Background refinement — dataset quality

Object-removal models leave afterimage or ghosting artifacts. The proposed AGBI applies source-appearance-guided inpainting with a two-stage mask (bounding-box, then object-level) to remove these artifacts from the background videos used as training supervision.

Example 1

Composite (input)
Object-removal model
After AGBI (refined)

Example 2

Composite (input)
Object-removal model
After AGBI (refined)

Background refinement — model reconstruction quality

Training on AGBI-refined background videos, rather than raw object-removal outputs, removes these artifacts from the supervision and prevents them from being propagated into the learned background reconstruction.

Example 1

Composite (input)
w/o refined dataset
w/ refined dataset

Example 2

Composite (input)
w/o refined dataset
w/ refined dataset
08 - Appendix: More Quantitative Results

Appendix — perceuptual evaluations

(1) Fréchet Inception Distance(↓)

DBL-Decompose achieves the superior perceptual fidelity for both foreground and background, as evidenced by consistently lower FID and more visually plausible results.

ModelKubric (BG)Movie (BG)LayeredVid-Benchmark (FG)
Gen-Omnimatte16.59816.34460.826
OmnimatteZero96.07736.42846.882
DBL-Decompose15.81514.36432.967

(2) Human preference study — video object insertion

DBL-Insert is consistently preferred over all competing methods across every evaluation criterion, demonstrating superior insertion quality and better alignment with the intended editing objective.

ModelSubject Consistency(↑)Insertion Rationality(↑)Prompt Alignment(↑)Overall Quality(↑)
AnyV2V1.0%1.0%1.5%1.5%
ReVideo0.5%1.0%1.0%1.5%
VACE-Inp17.5%10.0%21.0%15.5%
VACE-C15.0%11.5%18.5%13.0%
VACE-F9.5%8.5%7.5%7.0%
DBL-Insert56.5%68.0%50.5%61.5%

(3) Human preference study — video layer decomposition

DBL-Decompose receives the highest preference for both background and foreground reconstruction by a large margin, indicating that our explicit layer modeling produces more visually plausible decomposed layers than existing methods.

ModelBackground Quality(↑)Foreground Quality(↑)
Gen-Omnimatte8.0%19.0%
OmnimatteZero20.5%6.0%
DBL-Decompose71.5%75.0%

(4) Supplementary video

The supplementary video showcases more qualitative examples and ablations, along with various applications including layer-based video editing, scene-aware video editing, and other applications enabled by our layered video representation.

09 — Cite this work

BibTeX

@article{han2026explicitlayer,
  title     = {Explicit Layer Modeling for Video Object Insertion and Video Layer Decomposition},
  author    = {Han, kyujin and Shin, seungjoo and Cho, sunghyun},
  journal   = {arxiv},
  year      = {2026},
}