DNI: Dilutional Noise Initialization for Diffusion Video Editing

10citations

arXiv:2409.13037 PDF

citations

#1066

in ECCV 2024

of 2387 papers

Top Authors

Data Points

Top Authors

Sunjae Yoon Gwanhyeong Koo Ji Woo Hong Chang Yoo

Topics

diffusion video editing text-based editing non-rigid editing noise initialization latent noise manipulation motion change editing structural modification video generation

Abstract

Text-based diffusion video editing systems have been successful in performing edits with high fidelity and textual alignment. However, this success is limited to rigid-type editing such as style transfer and object overlay, while preserving the original structure of the input video. This limitation stems from an initial latent noise employed in diffusion video editing systems. The diffusion video editing systems prepare initial latent noise to edit by gradually infusing Gaussian noise onto the input video. However, we observed that the visual structure of the input video still persists within this initial latent noise, thereby restricting non-rigid editing such as motion change necessitating structural modifications. To this end, this paper proposes Dilutional Noise Initialization (DNI) framework which enables editing systems to perform precise and dynamic modification including non-rigid editing. DNI introduces a concept of `noise dilution' which adds further noise to the latent noise in the region to be edited to soften the structural rigidity imposed by input video, resulting in more effective edits closer to the target prompt. Extensive experiments demonstrate the effectiveness of the DNI framework.

Citation History

Jan 26, 2026

9+9

Jan 27, 2026

Feb 3, 2026

10+1

Feb 13, 2026