Digital zoom transitions between dual cameras often exhibit conspicuous discontinuities in geometric structure and chromatic consistency, degrading the user experience. While recent dual-camera smooth zoom (DCSZ) methods attempt to mitigate this by fine-tuning frame interpolation (FI) models on DCSZ data, they struggle with the large cross-view disparities and complex geometric transformations. Considering that the generative prior of diffusion models is suitable for addressing this problem, we explore their application to DCSZ. However, naively applying existing diffusion-based FI models still yields low-fidelity transitions due to insufficient conditional guidance, high-frequency information loss during VAE encoding, as well as inadequate temporal consistency. To address this, we propose ZoomDiff, a high-fidelity diffusion model that leverages dual-camera inputs in both latent and pixel spaces for photo-realistic transitions. Specifically, we first strengthen dual-image conditional guidance during the multi-step denoising process to improve geometric consistency. Then we inject flow-aligned multi-scale features from the VAE encoder into the VAE decoder to recover high-frequency details, where flow-guided temporal consistency supervision are introduced to produce more smooth transitions. Extensive experiments on both synthetic and real-world datasets demonstrate that ZoomDiff outperforms state-of-the-art methods quantitatively and qualitatively. Codes will be publicly available.
ZoomDiff leverages dual-camera inputs in both latent and image spaces for high-fidelity smooth zooming. We strengthen dual-image conditioning during diffusion denoising to improve geometric consistency, and inject flow-aligned multi-scale reference features into the VAE decoder to recover high-frequency details. A flow-guided temporal consistency objective further encourages smooth and temporally coherent transitions.
We present qualitative results on both real-world and synthetic datasets across different difficulty levels. ZoomDiff produces smooth zoom transitions while preserving geometric structures and fine-grained visual details.
We compare ZoomDiff with representative frame interpolation and diffusion-based methods. ZoomDiff better handles large cross-view disparities and complex geometric transformations, producing more consistent and visually pleasing smooth zoom transitions.
@article{zhang2026zoomdiff,
title = {ZoomDiff: A High-Fidelity Diffusion Model for Dual-Camera Smooth Zooming},
author = {Jiayi Zhang and Renlong Wu and Yukang Ding and Sibin Deng and Wangmeng Zuo},
year = {2026},
}