PhaseDiff: Conditional Latent Diffusion for Multi-Phase Renal CT Synthesis
Computed tomography (CT) is widely used for evaluating renal anatomy and pathology, with multi-phase contrastenhanced imaging providing critical diagnostic information. However, acquiring multiple contrast phases requires repeated scans, in turn increasing radiation exposure. In addition, the development of AI methods for medical imaging is limited by the scarcity of datasets, often due to privacy and data sensitivity concerns. To address both of these limitations, we propose PhaseDiff, which learns the relationship between non-contrast and contrast-enhanced renal CT scans to synthesize arterial, portal venous, and delayed phases directly from non-contrast inputs. The framework is based on a conditional latent diffusion model that operates in a learned latent space, where a variational autoencoder encodes images into compact latent representations. Diffusion is applied only to the target contrast-enhanced latent, while the non-contrast latent is incorporated via channel-wise concatenation to provide spatial guidance. In addition, the desired contrast phase is introduced through cross-attention, enabling phase-controlled image synthesis. Trained on paired multi-phase CT data, PhaseDiff generates anatomically and phase-consistent contrast-enhanced images, highlighting the potential to reduce radiation exposure in renal CT workflows while also creating realistic synthetic datasets.