Core paper:
High-Resolution Image Synthesis with Latent Diffusion Models
Robin Rombach et al., 2021 / 2022
Why it is essential:
This is the core paper behind Stable Diffusion. It applies diffusion in the latent space of a pretrained autoencoder, greatly reducing computation, and uses cross-attention for text and other conditioning signals. arXiv
Topics:
Why not diffuse directly in pixel space
VAE latent space
Cross-attention
Text conditioning
Inpainting, super-resolution, image-to-image
Key concept:
image → VAE encoder → latent
diffusion denoising in latent space
latent → VAE decoder → imageEngineering connection:
Stable Diffusion
SDXL
Image-to-image
Inpainting
ControlNet
Multi-image reference