YN/KZH/K & W are available.
Will be added: BP - R & JS
Core paper:
Learning Transferable Visual Models From Natural Language Supervision
Alec Radford et al., 2021
Why it is essential:
CLIP maps images and text into a shared semantic space, enabling text-driven visual control and multimodal retrieval.
Topics:
Image encoder
Text encoder
Contrastive learning
Shared embedding space
Why prompts can control images
Engineering connection:
Text-to-image generation
Image retrieval
Visual RAG-like retrieval
Multi-reference generation