The course provides an in-depth understanding of computer vision as the study of visual signals and learned representations. It maintains a strong foundation in classical methods while focusing on techniques relevant to state-of-the-art deep learning systems.

Students will gain knowledge of: image formation and camera and color models; Fourier analysis and convolution in the spatial and frequency domains; blurring, image derivatives, aliasing, and multi-scale representations; local feature detection and description; robust estimation of geometric transformations (matching, RANSAC, and homographies); camera models, calibration, stereo vision, triangulation, and epipolar geometry; convolutional neural networks, WaveNet, and inductive biases for vision; Vision Transformer and U-Net architectures for dense prediction; object detection and semantic segmentation; self-supervised representation learning (contrastive methods and DINOv3); vision-language models and parameter-efficient fine-tuning (PEFT) techniques; diffusion models; foundation models, multimodality, and vision-language-action models; robustness, uncertainty calibration, out-of-distribution detection, open-world recognition, person re-identification, and continual learning in vision systems.