TEACHING HOME

The Science of Deep Learning

CSCI-GA.3033-134 (5696) · Special Topics · Fall 2026
New York University, Courant Institute of Mathematical Sciences

Course description

Large language models generalize in ways we did not design and fail in ways we did not anticipate. This course is about understanding why.

This course will take an empirical approach to unpacking this question: building language models from scratch to develop hands-on fluency, then using that fluency to ask scientific questions about generalization. What does compute-optimal training actually mean? What do data quality, repetition, and domain diversity do to what a model learns? When and why does chain-of-thought help? And more.

Students will learn to build transformers from scratch, fit scaling laws, build data pipelines, and run fine-tuning experiments---not as engineering exercises, but as instruments for forming and testing hypotheses. By the end, students will be equipped to read the frontier literature critically, design experiments, and pursue research questions in one of the most consequential but incompletely understood areas of science.

Prerequisites

Machine learning, linear algebra, probability, and proficiency in Python and PyTorch.

Logistics

Lectures: Mondays, 2:45–4:45pm
Instructor: Pratyusha Sharma
Teaching Assistant: Jiahan Li
Contact: p.sharma@nyu.edu, jl17319@nyu.edu
Office Hours: Tuesday 9am–10am [Pratyusha Sharma], Friday 1pm-2pm [Jiahan Li]

Schedule

Schedule is tentative and will be updated as the semester progresses.

# Date Topic Papers
1 Sept 14 Framing the Course. LLMs as scientific objects; generalization as the organizing question.
2 Sept 21 Optimization. Guest lecture by John Langford. Adam. Muon. Dion.
3 Sept 28 Scaling Laws Kaplan scaling law. Hoffman scaling law. Data constrained scaling. Beating power law scaling [Bonus].
4 Oct 5 Training Dynamics. Grokking. Double descent. Edge of stability.
Oct 12 Fall Break — no class
5 Oct 14 (Legislative Monday) Post-Hoc Phenomena. Lottery ticket hypothesis. Linear mode connectivity. LASER / Better sub networks.
6 Oct 19 Data. Fineweb. Data mixing. Dataset difficulty.
7 Oct 26 Generalization. Lazy vs rich. Emergence. Sharpness aware minimization.
8 Nov 2 In-context learning. Influence of data. Induction heads. Role of demonstrations.
9 Nov 9 Reasoning. Chain-of-thought. DeepSeek-R1. Agentic RL.
10 Nov 16 Evaluation. HELM. Emergence mirage? Red-teaming.
11 Nov 23 Reinforcement Learning. Role of RL beyond SFT? RLxForgetting. Dr.GRPO.
12 Nov 30 Alignment InstructGPT. J-Space. Pacing the frontier.
13 Dec 7 Interpretability. Guest lecture by John Hewitt.
14 Dec 14 Project Presentations & Open Problems. What we still do not understand; student project presentations.