MIT 9.520
Statistical Learning Theory and Applications
A graduate course on the foundations of learning, from classical regularization and kernel methods to modern questions about deep networks, optimization, generalization, and the theory needed to understand today's AI systems.
Instructors
- Tomaso Poggio - Instructor
- Lorenzo Rosasco - Instructor
- Pierfrancesco Beneventano - Instructor
Teaching Assistants
- Yulu Gan - TA
- Federico V. Cortesi - TA
- Mahmoud Abdelmoneum - TA
Course Description
Learning theory as a route to understanding intelligence
Understanding intelligence, and how to replicate it in machines, is arguably one of the greatest problems in science. Learning, through its theory and computational implementations, lies at the core of intelligence.
Over the last two decades, AI systems have learned to solve complex tasks that were once the exclusive domain of biological organisms: computer vision, speech recognition, and natural language understanding and generation. These successes are driven by algorithms trained from examples rather than explicitly programmed to solve each task. This course, probably the oldest continuously running machine learning course at MIT, has been pushing toward this shift since its inception in 1992.
Yet a comprehensive theory of learning, especially one that explains the empirical puzzles raised by deep learning, remains incomplete. Such a theory could enable more powerful learning approaches, guide the use of learning algorithms in high-stakes settings, and inform our understanding of human intelligence.
Part I: Classical SLT
- Classical regularization and regularized least squares
- Kernel methods and support vector machines
- Logistic regression, squared loss, and exponential loss
- Large margin theory and minimum norm solutions
- Stochastic gradient methods
- Overparameterization and implicit regularization for linear models
- Approximation and estimation errors
Part II: Neural Networks
- Approximation: why deep networks can be universal parametric approximators and avoid the curse of dimensionality
- Optimization: how weights evolve in time and across layers during training
- Learning theory: if and how generalization in deep networks can be explained by implicit complexity control and sparse compositional structure
- What breaks and what survives as model, task, and data complexity increase
- How recent machine learning theory reconnects to the Statistical Learning Theory framework
- New paradigms in learning, including diffusion models and autoregressive transformers
Rules and Expectations
Project-centered grading and research practice
The current format removes traditional problem sets to give more time to projects and introduces an oral presentation. The goal is to understand how well students own their project, how clearly they can position it within Statistical Learning Theory, and how carefully they can connect theory, experiments, and implications.
Prerequisites
Part II is designed for students with a good background in ML. The course uses calculus, linear algebra, probability, basic optimization, and some functional or convex analysis. For course 6 students, expected background includes 6.041, 18.06, and an introductory ML course such as 6.036, 6.401, or 6.867.
AI Tools
Students are expected to use modern LLM-based tools when useful, but must still read the relevant papers and be able to explain, rework, and defend the work offline.
Teams
Projects may be individual or in teams of two. Groups of two are encouraged. Multiple teams may work on related problems, but authorship and submission plans should be coordinated with the staff.
Timeline
Deliverables are designed to move projects early
Deadlines are intended to make the project research process concrete: choose a problem, understand the literature, plan the path, show early evidence, present the work, and submit a paper.
September 25, 2026
Groups and proposals
Submit your group and indicate three project choices, or two listed projects plus one self-proposed project.
October 9, 2026
Literature reviews and implications
For each of the three indicated projects, submit 3-4 pages covering related work and consequences for theory and practice.
October 16, 2026
Project plan
Submit a concise plan explaining the chosen problem, expected result, proof or experiment strategy, and milestones.
October 30, 2026
Initial checkpoint
Submit early results: first plots, proof sketches, ablations, or a short account of what has been learned.
November 3-12, 2026
Project discussions
Meet during office hours to discuss progress, roadblocks, positioning, and next steps.
December 1-3, 2026
Oral presentation
Give an 8-minute presentation with up to 10 content slides covering motivation, related work, results, and implications.
December 10, 2026
Final paper
Submit the final paper and a link to a public code repository or runnable notebook.
Grading
Participation plus project work
The grading scheme is project-based: 10 points for participation and up to 90 points for project-related activities, with possible bonus points for a strong project plan.
Participation
Up to 10 points for active attendance, engagement, and project discussion.
Literature reviews
Up to 15 points across the three project reviews and implications documents.
Initial checkpoint
Up to 10 points for early evidence that the project is on track.
Presentation
Up to 25 points for motivation, results, clarity, organization, and answering questions.
Final paper
Up to 40 points for execution, positioning, clarity, novelty, limitations, and significance.
Projects
Research questions for the semester
The project area will stay visible on the course page, but we are leaving it empty for the moment while the project material is prepared.
Coming soon...
Calendar and Syllabus
Fall 2026 meeting calendar using the ordered 2025 lecture sequence
This schedule follows MIT's official fall 2026 class calendar for Tuesday/Thursday meetings and places the lecture decks from last year in the same order. TP = Tomaso Poggio, LR = Lorenzo Rosasco, and PB = Pierfrancesco Beneventano.
MIT calendar notes
The class starts on Thursday, September 10, 2026 because MIT's first day of classes is Wednesday, September 9, 2026.
There is no 9.520 meeting on Tuesday, October 13, 2026 because MIT holds a Monday schedule that day, and there is no class on Thursday, November 26, 2026 for Thanksgiving.
Slide archive
The lecture rows below link directly to the slide decks from last year whenever slides are available.
Thu, Oct 15
Learning bounds for ERM
MIT follows a Monday schedule on Tuesday, October 13, 2026, so there is no 9.520 meeting that day.
LR
Thu, Nov 19
Open
Course staff
Tue, Nov 24
Open
Course staff
Tue, Dec 1
Open
Course staff
Thu, Dec 3
Open
Course staff
Tue, Dec 8
Open
Course staff
Thu, Dec 10
Open
MIT's fall 2026 last day of classes and the current final paper deadline.
Course staff
References
Reading list and supporting resources
This section intentionally sits at the very end of the page. The references below follow last year's syllabus structure and have been cleaned up for consistency.
Notes covering the classes will be provided in the form of independent chapters from a draft set of lecture notes. The books and papers listed below are useful general references, especially from the theoretical viewpoint, and additional suggested readings can be attached to individual classes as needed.
Book (draft)
Machine Learning: a Regularization Approach, MIT 9.520 Lecture Notes
L. Rosasco and T. Poggio, manuscript, Dec. 2017 (provided).
Primary References
Understanding Machine Learning: From Theory to Algorithms
S. Shalev-Shwartz and S. Ben-David, Cambridge University Press, 2014.
Introduction to Statistical Learning Theory
O. Bousquet, S. Boucheron, and G. Lugosi. In Advanced Lectures on Machine Learning, LNCS 3176, pp. 169-207, Springer, 2004.
On The Mathematical Foundations of Learning
F. Cucker and S. Smale, Bulletin of the American Mathematical Society, 2002.
A Probabilistic Theory of Pattern Recognition
L. Devroye, L. Gyorfi, and G. Lugosi, Springer, 1997.
Regularization Networks and Support Vector Machines
T. Evgeniou, M. Pontil, and T. Poggio, Advances in Computational Mathematics, 2000.
The Mathematics of Learning: Dealing with Data
T. Poggio and S. Smale, Notices of the AMS, 2003.
Statistical Learning Theory
V. N. Vapnik, Wiley, 1998.
Papers of Interest
Why and When Can Deep-but Not Shallow-Networks Avoid the Curse of Dimensionality: A Review
T. Poggio, H. Mhaskar, L. Rosasco, B. Miranda, and Q. Liao, International Journal of Automation and Computing, 2017.
Compositional sparsity of learnable functions
T. Poggio and M. Fraser, Bulletin of the American Mathematical Society, 2024.
Dynamics in Deep Classifiers Trained with the Square Loss: Normalization, Low Rank, Neural Collapse, and Generalization Bounds
M. Xu, A. Rangamani, Q. Liao, T. Galanti, and T. Poggio, Research, 2023.
Deep learning
Y. LeCun, Y. Bengio, and G. Hinton, Nature, 521(7553):436-444, 2015.
Mastering the game of Go with deep neural networks and tree search
D. Silver et al., Nature, 529(7587):484-489, 2016.
Highly accurate protein structure prediction with AlphaFold
J. Jumper et al., Nature, 596(7873):583-589, 2021.
Attention Is All You Need
A. Vaswani et al., Advances in Neural Information Processing Systems 30, 2017.