MIT 9.520
Statistical Learning Theory and Applications
A graduate course on the foundations of learning, from classical regularization and kernel methods to modern questions about deep networks, optimization, generalization, and the theory needed to understand today's AI systems.
Instructors
- Tomaso Poggio - Instructor
- Lorenzo Rosasco - Instructor
- Pierfrancesco Beneventano - Instructor
Teaching Assistants
- Qianli Liao - TA
- Yulu Gan - TA
- Federico V. Cortesi - TA
- Mahmoud Abdelmoneum - TA
Course Description
Learning theory as a route to understanding intelligence
Understanding intelligence, and how to replicate it in machines, is arguably one of the greatest problems in science. Learning, through its theory and computational implementations, lies at the core of intelligence.
Over the last two decades, AI systems have learned to solve complex tasks that were once the exclusive domain of biological organisms: computer vision, speech recognition, and natural language understanding and generation. These successes are driven by algorithms trained from examples rather than explicitly programmed to solve each task. This course, probably the oldest continuously running machine learning course at MIT, has been pushing toward this shift since its inception in 1992.
Yet a comprehensive theory of learning, especially one that explains the empirical puzzles raised by deep learning, remains incomplete. Such a theory could enable more powerful learning approaches, guide the use of learning algorithms in high-stakes settings, and inform our understanding of human intelligence.
Part I: Classical SLT
- Classical regularization and regularized least squares
- Kernel methods and support vector machines
- Logistic regression, squared loss, and exponential loss
- Large margin theory and minimum norm solutions
- Stochastic gradient methods
- Overparameterization and implicit regularization for linear models
- Approximation and estimation errors
Part II: Neural Networks
- Approximation: why deep networks can be universal parametric approximators and avoid the curse of dimensionality
- Optimization: how weights evolve in time and across layers during training
- Learning theory: if and how generalization in deep networks can be explained by implicit complexity control and sparse compositional structure
- What breaks and what survives as model, task, and data complexity increase
- How recent machine learning theory reconnects to the Statistical Learning Theory framework
- New paradigms in learning, including diffusion models and autoregressive transformers
Calendar and Syllabus
Fall 2026 course schedule
This schedule follows the official Fall 2026 syllabus. Archived Fall 2025 lecture decks are linked where they match the scheduled topic. TP = Tomaso Poggio, LR = Lorenzo Rosasco, and PB = Pierfrancesco Beneventano.
MIT calendar notes
The class starts on Thursday, September 10, 2026 because MIT's first day of classes is Wednesday, September 9, 2026.
There is no 9.520 meeting on Tuesday, October 13, 2026 because MIT holds a Monday schedule that day, and there is no class on Thursday, November 26, 2026 for Thanksgiving.
Slide archive
The lecture rows below link directly to the slide decks from last year whenever slides are available.
Thu, Oct 15
From Classical to Modern Deep Learning
MIT follows a Monday schedule on Tuesday, October 13, 2026, so there is no 9.520 meeting that day.
PB + TP
Tue, Nov 3
Multi Index models: SGD and Computational/Statistical Gaps
TBA
Thu, Nov 5
Genericity and Training on Polynomial
TP + PB
Thu, Nov 12
Intro to LLM and Pretraining
Blake Woodworth
Tue, Nov 17
Post-training LLMs
PB + Yulu
Thu, Nov 19
Agentic Systems
PB + Mahmoud Abdelmoneum
Tue, Nov 24
Neural Networks in the Financial Markets
Atlas Wang
Tue, Dec 1
Project Presentations
Course staff
Wed, Dec 2
Panel
Course staff
Tue, Dec 8
TBA
TBA
Thu, Dec 10
TBA
MIT's fall 2026 last day of classes and the current final paper deadline.
TBA
Timeline
Deliverables are designed to move projects early
Deadlines are intended to make the project research process concrete: choose a problem, understand the literature, plan the path, show early evidence, present the work, and submit a paper.
Friday deadlines are due by 11:59 PM local time. After the first missed deadline, each additional missed deadline carries a 3-point penalty, plus 3 points when a submission is more than three days late. Late presentations and final papers are not accepted.
A project that overlaps with a group member's current or past research or coursework carries a 15-point penalty for the whole group. Contact the teaching staff if you have any concerns.
Throughout the semester
Attendance quizzes (5 points)
Attend class regularly and answer the in-class quizzes correctly. Completing quizzes for another student is a serious breach of MIT rules.
September 25, 2026
Form filled (no grade)
Complete the Google Form with your group of one or two people and indicate either three projects from the official list or two listed projects plus one self-proposed project. For a self-proposed project, attach a PDF proposal of approximately 0.5-1 page.
October 2, 2026
Literature reviews and implications (up to 15 points)
As a group, submit one 3-4 page document for each of the three indicated projects: 2-3 pages of substantial literature review and one final page with at least three detailed implications for machine learning theory and practice. Each document is graded from 1 to 5.
October 9, 2026
Plan (no grade, up to 5 bonus points)
Submit 1-2 pages stating the selected problem, what you plan to achieve, how you plan to achieve it through proofs or experiments, and a realistic timeline. Earn up to 3 bonus points for selecting a listed problem and up to 2 for a particularly strong plan.
October 30, 2026
Initial checkpoint (up to 10 points)
Submit a short 1-2 page commentary with initial results, such as first plots or a proof sketch, demonstrating that the project started early and is on track.
First two weeks of November 2026
Project discussion (5 points)
Attend office hours to discuss your work with the course staff. Sign up through Calendly.
December 1-4, 2026
Presentation (up to 25 points)
Give an 8-minute-sharp group presentation covering motivation, related work, the open question, results, and implications. Upload up to 10 content slides by the end of the presentation day. The presentation is worth 20 points and question answering 5 points; exceeding the time limit carries a 5-point penalty.
December 11, 2026
Final paper (up to 40 points)
Submit the final paper and a public code repository. The main text must be 5-9 pages, with approximately 8 pages expected, followed by references and an optional appendix. Provide a Python notebook runnable in Google Colab with a small reproducible experiment.
Projects
Research questions for the semester
The project area will stay visible on the course page, but we are leaving it empty for the moment while the project material is prepared.
Coming soon...
References
Reading list and supporting resources
This section intentionally sits at the very end of the page. The references below follow last year's syllabus structure and have been cleaned up for consistency.
Notes covering the classes will be provided in the form of independent chapters from a draft set of lecture notes. The books and papers listed below are useful general references, especially from the theoretical viewpoint, and additional suggested readings can be attached to individual classes as needed.
Book (draft)
Machine Learning: a Regularization Approach, MIT 9.520 Lecture Notes
L. Rosasco and T. Poggio, manuscript, Dec. 2017 (provided).
Primary References
Understanding Machine Learning: From Theory to Algorithms
S. Shalev-Shwartz and S. Ben-David, Cambridge University Press, 2014.
Introduction to Statistical Learning Theory
O. Bousquet, S. Boucheron, and G. Lugosi. In Advanced Lectures on Machine Learning, LNCS 3176, pp. 169-207, Springer, 2004.
On The Mathematical Foundations of Learning
F. Cucker and S. Smale, Bulletin of the American Mathematical Society, 2002.
A Probabilistic Theory of Pattern Recognition
L. Devroye, L. Gyorfi, and G. Lugosi, Springer, 1997.
Regularization Networks and Support Vector Machines
T. Evgeniou, M. Pontil, and T. Poggio, Advances in Computational Mathematics, 2000.
The Mathematics of Learning: Dealing with Data
T. Poggio and S. Smale, Notices of the AMS, 2003.
Statistical Learning Theory
V. N. Vapnik, Wiley, 1998.
Papers of Interest
Why and When Can Deep-but Not Shallow-Networks Avoid the Curse of Dimensionality: A Review
T. Poggio, H. Mhaskar, L. Rosasco, B. Miranda, and Q. Liao, International Journal of Automation and Computing, 2017.
Compositional sparsity of learnable functions
T. Poggio and M. Fraser, Bulletin of the American Mathematical Society, 2024.
Dynamics in Deep Classifiers Trained with the Square Loss: Normalization, Low Rank, Neural Collapse, and Generalization Bounds
M. Xu, A. Rangamani, Q. Liao, T. Galanti, and T. Poggio, Research, 2023.
Deep learning
Y. LeCun, Y. Bengio, and G. Hinton, Nature, 521(7553):436-444, 2015.
Mastering the game of Go with deep neural networks and tree search
D. Silver et al., Nature, 529(7587):484-489, 2016.
Highly accurate protein structure prediction with AlphaFold
J. Jumper et al., Nature, 596(7873):583-589, 2021.
Attention Is All You Need
A. Vaswani et al., Advances in Neural Information Processing Systems 30, 2017.