MIT 9.520

Statistical Learning Theory and Applications

A graduate course on the foundations of learning, from classical regularization and kernel methods to modern questions about deep networks, optimization, generalization, and the theory needed to understand today's AI systems.

Instructors

Teaching Assistants

  • Qianli Liao - TA
  • Yulu Gan - TA
  • Federico V. Cortesi - TA
  • Mahmoud Abdelmoneum - TA

Course Description

Learning theory as a route to understanding intelligence

Understanding intelligence, and how to replicate it in machines, is arguably one of the greatest problems in science. Learning, through its theory and computational implementations, lies at the core of intelligence.

Over the last two decades, AI systems have learned to solve complex tasks that were once the exclusive domain of biological organisms: computer vision, speech recognition, and natural language understanding and generation. These successes are driven by algorithms trained from examples rather than explicitly programmed to solve each task. This course, probably the oldest continuously running machine learning course at MIT, has been pushing toward this shift since its inception in 1992.

Yet a comprehensive theory of learning, especially one that explains the empirical puzzles raised by deep learning, remains incomplete. Such a theory could enable more powerful learning approaches, guide the use of learning algorithms in high-stakes settings, and inform our understanding of human intelligence.

Part I: Classical SLT

  • Classical regularization and regularized least squares
  • Kernel methods and support vector machines
  • Logistic regression, squared loss, and exponential loss
  • Large margin theory and minimum norm solutions
  • Stochastic gradient methods
  • Overparameterization and implicit regularization for linear models
  • Approximation and estimation errors

Part II: Neural Networks

  • Approximation: why deep networks can be universal parametric approximators and avoid the curse of dimensionality
  • Optimization: how weights evolve in time and across layers during training
  • Learning theory: if and how generalization in deep networks can be explained by implicit complexity control and sparse compositional structure
  • What breaks and what survives as model, task, and data complexity increase
  • How recent machine learning theory reconnects to the Statistical Learning Theory framework
  • New paradigms in learning, including diffusion models and autoregressive transformers

Calendar and Syllabus

Fall 2026 course schedule

This schedule follows the official Fall 2026 syllabus. Archived Fall 2025 lecture decks are linked where they match the scheduled topic. TP = Tomaso Poggio, LR = Lorenzo Rosasco, and PB = Pierfrancesco Beneventano.

MIT calendar notes

The class starts on Thursday, September 10, 2026 because MIT's first day of classes is Wednesday, September 9, 2026.

There is no 9.520 meeting on Tuesday, October 13, 2026 because MIT holds a Monday schedule that day, and there is no class on Thursday, November 26, 2026 for Thanksgiving.

Slide archive

The lecture rows below link directly to the slide decks from last year whenever slides are available.

Thu, Sep 10

Course overview, logistics, and why theory

TP / LR / PB

Tue, Sep 15

Statistical Learning Theory

LR

Thu, Sep 17

Least squares and overparameterization

LR

Tue, Sep 22

Logistic regression and SGD

LR

Thu, Sep 24

Implicit Regularization

LR

Tue, Sep 29

Neural Networks

LR

Thu, Oct 1

Random Features, NTK & RKHS

LR

Tue, Oct 6

Infinite-Width Neural Networks & RKBS

LR

Thu, Oct 8

Learning Theory: Approximation and Estimation Errors

LR

Thu, Oct 15

From Classical to Modern Deep Learning

MIT follows a Monday schedule on Tuesday, October 13, 2026, so there is no 9.520 meeting that day.

PB + TP

Tue, Oct 20

Deep Learning: Approximation Theory

TP

Thu, Oct 22

Sparse Compositionality and Generalization

TP

Tue, Oct 27

Overview of Convex + Non-convex optimization

PB

Thu, Oct 29

Training Neural Networks: How We Train

PB

Tue, Nov 3

Multi Index models: SGD and Computational/Statistical Gaps

TBA

No slides

Thu, Nov 5

Genericity and Training on Polynomial

TP + PB

No slides

Tue, Nov 10

Location of Convergence and Stability of Training

PB

Thu, Nov 12

Intro to LLM and Pretraining

Blake Woodworth

No slides

Tue, Nov 17

Post-training LLMs

PB + Yulu

No slides

Thu, Nov 19

Agentic Systems

PB + Mahmoud Abdelmoneum

No slides

Tue, Nov 24

Neural Networks in the Financial Markets

Atlas Wang

No slides

Tue, Dec 1

Project Presentations

Course staff

No slides

Wed, Dec 2

Panel

Course staff

No slides

Tue, Dec 8

TBA

TBA

No slides

Thu, Dec 10

TBA

MIT's fall 2026 last day of classes and the current final paper deadline.

TBA

No slides

Timeline

Deliverables are designed to move projects early

Deadlines are intended to make the project research process concrete: choose a problem, understand the literature, plan the path, show early evidence, present the work, and submit a paper.

Friday deadlines are due by 11:59 PM local time. After the first missed deadline, each additional missed deadline carries a 3-point penalty, plus 3 points when a submission is more than three days late. Late presentations and final papers are not accepted.

A project that overlaps with a group member's current or past research or coursework carries a 15-point penalty for the whole group. Contact the teaching staff if you have any concerns.

Throughout the semester

Attendance quizzes (5 points)

Attend class regularly and answer the in-class quizzes correctly. Completing quizzes for another student is a serious breach of MIT rules.

September 25, 2026

Form filled (no grade)

Complete the Google Form with your group of one or two people and indicate either three projects from the official list or two listed projects plus one self-proposed project. For a self-proposed project, attach a PDF proposal of approximately 0.5-1 page.

October 2, 2026

Literature reviews and implications (up to 15 points)

As a group, submit one 3-4 page document for each of the three indicated projects: 2-3 pages of substantial literature review and one final page with at least three detailed implications for machine learning theory and practice. Each document is graded from 1 to 5.

October 9, 2026

Plan (no grade, up to 5 bonus points)

Submit 1-2 pages stating the selected problem, what you plan to achieve, how you plan to achieve it through proofs or experiments, and a realistic timeline. Earn up to 3 bonus points for selecting a listed problem and up to 2 for a particularly strong plan.

October 30, 2026

Initial checkpoint (up to 10 points)

Submit a short 1-2 page commentary with initial results, such as first plots or a proof sketch, demonstrating that the project started early and is on track.

First two weeks of November 2026

Project discussion (5 points)

Attend office hours to discuss your work with the course staff. Sign up through Calendly.

December 1-4, 2026

Presentation (up to 25 points)

Give an 8-minute-sharp group presentation covering motivation, related work, the open question, results, and implications. Upload up to 10 content slides by the end of the presentation day. The presentation is worth 20 points and question answering 5 points; exceeding the time limit carries a 5-point penalty.

December 11, 2026

Final paper (up to 40 points)

Submit the final paper and a public code repository. The main text must be 5-9 pages, with approximately 8 pages expected, followed by references and an optional appendix. Provide a Python notebook runnable in Google Colab with a small reproducible experiment.

Projects

Research questions for the semester

The project area will stay visible on the course page, but we are leaving it empty for the moment while the project material is prepared.

Coming soon...

References

Reading list and supporting resources

This section intentionally sits at the very end of the page. The references below follow last year's syllabus structure and have been cleaned up for consistency.

Notes covering the classes will be provided in the form of independent chapters from a draft set of lecture notes. The books and papers listed below are useful general references, especially from the theoretical viewpoint, and additional suggested readings can be attached to individual classes as needed.

Book (draft)

Machine Learning: a Regularization Approach, MIT 9.520 Lecture Notes

L. Rosasco and T. Poggio, manuscript, Dec. 2017 (provided).

Primary References

Understanding Machine Learning: From Theory to Algorithms

S. Shalev-Shwartz and S. Ben-David, Cambridge University Press, 2014.

Introduction to Statistical Learning Theory

O. Bousquet, S. Boucheron, and G. Lugosi. In Advanced Lectures on Machine Learning, LNCS 3176, pp. 169-207, Springer, 2004.

On The Mathematical Foundations of Learning

F. Cucker and S. Smale, Bulletin of the American Mathematical Society, 2002.

A Probabilistic Theory of Pattern Recognition

L. Devroye, L. Gyorfi, and G. Lugosi, Springer, 1997.

Regularization Networks and Support Vector Machines

T. Evgeniou, M. Pontil, and T. Poggio, Advances in Computational Mathematics, 2000.

The Mathematics of Learning: Dealing with Data

T. Poggio and S. Smale, Notices of the AMS, 2003.

Statistical Learning Theory

V. N. Vapnik, Wiley, 1998.

Resources and links