MIT 9.520

Statistical Learning Theory and Applications

A graduate course on the foundations of learning, from classical regularization and kernel methods to modern questions about deep networks, optimization, generalization, and the theory needed to understand today's AI systems.

Instructors

Teaching Assistants

  • Yulu Gan - TA
  • Federico V. Cortesi - TA
  • Mahmoud Abdelmoneum - TA

Course Description

Learning theory as a route to understanding intelligence

Understanding intelligence, and how to replicate it in machines, is arguably one of the greatest problems in science. Learning, through its theory and computational implementations, lies at the core of intelligence.

Over the last two decades, AI systems have learned to solve complex tasks that were once the exclusive domain of biological organisms: computer vision, speech recognition, and natural language understanding and generation. These successes are driven by algorithms trained from examples rather than explicitly programmed to solve each task. This course, probably the oldest continuously running machine learning course at MIT, has been pushing toward this shift since its inception in 1992.

Yet a comprehensive theory of learning, especially one that explains the empirical puzzles raised by deep learning, remains incomplete. Such a theory could enable more powerful learning approaches, guide the use of learning algorithms in high-stakes settings, and inform our understanding of human intelligence.

Part I: Classical SLT

  • Classical regularization and regularized least squares
  • Kernel methods and support vector machines
  • Logistic regression, squared loss, and exponential loss
  • Large margin theory and minimum norm solutions
  • Stochastic gradient methods
  • Overparameterization and implicit regularization for linear models
  • Approximation and estimation errors

Part II: Neural Networks

  • Approximation: why deep networks can be universal parametric approximators and avoid the curse of dimensionality
  • Optimization: how weights evolve in time and across layers during training
  • Learning theory: if and how generalization in deep networks can be explained by implicit complexity control and sparse compositional structure
  • What breaks and what survives as model, task, and data complexity increase
  • How recent machine learning theory reconnects to the Statistical Learning Theory framework
  • New paradigms in learning, including diffusion models and autoregressive transformers

Rules and Expectations

Project-centered grading and research practice

The current format removes traditional problem sets to give more time to projects and introduces an oral presentation. The goal is to understand how well students own their project, how clearly they can position it within Statistical Learning Theory, and how carefully they can connect theory, experiments, and implications.

Prerequisites

Part II is designed for students with a good background in ML. The course uses calculus, linear algebra, probability, basic optimization, and some functional or convex analysis. For course 6 students, expected background includes 6.041, 18.06, and an introductory ML course such as 6.036, 6.401, or 6.867.

AI Tools

Students are expected to use modern LLM-based tools when useful, but must still read the relevant papers and be able to explain, rework, and defend the work offline.

Teams

Projects may be individual or in teams of two. Groups of two are encouraged. Multiple teams may work on related problems, but authorship and submission plans should be coordinated with the staff.

Timeline

Deliverables are designed to move projects early

Deadlines are intended to make the project research process concrete: choose a problem, understand the literature, plan the path, show early evidence, present the work, and submit a paper.

September 25, 2026

Groups and proposals

Submit your group and indicate three project choices, or two listed projects plus one self-proposed project.

October 9, 2026

Literature reviews and implications

For each of the three indicated projects, submit 3-4 pages covering related work and consequences for theory and practice.

October 16, 2026

Project plan

Submit a concise plan explaining the chosen problem, expected result, proof or experiment strategy, and milestones.

October 30, 2026

Initial checkpoint

Submit early results: first plots, proof sketches, ablations, or a short account of what has been learned.

November 3-12, 2026

Project discussions

Meet during office hours to discuss progress, roadblocks, positioning, and next steps.

December 1-3, 2026

Oral presentation

Give an 8-minute presentation with up to 10 content slides covering motivation, related work, results, and implications.

December 10, 2026

Final paper

Submit the final paper and a link to a public code repository or runnable notebook.

Grading

Participation plus project work

The grading scheme is project-based: 10 points for participation and up to 90 points for project-related activities, with possible bonus points for a strong project plan.

Participation

Up to 10 points for active attendance, engagement, and project discussion.

Literature reviews

Up to 15 points across the three project reviews and implications documents.

Initial checkpoint

Up to 10 points for early evidence that the project is on track.

Presentation

Up to 25 points for motivation, results, clarity, organization, and answering questions.

Final paper

Up to 40 points for execution, positioning, clarity, novelty, limitations, and significance.

Projects

Research questions for the semester

The project area will stay visible on the course page, but we are leaving it empty for the moment while the project material is prepared.

Coming soon...

Calendar and Syllabus

Fall 2026 meeting calendar using the ordered 2025 lecture sequence

This schedule follows MIT's official fall 2026 class calendar for Tuesday/Thursday meetings and places the lecture decks from last year in the same order. TP = Tomaso Poggio, LR = Lorenzo Rosasco, and PB = Pierfrancesco Beneventano.

MIT calendar notes

The class starts on Thursday, September 10, 2026 because MIT's first day of classes is Wednesday, September 9, 2026.

There is no 9.520 meeting on Tuesday, October 13, 2026 because MIT holds a Monday schedule that day, and there is no class on Thursday, November 26, 2026 for Thanksgiving.

Slide archive

The lecture rows below link directly to the slide decks from last year whenever slides are available.

Thu, Sep 10

Course overview, logistics, and why theory

TP / LR / PB

Tue, Sep 15

Statistical Learning Theory

LR

Thu, Sep 17

Least squares and overparameterization

LR

Tue, Sep 22

Logistic regression and SGD

LR

Thu, Sep 24

Implicit Regularization

LR

Tue, Sep 29

Neural networks

LR

Thu, Oct 1

Random Features, NTK & RKHS

LR

Tue, Oct 6

Infinite Width Neural Networks & RKBS

LR

Thu, Oct 8

Learning bounds for linear least squares

LR

Thu, Oct 15

Learning bounds for ERM

MIT follows a Monday schedule on Tuesday, October 13, 2026, so there is no 9.520 meeting that day.

LR

Tue, Oct 20

Sequential prediction as learning dynamical systems

LR

Thu, Oct 22

From classical to modern deep learning

PB + TP

Tue, Oct 27

Deep Learning: approximation theory

TP

Thu, Oct 29

Sparse Compositionality

TP

Tue, Nov 3

Deep Learning Theory: Optimization

PB

Thu, Nov 5

Training neural networks: trainability

PB

Tue, Nov 10

Training neural networks: how we train

PB

Thu, Nov 12

Where training goes and its stability

PB

Tue, Nov 17

Towards a Learning Theory of Grammars

Dan Mitropolsky (guest)

Thu, Nov 19

Open

Course staff

Open

Tue, Nov 24

Open

Course staff

Open

Tue, Dec 1

Open

Course staff

Open

Thu, Dec 3

Open

Course staff

Open

Tue, Dec 8

Open

Course staff

Open

Thu, Dec 10

Open

MIT's fall 2026 last day of classes and the current final paper deadline.

Course staff

Open

References

Reading list and supporting resources

This section intentionally sits at the very end of the page. The references below follow last year's syllabus structure and have been cleaned up for consistency.

Notes covering the classes will be provided in the form of independent chapters from a draft set of lecture notes. The books and papers listed below are useful general references, especially from the theoretical viewpoint, and additional suggested readings can be attached to individual classes as needed.

Book (draft)

Machine Learning: a Regularization Approach, MIT 9.520 Lecture Notes

L. Rosasco and T. Poggio, manuscript, Dec. 2017 (provided).

Primary References

Understanding Machine Learning: From Theory to Algorithms

S. Shalev-Shwartz and S. Ben-David, Cambridge University Press, 2014.

Introduction to Statistical Learning Theory

O. Bousquet, S. Boucheron, and G. Lugosi. In Advanced Lectures on Machine Learning, LNCS 3176, pp. 169-207, Springer, 2004.

On The Mathematical Foundations of Learning

F. Cucker and S. Smale, Bulletin of the American Mathematical Society, 2002.

A Probabilistic Theory of Pattern Recognition

L. Devroye, L. Gyorfi, and G. Lugosi, Springer, 1997.

Regularization Networks and Support Vector Machines

T. Evgeniou, M. Pontil, and T. Poggio, Advances in Computational Mathematics, 2000.

The Mathematics of Learning: Dealing with Data

T. Poggio and S. Smale, Notices of the AMS, 2003.

Statistical Learning Theory

V. N. Vapnik, Wiley, 1998.

Resources and links