Agenda

PhD defense Xiaolin Wang: Efficient Artificial Intelligence Based On Teacher-Student Paradigm

Wednesday 9 Sept., 2026, at 14:30 (Paris time) at Télécom Paris

Télécom Paris, 19 place Marguerite Perey F-91120 Palaiseau [getting there], amphi 2 and in videoconferencing

Jury

  • Sheng Yang, Professor, L2S, CentraleSupélec, Paris-Saclay University, France (Examiner)
  • François Malgouyres, Professor, IMT, University of Toulouse, France (Reviewer)
  • Helmut Bölcskei, Professor, ETH Zurich, Switzerland (Reviewer)
  • Anissa Mokraoui, Professor, L2TI, Institut Galilée, Sorbonne Paris North University, France (Examiner)
  • Michel Kieffer, Professor, L2S, CentraleSupélec, Paris-Saclay University (Examiner)
  • Vincent Corlay, PhD, Mitsubishi Electric R&D Centre Europe (Examiner)
  • Olivier Rioul, Professor, LTCI, Télécom Paris, Institut Polytechnique de Paris (PhD Supervisor)
  • Joseph Jean Boutros, Professor, Texas A&M University at Qatar (PhD Supervisor)
  • Pierre Duhamel, Professor, L2S, CentraleSupélec, Paris-Saclay University (Guest)

Abstract

This thesis addresses one question through the teacher-student paradigm: when can a small model do the work of a large one? Its central principle: the sample and parameter efficiency of a learned model is set by the approximation capacity of the task and by the route through which the student absorbs the teacher’s structural prior-implicitly via gradient descent on a factored parameterization, or explicitly via a task-aware preprocessing layer.

Learn more
The implicit route is worked out for shallow networks under low-rank teachers. For the quadratic activation, a closed-form optimal student at any prescribed rank exposes the irreducible bias of low-rank learning, and correlated inputs soften the sample-complexity threshold into a two-regime scaling law set by the teacher’s effective rank. For ReLU classification, a three-stage transition-random guessing, feature learning, rank-limited saturation-emerges with critical sample size linear in the input dimension and the shared effective rank. Capacity here reduces to the number of singular modes of the teacher.
The explicit route is worked out for neural channel decoders of binary linear codes. A syndrome-shifted near-origin reduction lowers the affine-piece complexity of the bit-MAP boundary from exponential in the code dimension to polynomial in the code length, with the degree set by the minimum distance. A code-aware preprocessing layer then closes the capacity gap through closed-form parity-check extrinsic statistics, yielding compact decoders-from a sigmoid perceptron to a pre-norm transformer-that approach the BCJR decoder on short and medium-length BCH codes and a new state of the art on a long BCH code.