Research

Learning and inference from survey data

My research aims to draw reliable conclusions from data collected through surveys. I am interested both in finite population quantities, such as totals, means or quantiles, and in superpopulation quantities, such as a regression function. This work is organized around four themes, described below: machine learning for finite population inference, statistical learning from survey data, missing data and high-dimensional inference.

Surveys remain the backbone of official statistics, from employment figures to health and economic indicators. They increasingly have to deliver reliable estimates in difficult settings: falling response rates, many auxiliary variables, highly nonlinear relationships, or domains containing few sampled units (small area estimation). Combining surveys with other sources of data (data integration) adds further challenges. In such settings, classical procedures may lose some of their properties, and newer ones, including methods from statistical learning, raise new questions.

Survey data add a further difficulty: they are often neither independent nor identically distributed. Units are typically drawn without replacement, which creates dependence. They are often selected with unequal probabilities, so that observations are not identically distributed, and the sampling design can even be informative. Classical survey theory was built to handle these features. Many of the newer methodologies, such as machine learning, were instead developed for independent and identically distributed data. They offer new ways to address classical survey problems, as well as emerging ones, provided that both the methods and their theory are adapted to the sampling design.

My aim is to understand and improve current methodologies in these settings: to derive the asymptotic properties of estimators, to quantify their uncertainty, and to identify when they are optimal.

Research themes

Machine learning for finite population inference

Estimating finite population quantities is the core task of survey sampling. Predictions from flexible models can make these estimates much more precise.

I am interested in estimators that incorporate such predictions, for example through model-assisted estimation. I aim to study when they keep the guarantees of design-based inference, whatever the quality of the predictions, and when they are efficient.

Related papers

Statistical learning from survey data

Survey data are increasingly used to fit predictive models. Yet standard learning methods assume independent and identically distributed data, and can be misleading under complex or informative designs.

I am interested in learning methods that estimate superpopulation quantities, such as a regression function, from survey data, and in their theoretical properties.

Related papers

Missing data

Nonresponse affects nearly every survey and keeps increasing. Left uncorrected, it can bias estimates.

I am interested in procedures that account for missing data, imputation and reweighting being two examples, possibly based on machine learning, and in how to obtain valid inference in their presence.

Related papers

High-dimensional inference

The number of available auxiliary variables keeps growing, sometimes to the point of being comparable with the sample size. Classical asymptotic results may then no longer apply.

I aim to understand how estimators and their variance estimators behave in such regimes, and to develop procedures that remain reliable.

Related papers