Skip to main content

From Data to Formulas: Statistical Learning for Scalable Symbolic Discovery

Meng Li

Associate Professor, Dept. of Statistics, Rice University

Meng Li

Abstract: Symbolic discovery, or learning compact mathematical expressions from data, offers a form of interpretability in which the fitted model is itself a scientific statement. In this talk, I treat symbolic discovery as a statistical learning problem and discuss scalable methods for searching large, structured expression spaces. The starting point is descriptor discovery in materials science, where candidate predictors are generated from primary physical features through compositions of algebraic operators, yielding a combinatorially large and highly correlated predictor space. An iterative nonparametric strategy exploits this compositional structure to identify, directly from the primary features, interpretable descriptors, or “materials genes,” associated with binding behavior in single-atom catalysis. The talk then turns to a key statistical component of scalable symbolic regression: variable selection with Bayesian tree ensembles. Here, selection accuracy can depend as much on posterior summarization as on the tree prior, and a simple, tuning-free posterior summary consistently improves existing approaches and substantially expands the reach of symbolic regression. A recent extension moves from regression functions to probability laws, where the discovered expression must satisfy the structural constraints of a valid distribution. These examples illustrate how statistical formulation through selection, summarization, and validity can make interpretable discovery feasible at the scale of modern scientific data.

Skip to content