Research

My research connects statistics, AI, and scientific discovery. I develop methods for learning from complex data, with particular attention to differences across individuals and data sources, the structure of learned representations, and uncertainty in scientific conclusions. My work spans Bayesian modeling and computation, reinforcement learning, graph and multimodal learning, and the analysis of dynamic systems. Collaborations in neuroscience, biomedicine, and health motivate these developments and provide settings in which to evaluate their scientific value.

Statistics and AI

At the interface of statistics and AI, I study how statistical principles can inform the construction of learning systems. How should a model account for differences across individuals? Which relationships should a graph representation capture? What information should be shared across different measurement types? These questions connect AI model design to statistical ideas about heterogeneity, dependence, and generalization.

In reinforcement learning, our work develops individualized policies from previously collected data on heterogeneous populations. By representing individual differences through latent variables, we learn decision rules that account for variation in how people respond to actions. This framework combines individualized policy estimation with theoretical guarantees on policy performance.

In graph learning, our MaGNET framework integrates information from local neighborhoods and more distant graph relationships. It identifies influential nodes, edges, and features, connecting predictions to interpretable graph structures. In multimodal learning, our work on variational autoencoders investigates how distinct measurement types can contribute to useful representations. Our recent Meta Fusion framework extends this direction through mutual learning across models, unifying early, intermediate, and late fusion.

Together, these projects investigate how the treatment of variation and structure within a model affects what it can learn and how its results can be interpreted.

Representative work

Bayesian modeling and computation

My Bayesian research develops flexible probability models and the computational methods needed to use them. A central question is how to represent complex relationships in data while keeping inference reliable and computationally feasible. My early work on Dirichlet process mixtures developed nonlinear prediction models that adapt their complexity to the data. Related work uses Gaussian processes and latent-variable models to describe dependence in neural and other biomedical measurements.

A sustained thread develops sampling methods that exploit the structure and geometry of posterior distributions. Split Hamiltonian Monte Carlo separates the Hamiltonian into components so that much of the simulation can be performed at lower computational cost. Spherical Hamiltonian Monte Carlo addresses constrained distributions, while wormhole Hamiltonian Monte Carlo facilitates movement between separated modes. Our work on distributed stochastic-gradient MCMC extends Bayesian computation to large datasets by coordinating sampling across workers.

More recently, I have investigated how neural networks can support Bayesian modeling and inference. This includes methods that approximate expensive computations within sampling workflows, a Calibration–Emulation–Sampling strategy for Bayesian neural networks, and fully Bayesian autoencoders with sparse Gaussian-process priors. These developments extend my broader interest in designing models and algorithms together, so that computational advances support richer statistical inference.

Representative work

Scientific discovery in neuroscience and health

My scientific collaborations investigate how biological systems represent information, change over time, and relate to health outcomes. I contribute to question formulation, study design, method development, and interpretation, connecting statistical advances to the scientific problems that motivate them.

In neuroscience, we study how populations of neurons represent sequences and support memory and decision-making. Our work on hippocampal ensembles showed that neural representations of sequential relationships extend to discrete, nonspatial events. Related projects use latent-factor Gaussian processes to model changing functional connectivity and optimal transport to align neuronal activity across heterogeneous datasets. Current work investigates multi-step planning and goal-directed choice, connecting learned representations to temporal and decision models.

My biomedical collaborations also span Alzheimer’s disease, stroke recovery, circadian biology, and maternal and infant health. Recent projects examine how low-burden clinical measurements can help identify dementia neuropathology and how transcriptomic data can reveal circadian time. Ongoing research on wearable and health data develops methods for individualized risk prediction and intervention. Across these applications, the aim is to connect statistical evidence to specific scientific questions and decisions.

Representative work

Research grants

PI / MPI / Co-PI
T32: Statistical Training to Enhance the Excellence of Research (STEER) in Biomedical Sciences
NIH–NIGMS · Lead PI: Shahbaba · 2025–2030 · $1.9M

Supports eight PhD students per year, joining advanced analytical training with biomedical research.

R01 / SCH: Individualized Learning and Prediction for Heterogeneous Multimodal Data from Wearable Devices
NIH · UCI PI: Shahbaba · 2024–2028 · $1.2M

Machine-learning and reinforcement-learning methods for individualized risk prediction and intervention from wearable and health data.

NCS-FR / DEJA-VU: Design of Joint 3D Solid-State Learning Machines for Various Cognitive Use-Cases
NSF · Co-PI (PI: Fortin) · 2023–2026 · $1.1M

Designing a new class of computer chips by mapping brain-like spatiotemporal signaling onto 3D integrated chips.

PIPE-LINE: Programs for Institutional Pathway Engagement — Accelerating Infrastructure and Education
California Education Learning Lab · UCI PI: Shahbaba · 2023–2027 · $1.3M

Building institutional pathways in data science across UCI, CSUF, and participating community colleges.

R25: Irvine Summer Institute in Biostatistics and Data Science
NIH · MPI: Shahbaba · 2022–2027 · $1.2M

Hands-on research training in biostatistics and data science, with exposure to careers in the field.

HDR DSC: Data Science Training and Practices — Preparing a Diverse Workforce via Academic and Industrial Partnership
NSF · Lead PI: Shahbaba · 2021–2025 · $1.5M

Data science training through curriculum development, mentored research, and academic–industry partnership.

R01: Scalable Bayesian Stochastic Process Models for Neural Data Analysis
NIH · PI: Shahbaba · 2018–2023 · $1.7M

A scalable class of Bayesian stochastic process models and efficient algorithms for multimodal neural data. Project notes on GitHub.

MODULUS: Data-Driven Mechanistic Modeling of Hierarchical Tissues
NSF · PI: Shahbaba · 2019–2022 · $800K

Statistical and mathematical models of how cells and molecules self-organize, with a focus on hematopoiesis.

Theory and Practice for Exploiting the Underlying Structure of Probability Models in Big Data Analysis
NSF · PI: Shahbaba · 2016–2019 · $250K

Combining geometric techniques with computational algorithms to scale up statistical methods for big data. Project notes on GitHub.

Efficient Bayesian Learning from Stochastic Gradients
NSF · Co-PI (PI: Welling) · 2012–2015 · $500K

A family of MCMC procedures requiring only a few data cases per update.

Co-I / Senior Personnel
MathBioSys: NSF-Simons Center for Multiscale Cell Fate
NSF / Simons Foundation · Senior Personnel (PI: Nie) · 2018–2023 · $10M

Investigating how cells differentiate into different cell types.

Bayesian Modeling and Data Integration in Infectious Disease Phylodynamics
NIH · Co-I (PI: Minin) · 2013–2019 · $1.7M

Methodology for population dynamics of infectious disease agents, integrating gene sequencing and surveillance data.

Brain Plasticity and Rehabilitation after Stroke
NIH · Senior Personnel (PI: Cramer) · 2013–2018 · $600K
Transcriptomic, Oxidative Stress, and Inflammatory Responses to Air Pollutants
NIH · Senior Personnel (PI: Delfino) · 2011–2016 · $3.2M
Fetal Programming of the Newborn and Infant Human Brain
NIH · Senior Personnel (PI: Buss) · 2010–2015 · $2.4M
Prenatal Stress Biology, Infant Body Composition and Obesity Risk
NIH · Senior Personnel (PI: Entringer) · 2010–2015 · $1.9M