Nhat Ho

Nhat Ho

Associate Professor
Department of Statistics and Data Sciences
The University of Texas at Austin

Other affiliations:
Core member, Machine Learning Laboratory
Senior personnel, Institute for Foundations of Machine Learning (IFML)

Email: minhnhat@utexas.edu
Office: WEL 5.242, 105 E 24th Street Austin, TX 78712

Brief Biography


I am an Associate Professor of Statistics and Data Sciences at the University of Texas at Austin, a core member of the Machine Learning Laboratory, and senior personnel of the Institute for Foundations of Machine Learning (IFML). My research develops statistical foundations for modern AI, with the goal of understanding how latent structure, architectural choices, and computational mechanisms determine the statistical efficiency, robustness, interpretability, and scalability of learning systems. A recurring theme of my work is to use statistical and mathematical structure not only to explain modern AI architectures, but also to guide the design of new ones.

Before joining UT Austin, I was a postdoctoral fellow in the Department of Electrical Engineering and Computer Sciences at UC Berkeley, where I was fortunate to be mentored by Professor Michael I. Jordan and Professor Martin J. Wainwright. I received my Ph.D. in Statistics from the University of Michigan in 2017, where I was advised by Professor XuanLong Nguyen and Professor Ya'acov Ritov.

I received the COPSS Emerging Leader Award from the Committee of Presidents of Statistical Societies.

I am also the founder of Trivita AI, a medical AI startup in Vietnam. In collaboration with hospitals and clinical partners, we develop AI systems aimed at reducing physicians' workload, supporting clinical decision making, and improving the accessibility and quality of healthcare. This effort also provides a real-world setting for studying multimodal learning, robustness, adaptation, and decision making in modern AI systems.

Research: Statistical Foundations of Modern AI


Modern AI systems are increasingly built from structured components—latent representations, experts, routers, attention mechanisms, prompts, modalities, and optimization procedures—whose interactions determine how effectively a model can learn from data. My research develops statistical foundations for these systems. I am particularly interested in how latent structure, architectural design, geometry, and computation shape identifiability, statistical efficiency, robustness, specialization, and generalization. A recurring goal is to connect classical ideas from statistics with the mechanisms that drive modern AI, and to use that understanding to suggest new principles for model design.

Latent Structure, Mixtures, and Identifiability

One foundation of my research is the statistical theory of latent-variable models. I study finite and infinite mixtures, hierarchical models, mixture-of-experts, and Bayesian nonparametric models, with particular emphasis on what can be learned about hidden components when the model is non-identifiable, singular, over-specified, weakly separated, or misspecified. Earlier work established how weak identifiability and singularity fundamentally alter parameter-estimation rates in finite mixtures (M.1, M.2) and developed refined geometric tools for mixture estimation (M.3) and Gaussian mixtures of experts (M.4).

More recently, we have been developing a finer theory of the geometry of latent structure. This includes identifying partial-differential-equation barriers to identifiability in infinite mixture models (M.5), deriving convergence rates for latent mixing measures in infinite homoscedastic location-scale mixtures (M.6), characterizing how separation changes the geometry of finite Gaussian mixtures (M.7), and using partial optimal transport to describe heterogeneous component-wise estimation rates (M.8). These problems require metrics and local geometries that reflect how individual latent components can merge, split, or become statistically indistinguishable.

This line of work provides part of the statistical language for my research on modern AI. Experts, modules, attention heads, and learned representations are all latent structures whose usefulness depends on whether they can become distinguishable and specialized from data. Thus, questions that first arise in mixture models—identifiability, separation, singularity, over-specification, and heterogeneous estimation—reappear in a new form when we study modular and conditional-computation architectures.

Modern AI Architectures and Conditional Computation

A major focus of my group is to understand how architectural choices change the statistical difficulty of learning. Mixture-of-experts (MoE) is a central example because routing, sparsity, expert specialization, and parameter sharing can be studied explicitly. Starting from the statistical theory of softmax-gated MoE (A.1), we developed results for top-K sparse routing (A.2), temperature and dense-to-sparse routing (A.3), and least-squares estimation (A.4). This program asks not only whether experts can be estimated, but how the routing mechanism itself changes identifiability, sample complexity, and the emergence of specialization.

We then began comparing architectural mechanisms directly. For example, our work shows that changing the gating rule from softmax to sigmoid can substantially change the statistical behavior of expert estimation (A.5), while alternative quadratic and cosine routers introduce different estimation geometries (A.6, A.7). We also study how contamination, heterogeneity, and expert structure affect minimax estimation (A.8). Together, these results move toward a statistical theory in which routing is not treated as an implementation detail, but as a structural component that can fundamentally alter the learnability of the model.

Our recent work extends these ideas from individual routing mechanisms to modular and shared architectures. We study the statistical benefits of shared experts and normalized sigmoid routing in DeepSeek-style MoE systems (A.9), develop a theory of gated attention through hierarchical mixtures of experts (A.10), and investigate when sigmoid self-attention can improve over softmax self-attention (A.11). These projects are motivated by a broader question: when does decomposition into multiple gates, experts, heads, or shared modules create a genuine statistical advantage rather than merely increasing model capacity?

A related direction studies efficient adaptation and multimodal learning through the same structural lens. We have used mixture and routing ideas to analyze visual prompt experts (A.12), low-rank adaptation through an MoE perspective (A.13), zero-initialized attention (A.14), and reparameterized prompt tuning (A.15). We also connect these ideas to multimodal and continual learning through FuseMoE (A.16) and prompt-based continual learning (A.17). The long-term objective is to develop statistical principles for routing, specialization, sharing, composition, and adaptation that can inform the design of future conditional-computation and modular AI systems.

Geometry, Computation, and Robustness

A second foundation of modern AI is the geometry used to compare complex data and learned representations. My work on optimal transport studies how to retain useful geometric information while making transport-based methods statistically and computationally viable in high dimensions. This includes entropic optimal transport (G.1), projection-robust Wasserstein distance (G.2), and a broad family of sliced-Wasserstein methods (G.3). Recent work develops semi-unbalanced Wasserstein barycenters for Gaussian distributions (G.4), scalable geometric dataset comparison through sliced optimal transport (G.5), and fairness-aware Wasserstein barycenters (G.6).

Geometry alone is not enough: the algorithms used to fit modern models can themselves determine what statistical accuracy is attainable. I therefore study statistical-computational trade-offs in optimization, including the interaction among instability, computation, and estimation error (G.7). This perspective also motivates work on robust decision making, where uncertainty in the data-generating distribution must be modeled rather than ignored. Our work on Bayesian nonparametrics and distributionally robust optimization (G.8) and hierarchical Dirichlet-process formulations of robust optimization (G.9) explores how statistical sharing across related populations can improve robustness.

These geometric, computational, and robustness questions increasingly interact with the architectural questions above. Large AI systems operate on heterogeneous, multimodal, and often shifting data, so the choice of representation, metric, optimization procedure, and robustness criterion can be as consequential as the architecture itself. My broader goal is to understand these ingredients jointly, rather than treating statistical accuracy, computation, and robustness as separate concerns.

Across these three directions, the common objective is to develop a statistical theory that explains when structure helps learning. I am particularly interested in when latent components become identifiable and specialized; when routing, sparsity, sharing, and modularity improve statistical efficiency; how geometry and computation constrain what can be learned; and how these insights can be converted into principles for designing more efficient, reliable, and adaptive AI systems. Current topics of particular interest include mixture-of-experts, modular and compositional AI, attention, multimodal learning, parameter-efficient adaptation, latent-variable models, optimal transport, and robust learning. For a complete list of papers, please see the publication pages organized by type, by topic, or by year.

Editorial Boards of Journals


Area Chairs of Conferences in Machine Learning and Artificial Intelligence


Recent News


Students' News


Selected Publications on Theory

(* = equal contribution )
(** = alphabetical order )
(† = co-last author )


Selected Publications on Method and Application


Codes


The official Github link for codes of research papers from our Data Science and Machine Learning (DSML) Lab is: https://github.com/UT-Austin-Data-Science-Group.

Media Coverage


Data Science, Machine Learning, Statistics, and Artifical Intelligence have become very important fields in Vietnam these days. However, as these fields are still very young in Vietnam, young Vietnamese generation often faces challenges to equip themselves with enough information, knowledge, and skills to pursue their career paths in these fields. For this reason, several leading newspapers and shows in Vietnam covered my path and story of becoming a professor in the leading US university as well as my opinion about these fields to inspire and provide necessary information to young generation in Vietnam that would like to pursue their careers in Data Science, Machine Learning, Statistics, and Artifical Intelligence, including: