Nhat Ho

Associate Professor
Department of Statistics and Data Sciences
The University of Texas at Austin
Other affiliations:
Core member, Machine Learning Laboratory
Senior personnel, Institute for Foundations of Machine Learning (IFML)
Email: minhnhat@utexas.edu
Office: WEL 5.242, 105 E 24th Street Austin, TX 78712
Brief Biography
I am an Associate Professor of Statistics and Data Sciences at the University of Texas at Austin, a core member of the Machine Learning Laboratory, and senior personnel of the Institute for Foundations of Machine Learning (IFML). My research develops statistical foundations for modern AI, with the goal of understanding how latent structure, architectural choices, and computational mechanisms determine the statistical efficiency, robustness, interpretability, and scalability of learning systems. A recurring theme of my work is to use statistical and mathematical structure not only to explain modern AI architectures, but also to guide the design of new ones.
Before joining UT Austin, I was a postdoctoral fellow in the Department of Electrical Engineering and Computer Sciences at UC Berkeley, where I was fortunate to be mentored by Professor Michael I. Jordan and Professor Martin J. Wainwright. I received my Ph.D. in Statistics from the University of Michigan in 2017, where I was advised by Professor XuanLong Nguyen and Professor Ya'acov Ritov.
I received the COPSS Emerging Leader Award from the Committee of Presidents of Statistical Societies.
I am also the founder of Trivita AI, a medical AI startup in Vietnam. In collaboration with hospitals and clinical partners, we develop AI systems aimed at reducing physicians' workload, supporting clinical decision making, and improving the accessibility and quality of healthcare. This effort also provides a real-world setting for studying multimodal learning, robustness, adaptation, and decision making in modern AI systems.
Research: Statistical Foundations of Modern AI
Modern AI systems are increasingly built from structured components—latent representations, experts, routers, attention mechanisms, prompts, modalities, and optimization procedures—whose interactions determine how effectively a model can learn from data. My research develops statistical foundations for these systems. I am particularly interested in how latent structure, architectural design, geometry, and computation shape identifiability, statistical efficiency, robustness, specialization, and generalization. A recurring goal is to connect classical ideas from statistics with the mechanisms that drive modern AI, and to use that understanding to suggest new principles for model design.
Latent Structure, Mixtures, and Identifiability
One foundation of my research is the statistical theory of latent-variable models. I study finite and infinite mixtures, hierarchical models, mixture-of-experts, and Bayesian nonparametric models, with particular emphasis on what can be learned about hidden components when the model is non-identifiable, singular, over-specified, weakly separated, or misspecified. Earlier work established how weak identifiability and singularity fundamentally alter parameter-estimation rates in finite mixtures (M.1, M.2) and developed refined geometric tools for mixture estimation (M.3) and Gaussian mixtures of experts (M.4).
More recently, we have been developing a finer theory of the geometry of latent structure. This includes identifying partial-differential-equation barriers to identifiability in infinite mixture models (M.5), deriving convergence rates for latent mixing measures in infinite homoscedastic location-scale mixtures (M.6), characterizing how separation changes the geometry of finite Gaussian mixtures (M.7), and using partial optimal transport to describe heterogeneous component-wise estimation rates (M.8). These problems require metrics and local geometries that reflect how individual latent components can merge, split, or become statistically indistinguishable.
This line of work provides part of the statistical language for my research on modern AI. Experts, modules, attention heads, and learned representations are all latent structures whose usefulness depends on whether they can become distinguishable and specialized from data. Thus, questions that first arise in mixture models—identifiability, separation, singularity, over-specification, and heterogeneous estimation—reappear in a new form when we study modular and conditional-computation architectures.
Modern AI Architectures and Conditional Computation
A major focus of my group is to understand how architectural choices change the statistical difficulty of learning. Mixture-of-experts (MoE) is a central example because routing, sparsity, expert specialization, and parameter sharing can be studied explicitly. Starting from the statistical theory of softmax-gated MoE (A.1), we developed results for top-K sparse routing (A.2), temperature and dense-to-sparse routing (A.3), and least-squares estimation (A.4). This program asks not only whether experts can be estimated, but how the routing mechanism itself changes identifiability, sample complexity, and the emergence of specialization.
We then began comparing architectural mechanisms directly. For example, our work shows that changing the gating rule from softmax to sigmoid can substantially change the statistical behavior of expert estimation (A.5), while alternative quadratic and cosine routers introduce different estimation geometries (A.6, A.7). We also study how contamination, heterogeneity, and expert structure affect minimax estimation (A.8). Together, these results move toward a statistical theory in which routing is not treated as an implementation detail, but as a structural component that can fundamentally alter the learnability of the model.
Our recent work extends these ideas from individual routing mechanisms to modular and shared architectures. We study the statistical benefits of shared experts and normalized sigmoid routing in DeepSeek-style MoE systems (A.9), develop a theory of gated attention through hierarchical mixtures of experts (A.10), and investigate when sigmoid self-attention can improve over softmax self-attention (A.11). These projects are motivated by a broader question: when does decomposition into multiple gates, experts, heads, or shared modules create a genuine statistical advantage rather than merely increasing model capacity?
A related direction studies efficient adaptation and multimodal learning through the same structural lens. We have used mixture and routing ideas to analyze visual prompt experts (A.12), low-rank adaptation through an MoE perspective (A.13), zero-initialized attention (A.14), and reparameterized prompt tuning (A.15). We also connect these ideas to multimodal and continual learning through FuseMoE (A.16) and prompt-based continual learning (A.17). The long-term objective is to develop statistical principles for routing, specialization, sharing, composition, and adaptation that can inform the design of future conditional-computation and modular AI systems.
Geometry, Computation, and Robustness
A second foundation of modern AI is the geometry used to compare complex data and learned representations. My work on optimal transport studies how to retain useful geometric information while making transport-based methods statistically and computationally viable in high dimensions. This includes entropic optimal transport (G.1), projection-robust Wasserstein distance (G.2), and a broad family of sliced-Wasserstein methods (G.3). Recent work develops semi-unbalanced Wasserstein barycenters for Gaussian distributions (G.4), scalable geometric dataset comparison through sliced optimal transport (G.5), and fairness-aware Wasserstein barycenters (G.6).
Geometry alone is not enough: the algorithms used to fit modern models can themselves determine what statistical accuracy is attainable. I therefore study statistical-computational trade-offs in optimization, including the interaction among instability, computation, and estimation error (G.7). This perspective also motivates work on robust decision making, where uncertainty in the data-generating distribution must be modeled rather than ignored. Our work on Bayesian nonparametrics and distributionally robust optimization (G.8) and hierarchical Dirichlet-process formulations of robust optimization (G.9) explores how statistical sharing across related populations can improve robustness.
These geometric, computational, and robustness questions increasingly interact with the architectural questions above. Large AI systems operate on heterogeneous, multimodal, and often shifting data, so the choice of representation, metric, optimization procedure, and robustness criterion can be as consequential as the architecture itself. My broader goal is to understand these ingredients jointly, rather than treating statistical accuracy, computation, and robustness as separate concerns.
Across these three directions, the common objective is to develop a statistical theory that explains when structure helps learning. I am particularly interested in when latent components become identifiable and specialized; when routing, sparsity, sharing, and modularity improve statistical efficiency; how geometry and computation constrain what can be learned; and how these insights can be converted into principles for designing more efficient, reliable, and adaptive AI systems. Current topics of particular interest include mixture-of-experts, modular and compositional AI, attention, multimodal learning, parameter-efficient adaptation, latent-variable models, optimal transport, and robust learning. For a complete list of papers, please see the publication pages organized by type, by topic, or by year.
Editorial Boards of Journals
Area Chairs of Conferences in Machine Learning and Artificial Intelligence
Recent News
[08/2026] I am glad to serve as an Area Chair of the International Conference on Learning Representations (ICLR), 2027 and Senior Program Committee (Area Chair) of AAAI, 2027
[06/2026] I am glad to serve as an Associate Editor of the Transactions on Machine Learning Research (TMLR) .
[4/2026] I am glad to serve as an Area Chair of the Advances in Neural Information Processing Systems (NeurIPS), 2026
[02/2026] I am promoted to Associate Professor with tenure starting from August, 2026.
[02/2026] I am honored to receive the " 2026 COPSS Emerging Leader Award ", a recognition from the Committee of Presidents of Statistical Societies. This award celebrates early-career leaders who are shaping the future of Statistics and Data Science, and I am deeply grateful to the community for this recognition.
[01/2026] Several papers (1, 2, 3, 4, 5, 6, 7, 8, 9) were accepted to Advances in Neural Information Processing Systems (NeurIPS) 2025, International Conference on Computer Vision (ICCV) 2025, International Conference on Artificial Intelligence and Statistics (AISTATS) 2026, and International Conference on Learning Representations (ICLR) 2026.
[01/2026] I am glad to serve as an Area Chair of the International Conference on Machine Learning (ICML), 2026
[01/2026] The paper " Convergence rates for softmax gating mixture of experts " , coauthored with Huy Nguyen and Alessandro Rinaldo was accepted to IEEE Transactions on Information Theory
[08/2025] I am glad to serve as an Area Chair of the International Conference on Artificial Intelligence and Statistics (AISTATS), 2026
[08/2025] I am glad to serve as an Area Chair of the International Conference on Learning Representations (ICLR), 2026 and Senior Program Committee (Area Chair) of AAAI, 2026
[08/2025] The paper " MGPATH: A vision-language model with multi-granular prompt learning for few-shot whole slide pathology classification " , coauthored with Anh-Tien Nguyen, Duy Minh Ho Nguyen, Nghiem Tuong Diep, Trung Quoc Nguyen, Jacqueline Michelle Metsch, Miriam Cindy Maurer, Daniel Sonntag, Hanibal Bohnenberger, and Anne-Christin Hauschild was accepted to Transactions on Machine Learning Research (TMLR)
[05/2025] Four papers (1, 2, 3, 4) were accepted to the International Conference on Machine Learning (ICML), 2025
[4/2025] I am glad to serve as an Area Chair of the Advances in Neural Information Processing Systems (NeurIPS), 2025
[2/2025] The paper " Instability, computational efficiency, and statistical accuracy " , coauthored with Raaz Dwivedi, Koulik Khamaru, Martin J. Wainwright, Michael I. Jordan, and Bin Yu was accepted to Journal of Machine Learning Research (JMLR)
[01/2025] Four papers (1, 2, 3, 4) were accepted to International Conference on Learning Representations (ICLR), 2025
[01/2025] One paper (1) was accepted to International Conference on Artificial Intelligence and Statistics (AISTATS), 2025
[11/2024] I am glad to serve as an Area Chair of the International Conference on Machine Learning (ICML), 2025
[09/2024] Six papers (1, 2, 3, 4, 5, 6) were accepted to Conference on Neural Information Processing Systems (NeurIPS), 2024
[09/2024] I am glad to serve as an Area Chair of the International Conference on Artificial Intelligence and Statistics (AISTATS), 2025
[08/2024] I am glad to serve as an Area Chair of the International Conference on Learning Representations (ICLR), 2025 and Senior Program Committee (Area Chair) of AAAI, 2025
[07/2024] The paper " Global optimality of the EM Algorithm for mixtures of two component linear regressions " , coauthored with Jeongyeol Kwon, Wei Qian, Constantine Caramanis, Yudong Chen, Damek Davis was accepted to IEEE Transactions on Information Theory
[05/2024] Seven papers (1, 2, 3, 4, 5, 6, 7) were accepted to International Conference on Machine Learning (ICML), 2024
[04/2024] The paper " Statistical and computational complexities of BFGS quasi-Newton method for generalized linear models " , coauthored with Qiujiang Jin, Tongzheng Ren, Aryan Mokhtari was accepted to Transactions on Machine Learning Research (TMLR)
[02/2024] The paper " On the computational and statistical complexity of over-parameterized matrix sensing " , coauthored with Jiacheng Zhuo, Jeongyeol Kwon, Constantine Caramanis was accepted to Journal of Machine Learning Research (JMLR)
[02/2024] The paper " On integral theorems and their statistical properties " , coauthored with Stephen Walker was accepted to Proceedings of the Royal Society A: Mathematical, Physical and Engineering Sciences
[02/2024] One paper (1) was accepted to Conference on Computer Vision and Pattern Recognition (CVPR), 2024
[01/2024] I am glad to serve as an Area Chair of the International Conference on Machine Learning (ICML), 2024
[01/2024] Six papers (1, 2, 3, 4, 5, 6) were accepted to International Conference on Learning Representations (ICLR), 2024
[01/2024] Two papers (1, 2) were accepted to International Conference on Artificial Intelligence and Statistics (AISTATS), 2024
[12/2023] The paper " A diffusion process perspective on the posterior contraction rates for parameters ", coauthored with Wenlong Mou, Martin J. Wainwright, Peter L. Bartlett, Michael I. Jordan was accepted to SIAM Journal on Mathematics of Data Science (SIMODS)
[09/2023] Six papers (1, 2, 3, 4, 5, 6) were accepted to Conference on Neural Information Processing Systems (NeurIPS), 2023
[08/2023] I am glad to serve as an Area Chair of the International Conference on Artificial Intelligence and Statistics (AISTATS), 2024
[04/2023] Four papers (1, 2, 3, 4) were accepted to International Conference on Machine Learning (ICML), 2023
[04/2023] The paper " A Bayesian perspective of convolutional neural networks through a deconvolutional generative model " , coauthored with Tan Nguyen, Ankit Patel, Richard Baraniuk, Anima Anandkumar, Michael I. Jordan was accepted to Journal of Machine Learning Research (JMLR)
[01/2023] I am glad to serve as an Associate Editor of the Electronic Journal of Statistics , a top journal in Statistics and Data Science
[01/2023] Three papers (1, 2, 3) are accepted to International Conference on Learning Representations (ICLR) and International Conference on Artificial Intelligence and Statistics (AISTATS)
[09/2022] Six papers (1, 2, 3, 4, 5, 6) were accepted to Conference on Neural Information Processing Systems (NeurIPS), 2022
[09/2022] The paper " Bayesian consistency with the supremum metric", coauthored with Stephen Walker, was accepted to Statistica Sinica
[08/2022] I am glad to serve as an Area Chair of the International Conference on Artificial Intelligence and Statistics (AISTATS), 2023
[08/2022] The paper " Convergence rates for Gaussian mixtures of experts " , coauthored with Chiao-Yu Yang and Michael I. Jordan, was accepted to Journal of Machine Learning Research (JMLR)
[05/2022] Six papers (1, 2, 3, 4, 5, 6) were accepted to International Conference on Machine Learning (ICML), 2022
[05/2022] The paper " On the efficiency of entropic regularized algorithms for optimal transport " ,coauthored with Tianyi Lin and Michael I. Jordan, was accepted to Journal of Machine Learning Research (JMLR), 2022
[02/2022] The paper " On the complexity of approximating multi-marginal optimal transport" ,coauthored with Tianyi Lin, Marco Cuturi, and Michael I. Jordan, was accepted to Journal of Machine Learning Research (JMLR), 2022
[01/2022] Four papers (1, 2, 3, 4) were accepted to International Conference on Artificial Intelligence and Statistics (AISTATS), 2022
Students' News
[02/2026] Congratulations to Huy Nguyen for winning the UT Outstanding Graduate Research Fellowship from the Graduate School for 2026-2027!
[02/2026] Congratulations to Khai Nguyen who just accepted the tenure-track Assistant Professor position in Department of Statistics, Texas A&M!
[02/2025] Congratulations to Khai Nguyen for winning the UT Outstanding Graduate Research Fellowship from the Graduate School for 2025-2026!
Selected Publications on Theory
(* = equal contribution )
(** = alphabetical order )
(† = co-last author )
Partial differential equation barriers to identifiability in infinite mixture models . Under review.
Dung Le, Nicola Bariletto, Alessandro Rinaldo, Nhat Ho.
Convergence rates for latent mixing measures in infinite homoscedastic location-scale mixture models . Under review.
Nicola Bariletto*, Dung Le*, Alessandro Rinaldo, Nhat Ho.
On the geometry of separation in finite Gaussian mixtures . Under review.
Huy Nguyen*, Dung Le*, Alessandro Rinaldo, Nhat Ho.
Characterizing heterogeneous rates in finite mixture estimation via partial optimal transport . Under review.
Dung Le*, Huy Nguyen*, Trang Pham, Alessandro Rinaldo, Nhat Ho.
Instability, computational efficiency, and statistical accuracy . Journal of Machine Learning Research (JMLR), 2025.
Nhat Ho*, Raaz Dwivedi*, Koulik Khamaru*, Martin J. Wainwright, Michael I. Jordan, Bin Yu.
On DeepSeekMoE: Statistical benefits of shared experts and normalized sigmoid gating . Under review.
Huy Nguyen, Thong Doan, Quang Pham, Nghi Bui, Nhat Ho†, Alessandro Rinaldo†.
Demystifying softmax gating in Gaussian mixture of experts . Advances in NeurIPS, 2023 (Spotlight).
Huy Nguyen, Tin Nguyen, Nhat Ho.
A statistical theory of gated attention through the lens of hierarchical mixture of experts . Under review.
Nguyen Hoang Viet, Tuan Minh Pham, Thinh Cao Huy, Tan Dinh, Huy Nguyen, Nhat Ho†, Alessandro Rinaldo†.
Sigmoid self-attention is better than softmax self-attention: A mixture-of-experts perspective . Under review.
Fanqi Yan*, Huy Nguyen*, Pedram Akbarian, Nhat Ho†, Alessandro Rinaldo†.
Is temperature sample efficient for softmax Gaussian mixture of experts? . Proceedings of the ICML, 2024.
Huy Nguyen, Pedram Akbarian, Nhat Ho.
Bayesian nonparametrics meets data-driven robust optimization . Advances in NeurIPS, 2024.
Nicola Bariletto, Nhat Ho.
Sigmoid gating is more sample efficient than softmax gating in mixture of experts . Advances in NeurIPS, 2024.
Huy Nguyen, Nhat Ho†, Alessandro Rinaldo†.
On expert estimation in hierarchical mixture of experts: Beyond softmax gating functions . Under review.
Huy Nguyen*, Xing Han*, Carl Harris, Suchi Saria, Nhat Ho.
Quadratic gating functions in mixture of experts: A statistical insight . Under review.
Pedram Akbarian*, Huy Nguyen*, Xing Han*, Nhat Ho.
Revisiting prefix-tuning: Statistical benefits of reparameterization among prompts . ICLR, 2025.
Minh Le*, Chau Nguyen*, Huy Nguyen*, Quyen Tran, Trung Le, Nhat Ho.
Borrowing strength in distributionally robust optimization via hierarchical Dirichlet processes . Under review.
Nicola Bariletto, Khai Nguyen, Nhat Ho.
Statistical advantages of perturbing cosine router in sparse mixture of experts . ICLR, 2025.
Huy Nguyen, Pedram Akbarian*, Trang Pham*, Trang Nguyen*, Shujian Zhang Nhat Ho.
On least square estimation in softmax gating mixture of experts . Proceedings of the ICML, 2024.
Huy Nguyen, Nhat Ho†, Alessandro Rinaldo†.
Statistical perspective of top-K sparse softmax gating mixture of experts . ICLR, 2024.
Huy Nguyen, Pedram Akbarian, Fanqi Yan, Nhat Ho.
Multivariate smoothing via the Fourier integral theorem and Fourier kernel . Under revision.
Nhat Ho**, Stephen G. Walker**.
A diffusion process perspective on the posterior contraction rates for parameters. SIAM Journal on Mathematics of Data Science (SIMODS), 2023.
Wenlong Mou, Nhat Ho, Martin J. Wainwright, Peter L. Bartlett, Michael I. Jordan.
Minimax optimal rate for parameter estimation in multivariate deviated models. Advances in NeurIPS, 2023.
Dat Do*, Huy Nguyen*, Khai Nguyen, Nhat Ho.
Neural collapse in deep linear network: from balanced to imbalanced data . Proceedings of the ICML, 2023.
Hien Dang*, Tho Tran*, Hung Tran, Tan Nguyen†, Nhat Ho†.
Bayesian consistency with the supremum metric . Statistica Sinica, 2022.
Nhat Ho**, Stephen G. Walker**.
An exponentially increasing step-size for parameter estimation in statistical models. Under review.
Nhat Ho**, Tongzheng Ren**, Purnamrita Sarkar**, Sujay Sanghavi**, Rachel Ward**.
On integral theorems and their statistical properties . Proceedings of the Royal Society A: Mathematical, Physical and Engineering Sciences, 2024.
Nhat Ho**, Stephen G. Walker**.
On excess mass behavior in Gaussian mixture models with Orlicz Wasserstein distances . Proceedings of the ICML, 2023.
Aritra Guha, Nhat Ho, Long Nguyen.
Beyond black box densities: Parameter learning for the deviated components. Advances in NeurIPS, 2022.
Dat Do*, Nhat Ho*, Long Nguyen.
Refined convergence rates for maximum likelihood estimation under finite mixture models. Proceedings of the ICML, 2022 (Long Presentation).
Tudor Manole, Nhat Ho.
Entropic Gromov-Wasserstein between Gaussian distributions. Proceedings of the ICML, 2022.
Khang Le*, Dung Le*, Huy Nguyen*, Dat Do, Tung Pham, Nhat Ho.
Towards statistical and computational complexities of Polyak step size gradient descent . AISTATS, 2022.
Tongzheng Ren*, Fuheng Cui*, Alexia Atsidakou*, Sujay Sanghavi, Nhat Ho.
On the minimax optimality of the EM algorithm for learning two-component mixed linear regression. AISTATS, 2021.
Jeong Y. Kwon, Nhat Ho, Constantine Caramanis.
Beyond EM algorithm on over-specified two-component location-scale Gaussian mixtures. Under review.
Tongzheng Ren*, Fuheng Cui*, Sujay Sanghavi, Nhat Ho.
On the efficiency of entropic regularized algorithms for optimal transport . Journal of Machine Learning Research (JMLR), 2022.
Tianyi Lin, Nhat Ho, Michael I. Jordan.
Convergence rates for Gaussian mixtures of experts. Journal of Machine Learning Research (JMLR), 2022.
Nhat Ho, Chiao-Yu Yang, Michael I. Jordan.
On the computational and statistical complexity of over-parameterized matrix sensing . Journal of Machine Learning Research (JMLR), 2024.
Jiacheng Zhuo, Jeongyeol Kwon, Nhat Ho, Constantine Caramanis.
Projection robust Wasserstein distance and Riemannian optimization . Advances in NeurIPS, 2020 (Spotlight).
Tianyi Lin*, Chenyou Fan*, Nhat Ho, Marco Cuturi, Michael I. Jordan.
Improving computational complexity in statistical models with second-order information . Proceedings of the ICML, 2024.
Tongzheng Ren, Jiacheng Zhuo, Sujay Sanghavi, Nhat Ho.
On posterior contraction of parameters and interpretability in Bayesian mixture modeling . Bernoulli 27 (4), 2159-2188, 2021.
Aritra Guha, Nhat Ho, XuanLong Nguyen.
Singularity, misspecification, and the convergence rate of EM. Annals of Statistics, 48(6), 3161-3182, 2020.
Raaz Dwivedi*, Nhat Ho*, Koulik Khamaru*, Martin J. Wainwright, Michael I. Jordan, Bin Yu.
Singularity structures and impacts on parameter estimation behavior in finite mixtures of distributions. SIAM Journal on Mathematics of Data Science (SIMODS), 1(4), 730–758, 2019.
Nhat Ho and XuanLong Nguyen.
Convergence rates of parameter estimation for some weakly identifiable finite mixtures. Annals of Statistics, 44(6), 2726-2755, 2016.
Nhat Ho and XuanLong Nguyen
Selected Publications on Method and Application
A Bayesian perspective of convolutional neural networks through a deconvolutional generative model. Journal of Machine Learning Research (JMLR), 2023.
Nhat Ho*, Tan Nguyen*, Ankit Patel, Anima Anandkumar, Michael I. Jordan, Richard Baraniuk.
Quasi-Monte Carlo for 3D sliced Wasserstein . ICLR, 2024 (Spotlight).
Khai Nguyen, Nicola Bariletto, Nhat Ho.
Sliced Wasserstein estimation with control variates . ICLR, 2024.
Khai Nguyen, Nhat Ho.
FuseMoE: Mixture-of-experts Transformers for fleximodal fusion . Advances in NeurIPS, 2024.
Xing Han, Huy Nguyen, Carl William Harris, Nhat Ho†, Suchi Saria†.
Hierarchical hybrid sliced Wasserstein: A scalable metric for heterogeneous joint distributions . Advances in NeurIPS, 2024.
Khai Nguyen, Nhat Ho.
Lightspeed geometric dataset distance via sliced optimal transport. Proceedings of the ICML, 2025.
Khai Nguyen*, Hai Nguyen*, Tuan Pham, Nhat Ho.
CompeteSMoE - Effective training of sparse mixture of experts via competition . Under review.
Quang Pham, Truong Giang Do, Huy Nguyen, TrungTin Nguyen, Chenghao Liu, Mina Sartipi, Binh T. Nguyen, Savitha Ramasamy, Xiaoli Li, Steven Hoi†, Nhat Ho†.
Mixture of experts meets prompt-based continual learning . Advances in NeurIPS, 2024.
Minh Le, An Nguyen*, Huy Nguyen*, Trang Nguyen*, Trang Pham*, Linh Van Ngo, Nhat Ho.
Marginal fairness sliced Wasserstein barycenter . ICLR, 2025 (Spotlight).
Khai Nguyen, Hai Nguyen, Nhat Ho.
LoGra-Med: Long context multi-graph alignment for medical vision-language model . Under review.
Duy M. H. Nguyen, Nghiem T. Diep, Trung Q. Nguyen, Hoang-Bao Le, Tai Nguyen, Tien Nguyen, TrungTin Nguyen, Nhat Ho, Pengtao Xie, Roger Wattenhofer, James Zhou, Daniel Sonntag, Mathias Niepert.
Energy-based sliced Wasserstein distance . Advances in NeurIPS, 2023.
Khai Nguyen, Nhat Ho.
FourierFormer: Transformer meets generalized Fourier integral attentions . Advances in NeurIPS, 2022.
Tan Nguyen*, Minh Pham*, Tam Nguyen, Khai Nguyen, Stanley Osher, Nhat Ho.
Improving Transformers with probabilistic attention keys . ICML, 2022.
Tam Nguyen*, Tan Nguyen*, Dung Le, Khuong Nguyen, Anh Tran, Richard Baraniuk, Nhat Ho†, Stanley Osher†.
Amortized projection optimization for sliced Wasserstein generative models. Advances in NeurIPS, 2022.
Khai Nguyen, Nhat Ho.
Hierarchical sliced Wasserstein distance. ICLR, 2023.
Khai Nguyen, Tongzheng Ren, Huy Nguyen, Litu Rout, Tan Nguyen, Nhat Ho.
Designing robust transformers using robust kernel density estimation. Advances in NeurIPS, 2023.
Xing Han, Tongzheng Ren, Tan Minh Nguyen, Khai Nguyen, Joydeep Ghosh, Nhat Ho.
A primal-dual framework for transformers and neural networks. ICLR, 2023 (Spotlight).
Tan Minh Nguyen, Tam Minh Nguyen, Nhat Ho, Andrea L. Bertozzi, Richard Baraniuk, Stanley Osher.
Improving Transformer with an admixture of attention heads . Advances in NeurIPS, 2022.
Tam Nguyen*, Tan Nguyen*, Hai Do, Khai Nguyen, Vishwanath Saragadam, Minh Pham, Khuong Nguyen,Stanley Osher†, Nhat Ho†.
Generative models from the multivariate Fourier integral theorem . Under review.
Nhat Ho**, Stephen G. Walker**.
Probabilistic best subset selection via gradient-based optimization. Under review.
Mingzhang Yin, Nhat Ho, Bowei Yan, Xiaoning Qian, Mingyuan Zhou.
Codes
The official Github link for codes of research papers from our Data Science and Machine Learning (DSML) Lab is: https://github.com/UT-Austin-Data-Science-Group.
Media Coverage
Data Science, Machine Learning, Statistics, and Artifical Intelligence have become very important fields in Vietnam these days. However, as these fields are still very young in Vietnam, young Vietnamese generation often faces challenges to equip themselves with enough information, knowledge, and skills to pursue their career paths in these fields. For this reason, several leading newspapers and shows in Vietnam covered my path and story of becoming a professor in the leading US university as well as my opinion about these fields to inspire and provide necessary information to young generation in Vietnam that would like to pursue their careers in Data Science, Machine Learning, Statistics, and Artifical Intelligence, including:
Thanh Niên Newspaper (in Vietnamese): Giáo sư 34 tuổi giúp người trẻ học tiến sĩ tại các trường hàng đầu thế giới
Thanh Niên Newspaper (in Vietnamese): Nhà khoa học với 100 công trình AI: 'Muốn Việt Nam là điểm sáng của thế giới'
Vietnamnet Newspaper (in Vietnamese): Giáo sư 36 tuổi người Việt đoạt giải thưởng quốc tế danh giá về thống kê và AI
VnExpress Newspaper (in Vietnamese): Giáo sư Việt nhận giải thưởng danh giá ngành Thống kê
CafeF/ VTCNews Newspaper (in Vietnamese): Giáo sư Việt đầu tiên nhận giải thưởng danh giá ngành thống kê, AI toàn cầu
VTCNews Newspaper (in Vietnamese): Người Việt đứng sau 110 công trình quốc tế giải mã 'hộp đen' AI
VTCNews Newspaper (in Vietnamese): Những học giả, trí thức Việt tạo dấu ấn trên thế giới năm qua
Diễn Đàn Doanh Nghiệp Newspaper (in Vietnamese): Tâm huyết kiến tạo thế hệ AI Việt
VnExpress Newspaper (in Vietnamese): Cơn sốt du học ngành AI, bán dẫn
Công Thương Newspaper (in Vietnamese): Xây dựng hệ sinh thái AI mang dấu ấn Việt Nam
Thanh Niên Newspaper (in Vietnamese): Ngành nghề của tương lai: Nhiều cơ hội việc làm về khoa học dữ liệu
Sài Gòn Giải Phóng Newspaper (in Vietnamese): Báo Xuân Giáp Thìn "Những Cánh Thiên Di của Trí Tuệ Việt"; Outstanding Vietnamese conquering pinnacle of knowledge overseas
Thanh Niên Newspaper (in Vietnamese): Việc làm lĩnh vực trí tuệ nhân tạo: Cơ hội việc làm rộng mở
Tuổi Trẻ Newspaper (in Vietnamese): Muốn trở thành kỹ sư trí tuệ nhân tạo cần học tốt những môn nào?
Thanh Niên Newspaper (in Vietnamese): Trang bị kỹ năng sử dụng AI để có thêm nhiều cơ hội trong cuộc sống; Muốn sử dụng AI, bắt đầu từ đâu?; Muốn sử dụng AI, bắt đầu từ đâu?: Những người trẻ không... lỗi nhịp; Muốn sử dụng AI, bắt đầu từ đâu?: Nhờ AI, khởi nghiệp đột phá
Diễn Đàn Doanh Nghiệp Newspaper (in Vietnamese): AI, dữ liệu và tương lai chuỗi cung ứng Việt
Diễn Đàn Doanh Nghiệp Newspaper (in Vietnamese): Những yếu tố định hình tương lai nguồn nhân lực
Tiền Phong Newspaper (in Vietnamese): Chàng trai Bạc Liêu chia sẻ hành trình trở thành giáo sư tại Mỹ về AI và Khoa học công nghệ
VnExpress Newspaper (in Vietnamese): Từ học sinh tỉnh lẻ thành giáo sư đại học Mỹ
Thanh Niên Newspaper (in Vietnamese): Người trẻ thiếu trầm trọng kỹ năng sống; Người trẻ thiếu trầm trọng kỹ năng sống: Nguy cơ trở thành tội phạm; Người trẻ thiếu trầm trọng kỹ năng sống: 5 cách để tự lấp đầy 'lỗ hổng'
Sài Gòn Giải Phóng Newspaper (in Vietnamese): Đi vào vùng chưa từng khám phá
Giáo Dục Việt Nam Newspaper (in Vietnamese): Bài học kinh nghiệm nhìn từ cách huy động nguồn lực tài chính ở một số ĐH của Mỹ
VTV1 Television Show "Cất Cánh" (in Vietnamese): Thời khắc để bắt đầu
VTV1 Television Show "Cất Cánh" (in Vietnamese): Tiếng gọi quê hương
VTV1 Television Show "Toàn Cảnh Báo Xuân" (in Vietnamese): Tự hào Trí Tuệ Việt