I am an Assistant Professor in Statistics in the Donald Bren School of Information and Computer Sciences at University of California Irvine. I obtained my Ph.D. degree in Statistics at North Carolina State University (NCSU), co-advised by Dr. Wenbin Lu and Dr. Rui Song. Prior to that, I obtained a B.S. in Statistics from Zhejiang University in July 2017.
I have broad research interests in methodology and theory in reinforcement learning, causal inference, natural language processing, and their interchanges, to establish reliable, powerful, and interpretable solutions to wide real-world problems. Currently, my main research work includes individualized optimal decision making with complex data, causal reasoning for explainable artificial intelligence, large language model (LLM) alignment, policy evaluation in reinforcement learning (RL) and bandits.
Click here for my name in Chinese and how to pronounce it.
Contact me: hengrc1@uci.edu
PhD in Statistics, 2022
North Carolina State University
B.S. in Statistics, 2017
Zhejiang University
June 2026: Congratulations to my student Wenbo Zhang on earning his Ph.D.! Wenbo will join Alibaba’s Post-Training LLM Research Group, and was also a recipient of the UCI Statistics Fellowship Award for Methodology Research!
April 2026: Our paper Reasoning Is Not Free: Robust Adaptive Cost-Efficient Router for LLM-as-a-Judge is accepted at ICML 2026.
April 2026: Our paper Position: Prompting Intent Should Be Audited in LLM-Assisted Peer Review is accepted at ICML 2026.
April 2026: Our paper Beyond the Singular: Revealing the Value of Multiple Generations in Benchmark Evaluation is accepted at ACL 2026 Findings.
March 2026: I will serve as Area Chair for The Fortieth Annual Conference on Neural Information Processing Systems (NeurIPS 2026).
Jan 2026: Congratulations to my student Lijinghua (Lizzie) Zhang who won the ASA Student Paper Award in the Section on Text Analysis!
Dec 2025: Our work is supported by Thinking Machines Lab. Thanks Thinking Machines Lab!
Dec 2025: Our paper Sequential Knockoffs for Variable Selection in Reinforcement Learning is accepted at Journal of the American Statistical Association.
Dec 2025: Our paper A Review of Causal Decision Making is accepted at Journal of Artificial Intelligence Research.
Nov 2025: I will serve as Area Chair for International Conference on Machine Learning (ICML 2026).
Oct 2025: I will serve as Area Chair for ACL Rolling Review (ARR).
Aug 2025: I will serve as Area Chair for International Conference on Learning Representations (ICLR 2026).
Aug 2025: Our paper Recognizing Limits: Investigating Infeasibility in Large Language Models is accepted at EMNLP 2025.
June 2025: Our paper Where to Intervene: Action Selection in Deep Reinforcement Learning is accepted at Transactions on Machine Learning Research.
June 2025: Congratulations to my student Liner Xiang on passing the Ph.D. Advancement to Candidacy with the title Decisions in the Wild: Foresight, Adversaries, and Human Feedback in Online Policy Optimization and Evaluation !
(^ corresponding author, ___ graduate student author, * co-first author)
Zhang, W., Zhou, W., Cai, H.^, & Qi, Z. (2026). Towards Bridging the Gap Between Offline and Iterative Alignment via Preference Distillation. Transactions on Machine Learning Research.
Zhang, W.*, Zhang, L.*, Xiang, L.*, & Cai, H.^ (2026). Reasoning Is Not Free: Robust Adaptive Cost-Efficient Router for LLM-as-a-Judge. In International Conference on Machine Learning (ICML 2026).
Zhang, L.*, Bang, M.*, & Cai, H.^ (2026). Position: Prompting Intent Should Be Audited in LLM-Assisted Peer Review. In International Conference on Machine Learning (ICML 2026).
Zhang, W., Cai, H.^, & Chen, W. (2026). Beyond the Singular: Revealing the Value of Multiple Generations in Benchmark Evaluation. The 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026 Findings).
Ma, T., Zhu, J., Cai, H., Qi, Z., Chen, Y., Shi, C., & Laber, E. B. (2026). Sequential Knockoffs for Variable Selection in Reinforcement Learning. Journal of the American Statistical Association.
Ge, L. *, Cai, H. *, Wan, R. *, Xu, Y. *, & Song, R. (2026). A Review of Causal Decision Making. Journal of Artificial Intelligence Research.
Zhang, W., Xu, Z., & Cai, H.^ (2025). Recognizing Limits: Investigating Infeasibility in Large Language Models. Empirical Methods in Natural Language Processing (EMNLP 2025 Findings).
Zhang, W., & Cai, H.^ (2025) . Where to Intervene: Action Selection in Deep Reinforcement Learning Transactions on Machine Learning Research.
Cai, H., Wang, Y., Jordan, M., & Song, R. (2024). On Learning Necessary and Sufficient Causal Graphs. Advances in Neural Information Processing Systems (NeurIPS).
Shen, Y. *, Cai, H. *, & Song, R. (2024). Doubly Robust Interval Estimation for Optimal Policy Evaluation in Online Learning. Journal of the American Statistical Association.
Zhang, W., Wu, T., Wang, Y., Cai, Y., & Cai, H.^ (2023) Towards Trustworthy Explanation: On Causal Rationalization. In International Conference on Machine Learning (ICML 2023).
Watson, RA. *, Cai, H. *, An, X., McLean, S., & Song, R. (2023). On Heterogeneous Treatment Effects in Heterogeneous Causal Graphs. In International Conference on Machine Learning (ICML 2023).
Cai, H. *, Shi, C. *, Song, R., & Lu, W. (2023). Jump Q-Learning for Individualized Decision Making with Continuous Treatments. Journal of Machine Learning Research.
Cai, H.^, Lu, W., Marceau West R., Mehrotra DV., & Huang, L. (2022). CAPITAL: Optimal Subgroup Identification via Constrained Policy Tree Search. Statistics in Medicine.
Cai, H. *, Shi, C *., Song, R., & Lu, W. (2021). Deep Jump Learning for Off-Policy Evaluation in Continuous Treatment Settings. Advances in Neural Information Processing Systems (NeurIPS).
Cai, H., Song, R., & Lu, W. (2021). ANOCE: Analysis of Causal Effects with Multiple Mediators via Constrained Structural Learning. International Conference on Learning Representations (ICLR).
Cai, H., Lu, W., & Song, R. (2020). On Validation and Planning of An Optimal Decision Rule with Application in Healthcare Studies. International Conference on Machine Learning (ICML).
Xiang, L., Wang, Y., & Cai, H.^ (2026+). Policy Optimization and Statistical Inference for Online Contextual Matrix Games.. arXiv preprint arXiv:2608.17173.
Zhang, L.*, Xiang, L.*, Zhang, W.*, Zheng, Y., & Cai, H.^ (2026+). Large Language Model Alignment with Complex Feedback: A Survey. Preprints: 202608.0674.
Zhang, L., & Cai, H.^ (2026+). Text Rationalization for Robust Causal Effect Estimation.. arXiv preprint arXiv:2512.05373.
Xiang, L., Wang, J., & Cai, H.^ (2026+). Foresighted Online Policy Optimization with Interference.. arXiv preprint arXiv:2510.15273.
Cai, H. *, Jin, H. *, Li, L. (2025+). Conformal Diffusion Models for Individual Treatment Effect Estimation and Inference. arXiv preprint arXiv:2408.01582.
Cai, H. *, Liu, S. *, & Song, R. (2025+). Enhancing Causal Reasoning in Large Language Models: A Causal Attribution Model for Precision Fine-Tuning. arXiv preprint arXiv:2401.00139.
Current Ph.D. Students
Graduated Ph.D. Students