Hi! I’m Kaixuan Ji, a fourth-year Ph.D. student in Computer Science at UCLA, fortunately advised by Professor Quanquan Gu. Before coming to UCLA, I completed my undergraduate studies in Department of Computer Science and Technology at Tsinghua University, where I had the privilege to work with Professors Jie Tang and Juanzi Li. My research studies the algorithms and fundamental limits of sequential decision-making, including bandits and reinforcement learning, and applies them to the practice of large language models. Here is my latest CV.
[09/2026] Our work on the optimal sample complexity of learning an offline multi-armed bandit with KL-regularized objective is accepted by NeurIPS 2026.
[09/2026] Our work showing that a faster rate of learning offline forward-KL regularized contextual bandits is achievable under single-policy concentrability is accepted by NeurIPS 2026 and selected as a Spotlight.
On the Optimal Sample Complexity of Offline Multi-Armed Bandits with KL Regularization
Kaixuan Ji*, Qiwei Di*, Heyang Zhao, Qingyue Zhao, Quanquan Gu, NeurIPS 2026
Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability
Qingyue Zhao*, Kaixuan Ji*, Heyang Zhao, Quanquan Gu, NeurIPS 2026, Spotlight