Zhuohan Wang

M.S. Student, Harvard University
Ex-MERL, Ex-SenseTime

Email: zhuohan_wang [at] g.harvard.edu

Photo of Zhuohan Wang

I am a master's student in Computational Science and Engineering at Harvard University. I received my B.Eng. in Computer and Data Engineering from City University of Hong Kong, ranked first in my class and graduating summa cum laude, and worked with Prof. Lai-Man Po on parameter-efficient fine-tuning. I spent a year as an intern on the LLM post-training team at SenseTime Research, mentored by Dr. Litong Feng. I also worked with Dr. Mengyue Yang on causal reasoning in LLMs. In 2026 I was a research intern at MERL with Dr. Pu (Perry) Wang and Dr. Anoop Cherian.

My long-term goal is to build systems that can think at least as well as the best human researchers, across intellectual domains, in a way that keeps improving with scale. I am interested in automated discovery and in reasoning in language, vision and other modalities. I also work on efficient training and decoding.

I am also exploring a startup that applies AI to a real e-commerce business. Outside research, I enjoy investing and playing badminton. I am always happy to discuss shared interests, so feel free to reach out.

News

Selected Publications

Reasoning

  • Figure from the paper
    NeurIPS 2026Oral nomination

    Solving Every Step Is Not Enough: Milestone Oracles Reveal a Composition Gap in LLM Math Reasoning

    Zhuohan Wang (project lead), Haoran Ma, Tianyu Wu, Yuanlin Duan, Zichun Liao, Jieming Yu

    LLMs can solve every step of a math problem on its own and still fail the whole problem, even when given a roadmap and every step's answer. OracleLadder locates the failure with increasing levels of oracle help. This composition gap is the largest of five failure types for all six models (8B to 671B), covering 33 to 48% of problems.

  • Figure from the paper
    NeurIPS 2025

    Causal Sufficiency and Necessity Improves Chain-of-Thought Reasoning

    Xiangning Yu*, Zhuohan Wang*, Linyi Yang, Haoxuan Li, Anjie Liu, Xiao Xue, Jun Wang, Mengyue Yang

    A causal view of chain-of-thought (probability of sufficiency and necessity) that identifies which reasoning steps are essential. Pruning the rest cuts 20 to 60% of tokens on hard benchmarks at comparable accuracy.

Efficient Training and Inference

  • Figure from the paper
    NeurIPS 2026Spotlight

    D-PACE: Dynamic Position-Aware Cross-Entropy for Parallel Speculative Drafting

    Tianyu Wu, Yu Yao, Zhenting Qi, Han Zheng, Chengxi Zhang, Zhuohan Wang, Haoran Ma, Zichun Liao, Himabindu Lakkaraju, Ju Li, Yilun Du

    A training loss for parallel speculative drafters that weights each draft position by how much it currently limits the accepted length. It improves wall-clock speedup with 2.3% training overhead.

  • Figure from the paper
    ICONIP 2024

    SBoRA: Low-Rank Adaptation with Regional Weight Updates

    Lai-Man Po, Y. Liu, H. Wu, T. Zhang, W.-Y. Yu, Zhuohan Wang, Z. Jiang, K. Li

    A LoRA variant that initializes one low-rank matrix with orthogonal standard-basis vectors, so only a few rows or columns of each weight change. It halves LoRA's trainable parameters, or doubles its rank at the same budget, and improves commonsense and arithmetic reasoning.

* equal contribution. Full list on Google Scholar.