韩卓衡 / Ph.D. student at Peking University

Zhuoheng Han

I study how language models fail under iteration, evaluation, and tool use.

I test reliability with trace evidence.

I am a Ph.D. student in the Institute of Computational Linguistics at Peking University, advised by Prof. Houfeng Wang. My work focuses on reliable language-model evaluation, agent failure diagnosis, and code intelligence.

News

NeurIPS 2025 selected Sheetpedia for a Datasets & Benchmarks Spotlight.

I began my Ph.D. in Natural Language Processing at Peking University.

Tracing how agents fail

I study long-horizon tool-using agents as they carry state across tasks, act through external tools, and report completion.

The project maps mismatches between task state, external actions, and claimed completion. I am building trace-based diagnosis and audit procedures that separate tool-chain faults from failures inside the agent.

How does an agent lose or overwrite task state?

Which trace evidence supports a failure diagnosis?

Study environments include CLAW, TerminalBench, AppWorld, tau2-bench, SWE-multilingual, and DeepSWE.

Selected publications

I work on model stability and evaluation. Sheetpedia extends that work to spreadsheet intelligence.

Examples from the Sheetpedia spreadsheet corpus
NeurIPS 2025 D&B Spotlight Second author

Sheetpedia: A 300K-Spreadsheet Corpus for Spreadsheet Intelligence and LLM Fine-Tuning

Zailong Tian, Zhuoheng Han, Houfeng Wang, Lizi Liao

We build a corpus of more than 290,000 real-world spreadsheets and two benchmarks for semantic-range understanding and formula generation.

arXiv 2025 Co-first author

Overconfidence in LLM-as-a-Judge: Diagnosis and Confidence-Driven Solution

Zailong Tian*, Zhuoheng Han*, Yanzhe Chen, Haozhe Xu, Xi Yang, Richeng Xuan, Houfeng Wang, Lizi Liao

We measure confidence-accuracy mismatch with TH-Score and combine model judgments through LLM-as-a-Fuser.

* Equal contribution

Background

Education

Peking University

Ph.D. student in Natural Language Processing, Institute of Computational Linguistics

Peking University

B.S. in Computer Science and Technology, Yuanpei College

Experience

BAAI FlagEval

Research intern working on VLM text recognition and understanding benchmarks. I also helped build the TRUE dataset and its annotation workflow.