COLM 2026 accepted Fragility Under Pressure.
韩卓衡 / Ph.D. student at Peking University
Zhuoheng Han
I study how language models fail under iteration, evaluation, and tool use.
I test reliability with trace evidence.
I am a Ph.D. student in the Institute of Computational Linguistics at Peking University, advised by Prof. Houfeng Wang. My work focuses on reliable language-model evaluation, agent failure diagnosis, and code intelligence.
News
NeurIPS 2025 selected Sheetpedia for a Datasets & Benchmarks Spotlight.
I began my Ph.D. in Natural Language Processing at Peking University.
Current research
Tracing how agents fail
I study long-horizon tool-using agents as they carry state across tasks, act through external tools, and report completion.
The project maps mismatches between task state, external actions, and claimed completion. I am building trace-based diagnosis and audit procedures that separate tool-chain faults from failures inside the agent.
How does an agent lose or overwrite task state?
Which trace evidence supports a failure diagnosis?
Study environments include CLAW, TerminalBench, AppWorld, tau2-bench, SWE-multilingual, and DeepSWE.
Selected publications
I work on model stability and evaluation. Sheetpedia extends that work to spreadsheet intelligence.
Fragility Under Pressure: Evaluating the Iterative Stability of LLMs in Constrained Interactive Coding
We introduce STRIDE to test whether coding agents preserve instructions and functionality through repeated edit-verify-debug cycles.
Overconfidence in LLM-as-a-Judge: Diagnosis and Confidence-Driven Solution
We measure confidence-accuracy mismatch with TH-Score and combine model judgments through LLM-as-a-Fuser.
* Equal contribution
Background
Education
Peking University
Ph.D. student in Natural Language Processing, Institute of Computational Linguistics
Peking University
B.S. in Computer Science and Technology, Yuanpei College
Experience
BAAI FlagEval
Research intern working on VLM text recognition and understanding benchmarks. I also helped build the TRUE dataset and its annotation workflow.