Zhuoheng Han 韩卓衡

Ph.D. student · Peking University

I work with Prof. Houfeng Wang at the Institute of Computational Linguistics, Peking University.

Research

I study LLM evaluation and agent reliability, alongside iterative coding and multimodal benchmarks. My current focus is trace-based diagnosis of long-horizon agent failures. I also collaborate on post-training refinement of LoRA adapters.

Portrait of Zhuoheng Han
Beijing, China

News

Work in progress

Under reviewFirst author

Out of Step with Evidence

I study when tool-using agents report task progress that is not supported by the available evidence. I develop trace-based methods to diagnose mismatches between observations, actions, and completion claims, and to distinguish these from tool or execution-environment failures.

I lead the research design, implementation, experiments, analysis, and writing.

Under reviewThird author

Learn the Directions, Normalize the Gains

I collaborate on post-training refinement of LoRA adapters, studying how redistributing adaptation gains affects task specialization and the retention of broader capabilities.

My contribution focuses on running a subset of the experiments.

Selected publications

* Equal contribution

2026
NLPCC 2026Co-first author

Overconfidence in LLM-as-a-Judge: Diagnosis and Confidence-Driven Solution

Zailong Tian*, Zhuoheng Han*, Yanzhe Chen, Haozhe Xu, Xi Yang, Richeng Xuan, Houfeng Wang, Lizi Liao

A study of nine LLM judges distinguishes confidence calibration from its usefulness for ranking judgments. LLM-as-a-Fuser combines decisions, rationales, and confidence signals.

My contribution

I co-designed the evaluation methodology and conducted experiments on confidence estimation and confidence-aware judgment fusion.

2025
NeurIPS 2025 · D&BSpotlightSecond author

Sheetpedia: A 300K-Spreadsheet Corpus for Spreadsheet Intelligence and LLM Fine-Tuning

Zailong Tian, Zhuoheng Han, Houfeng Wang, Lizi Liao

A corpus of more than 290,000 real-world spreadsheets, with benchmarks for semantic-range understanding and formula generation.

Examples of real-world spreadsheets in the Sheetpedia corpus
My contribution

I worked on data cleaning and standardization, language and quality filtering, and LSH-based deduplication. I also contributed to fine-tuning experiments, analysis, and writing.

Dataset work

Background

Education

Peking University

Ph.D. student in Natural Language Processing
Institute of Computational Linguistics

Advisor: Prof. Houfeng Wang

Peking University

B.S. in Computer Science and Technology
Yuanpei College

Experience

Beijing Academy of Artificial Intelligence

Research intern · FlagEval

Worked on evaluation datasets for vision-language models, focusing on text recognition, text understanding, and annotation quality.