Towards Data-Efficient and Trustworthy Reasoning.

I am an undergraduate student at Tongji University, advised by Prof. Miaomiao Zhang.

Since April 2025, I have been a visiting student at Westlake University, advised Prof. Yue Zhang. I also work closely with Prof. Qi Zhu at Zhejiang University, where I am researching the attention mechanisms of Vision-Language Models and working to enhance their long-context reasoning abilities. Since April 2026, I have been a research intern at the University of Virginia, advised by Prof. Yu Meng. My research focuses on post-training of LLM.

🔥 News

  • 📌 Pinned: I am looking for PhD opportunities for Fall 2027!
  • 2026.09:  🚀 We propose NSD, our new work on improving LLM reasoning through negative self-distillation!
  • 2026.08:  🌍✈️ Attended IJCAI 2026 in beautiful Bremen, Germany!
  • 2026.07:  🎉 Our paper “DeepInstructor: An Agentic AI Instructor for Experience-Driven Idea Evaluation” has been accepted by COLM 2026 Workshop LM4Sci!
  • 2026.05:  🎉 Our paper “VERA: Identifying and Leveraging Visual Evidence Retrieval Heads in Long-Context Understanding” has been accepted by IJCAI 2026!
  • 2026.04:  🎉 Our paper “Augmenting Multi-Technique Static Analysis with Large Language Models” has been accepted by ISSTA 2026!

🌳 My Experience Tree

I am committed to comprehensively understanding and researching the full pipeline of LLM training and inference. Below is my experience tree.

📝 Publications

arXiv 2026
NSD

Negative Self-Distillation: Learning to Reason by Avoiding Flaws

Rongcan Pei, Zhepei Wei, Shuyao Xu, Xinyu Zhu, Wei-Lin Chen, Yu Meng

arXiv

  • We propose Negative Self-Distillation (NSD), which improves LLM reasoning by learning to avoid errors produced by a negatively conditioned version of the model, without requiring reference answers or an external teacher. A dynamic token-level gate directs training toward reasoning errors while protecting basic language abilities. Across seven mathematical reasoning benchmarks, NSD improves average accuracy by 2.3, 7.5, and 6.0 percentage points for 1.7B, 4B, and 8B models, respectively, while retaining self-correction and improving training efficiency.
IJCAI 2026
VERA

VERA: Identifying and Leveraging Visual Evidence Retrieval Heads in Long-Context Understanding

Rongcan Pei, Huan Li, Fang Guo, Qi Zhu

Code | arXiv

  • Through attention analysis, we identify specific Visual Evidence Retrieval (VER) Heads — a sparse, dynamic set of attention heads critical for locating visual cues during reasoning. We propose VERA, a training-free framework that detects model uncertainty to trigger explicit verbalization of visual evidence. VERA achieves 21.3% relative improvement on Qwen3-VL-8B-Instruct and 20.1% on GLM-4.1V-Thinking across five benchmarks.
COLM 2026 Workshop LM4Sci
DeepInstructor

DeepInstructor: An Agentic AI Instructor for Experience-Driven Idea Evaluation

Rongcan Pei, Fang Guo, Qinglin Qi, Qi Zhu, Yun Luo, Jianhao Yan, Minjun Zhu, Qiujie Xie, Dehong Zheng, Yue Zhang

Paper | Code & Prompts

  • We propose DeepInstructor, the first agentic AI instructor that emulates how human mentors leverage past scholarly experience to guide research ideation. DeepInstructor extracts experience triples from large-scale review corpora, constructs an Experience Graph, and employs an agentic reasoning framework for experience-driven idea evaluation.
arXiv 2026
PaperIgnition

PaperIgnition & Sci-Surf Project

Qi Zhu, Fang Guo, Rongcan Pei*, Shuqi He, Hui Chen, Yue Zhang

arXiv | Website | Code

  • An intelligent paper recommendation and digestion platform that adopts a temporally decoupled architecture. The system autonomously fetches, indexes, and digests newly released arXiv papers through a daily pipeline, with multimodal LLM-based summarization executed offline.
ISSTA 2026
ELSA

Augmenting Multi-Technique Static Analysis with Large Language Models: A Neuro-Symbolic Approach to Smart Contract Vulnerability Detection

Junxiang Wang, Fu Song, Miaomiao Zhang, Bowen Du, Rongcan Pei

  • We propose ELSA, a neuro-symbolic approach that augments static analysis with LLM-assisted constraint-guided reasoning and resolves conflicting outputs through analyzer ensemble. ELSA significantly outperforms baselines across three open-source benchmarks and self-constructed ZKP-based smart contracts.

🎖 Honors and Awards

  • 2025, Social Activity Scholarship, Tongji University.
  • 2024, Guo Xie Birong Scholarship (equivalent to University First-Class Scholarship), Tongji University.

📖 Educations

  • 2023.09 - Present, Undergraduate Student, Tongji University. Advised by Prof. Miaomiao Zhang.

💻 Service

  • Invited as reviewer for KDD 2026.

💻 Experience

  • 2026.04 - Present, Research Intern, University of Virginia. Advised by Prof. Yu Meng.
  • 2025.04 - 2026.05, Visiting Student, Westlake University. Advised by Prof. Yue Zhang.
  • 2025.04 - Present, Research Collaborator, Zhejiang University. Advised by Prof. Qi Zhu.