About me
Iโm a second year master student from Tsinghua University (THU), Shenzhen Graduate School. My advisor is Prof. Yujiu Yang. Prior to this, I obtained my bachelorโs degree from the Department of Electronic Engineering at Tsinghua University in June 2025.
I am currently interning in Tencent AI Lab, now Hunyuan Frontier, focus on memory compression in long context scenarios. Before this, I interned in ByteDance, focusing on post-training of MLLM for e-commerce scenarios. Prior to this, I interned in TeleAI, working on the post-training of the Telechat series foundation models. I have also interned in Tencent before for kernel development.
My research interest includes Reward Model/LLM-as-a-Judge and LLM/MLLM post training.
Activities
- 2026.9 ๐ Think-with-rubrics: From external evaluator to internal reasoning guidance is accepted by NeurIPS 2026 Main Track!
- 2026.8 ๐ Two papers are accepted by EMNLP 2026 Findings!
- 2026.6 ๐ Our new work FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention has been posted on arxiv!
- 2025.12 ๐ง I start my internship in Tencent AI Lab, focusing on long-context compression.
- 2025.9 ๐ I enroll in the Shenzhen Graduate School of Tsinghua University to pursue a Master degree, with an expected graduation date in 2028.
- 2025.8 ๐ Improve llm-as-a-judge ability as a general ability is accepted by EMNLP 2025 Main Conference. Welcome to use our SOTA LLM-as-a-Judge model RISE-Judge!
- 2025.6 ๐ I graduate from Department of Electronic Engineering, Tsinghua University and get bachelor degree.
- 2025.4 ๐ผ I begin an internship at ByteDance, focusing on MLLM post-training research.
Publications
- Think-with-rubrics: From external evaluator to internal reasoning guidance
- Jiachen Yu*, Zhihao Xu*, Junjie Wang, Yujiu Yang
- NeurIPS 2026 Main Track
- S2J: Bridging the Gap Between Solving and Judging Ability in Generative Reward Models
- Shaoning Sun*, Jiachen Yu*, Zongqi Wang, Xuewei Yang, Tianle Gu, Yujiu Yang
- EMNLP 2026 Findings
- Internalize the Temperature: On-Policy Self-Distillation as Policy Reheater for Reinforcement Learning
- Xuewei Yang*, Jiachen Yu*, Jie Wu, Shaoning Sun, Junjie Wang, Yujiu Yang
- EMNLP 2026 Findings
- Improve llm-as-a-judge ability as a general ability
- Jiachen Yu*, Shaoning Sun*, Xiaohui Hu, Jiaxu Yan, Kaidong Yu, Xuelong Li
- EMNLP 2025 Main Conference
- Finished during internship in TeleAI.
Technical Reports
- FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention
- Yan Wang*, Qifan Zhang*, Jiachen Yu*, Tian Liang*, Dongyang Ma*, Xiang Hu, Zibo Lin, Chunyang Li, Zhichao Wang, Miao Peng, Nuo Chen, Jia Li, Yujiu Yang, Haitao Mi, Dong Yu
- Finished during internship in Tencent Hunyuan Frontier
Preprints
- CoWork-X: Experience-Optimized Co-Evolution for Multi-Agent Collaboration System
- Zexin Lin*, Jiachen Yu*, Haoyang Zhang, Yuzhao Li, Zhonghang Li, Yujiu Yang, Junjie Wang, Xiaoqiang Ji
- VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories?
- Jiachen Yu*, Yufei Zhan*, Ziheng Wu, Yujiu Yang, Yousong Zhu, Jinqiao Wang
- Finished during internship in ByteDance.
Internship
- 2025.12-2026.9: Tencent, AI Lab -> Hunyuan Frontier
- 2025.4-2025.11: Bytedance, Data, Global E-commerce
- 2024.10-2025.3: TeleAI
- 2024.6-2024.8: Tencent, TEG
Interests
- Riding, city walk, Running
- Computer Card Games (HearthStone et al.)
