About

I work on post-training for large language models at Tencent Hunyuan. In 2023, I joined Tencent through its top technical-talent recruitment program, now known as the Qingyun Program. My research aims to improve the general capabilities of LLMs by constructing synthetic RL data and developing effective RL recipes. I focus on capability dimensions including complex instruction following, long-context understanding and hallucination reduction in long-form outputs, and complex multi-turn interaction (including agents).

Before joining Tencent, I received my B.S., M.S., and Ph.D. degrees from Sun Yat-sen University, completing my Ph.D. in 2023. During my Ph.D., my research centered on natural language processing, pre-trained models, graph representation learning, and applying graph-based modeling to NLP and multimodal reasoning. I also worked as a research intern at Microsoft Research Asia and Huawei Noah's Ark Lab during this period, where I explored pre-trained models, language understanding, and related NLP problems.

Research Interests

  • Synthetic RL data for improving general LLM capabilities, including:
    • Complex instruction following under complex constraints/system-prompt settings.
    • Improving understanding of long and complex texts, and reducing hallucinations in long-form outputs.
    • Long or complex multi-turn interaction (including agents).
  • RL recipes for more effective post-training, including:
    • Reward design and optimization strategies for stable and efficient RL training.
    • Sampling, filtering, and data selection for higher-quality RL data.
  • Earlier work on pre-trained models and graph-based NLP.

Selected Publications

Recent LLM work: data construction/evaluation and RL recipes. * denotes equal contribution.

Data Construction and Capability Evaluation

  1. PlanningBench: Generating Scalable and Verifiable Planning Data for Evaluating and Training Large Language Models.
    Zenan Xu* et al. arXiv 2026 [paper] [code]
  2. CL-bench Life: Can Language Models Learn from Real-Life Context?
    ... Zenan Xu ... arXiv 2026 [paper]
  3. CL-bench: A Benchmark for Context Learning.
    ... Zenan Xu ... arXiv 2026 [paper] [code]
  4. Reinforcement Learning on Pre-Training Data.
    ... Zenan Xu ... ACL 2026 [paper]
  5. Hunyuan-TurboS: Advancing Large Language Models through Mamba-Transformer Synergy and Adaptive Chain-of-Thought.
    Hunyuan Team, including Zenan Xu. arXiv 2025 [paper]
  6. RMCBench: Benchmarking Large Language Models' Resistance to Malicious Code.
    ... Zenan Xu ... ASE 2024 [paper]

RL Recipes, Reasoning, and Training Tricks

  1. Search-R2: Enhancing Search-Integrated Reasoning via Actor-Refiner Collaboration.
    Zenan Xu* et al. ICM 2026 [paper]
  2. Adaptive Termination for Multi-round Parallel Reasoning: An Universal Semantic Entropy-Guided Framework.
    Zenan Xu et al. arXiv 2025 [paper]
  3. The Art of Efficient Reasoning: Data, Reward, and Optimization.
    Zenan Xu* et al. arXiv 2026 [project] [paper]
  4. Discovery and Reinforcement of Tool-Integrated Reasoning Chains via Rollout Trees.
    ... Zenan Xu ... arXiv 2026 [paper]
  5. Segmental Advantage Estimation: Enhancing PPO for Long-Context LLM Training.
    ... Zenan Xu ... arXiv 2026 [paper]
  6. ConMax: Confidence-Maximizing Compression for Efficient Chain-of-Thought Reasoning.
    Zenan Xu* et al. arXiv 2026 [paper]
  7. GMoE: Empowering LLMs Fine-Tuning via MoE Graph Collaboration.
    ... Zenan Xu ... arXiv 2024 [paper]

Earlier Research before LLMs

  1. Learning Summary-Worthy Visual Representation for Abstractive Summarization in Video.
    Zenan Xu et al. IJCAI 2023 [paper]
  2. A Graph Fusion Approach for Cross-Lingual Machine Reading Comprehension.
    Zenan Xu et al. AAAI 2023 [paper]
  3. Analytical Reasoning of Text.
    ... Zenan Xu ... NAACL Findings 2022 [paper]
  4. Syntax-Enhanced Pre-trained Model.
    Zenan Xu et al. ACL-IJCNLP 2021 [paper] [code]
  5. Embedding Dynamic Attributed Networks by Modeling the Evolution Processes.
    Zenan Xu et al. COLING 2020 [paper]
  6. Reasoning Over Semantic-Level Graph for Fact Checking.
    ... Zenan Xu ... ACL 2020 [paper]
  7. A Deep Neural Information Fusion Architecture for Textual Network Embeddings.
    Zenan Xu et al. EMNLP-IJCNLP 2019 [paper]

Experience

Tencent (2023 - present)

  • Work on large language model post-training.
  • Focus on synthetic RL data, RL recipes, general capability improvement, complex instruction following, long-context learning, and multi-turn interactions.

Microsoft Research Asia (2020.03 - 2021.01)

  • Research Intern, Natural Language Computing Group.
  • Worked on pre-trained models and language understanding.
  • Supervisors: Duyu Tang, Linjun Shou, and Ming Gong.

Huawei Noah's Ark Lab

  • Research Intern.
  • Worked on NLP and pre-trained model related problems.

Education

Sun Yat-sen University

  • Ph.D. in Computer Science, School of Computer Science and Engineering, 2019 - 2023.
  • M.S. in Computer Science, School of Computer Science and Engineering, 2017 - 2019.
  • B.S. in Computer Science, 2013 - 2017.