Minimal orbital geometry render

Build powerful and efficient AI for our future

Shengji Tang 唐圣汲

I am a PhD student in Information Engineering at The Chinese University of Hong Kong, advised by Prof. Wanli Ouyang. My research focuses on generalizable agentic reasoning, LLM post-training, multi-agent routing and aggregation, and efficient AI models. Don't hesitate to contact me for discussion and cooperation!

01

Academic Path

CUHK logo

The Chinese University of Hong Kong

  • PhDInformation Engineering, MM Lab
  • ResearchGeneralizable agent reasoning
  • AdvisorProf. Wanli Ouyang
Fudan University logo

Fudan University

  • M.Eng.Electronic Information Engineering, EDL Lab
  • ResearchEfficient and high-performance AI models
  • AdvisorProf. Tao Chen
  • HonorShanghai Outstanding Graduate
Fudan University logo

Fudan University

  • B.Eng.Electronic Information Engineering
  • GPA3.6/4.0, ranked 5/76
  • HonorsExcellent Engineer Class; Huawei Scholarship, top 5%
02

Research Vectors

01

Agentic Reasoning

Generalizable reasoning agents, agentic trajectory construction, tool-use reasoning, and multi-stage post-training.

02

Multi-agent Systems

Routing and aggregation over heterogeneous open-source LLMs for math, code, science, commonsense, and high-difficulty benchmarks.

03

Efficient AI Models

Subnet-based enhanced training, pruning, distillation, and reproducible model compression pipelines.

03

Publications

arXiv logoArxiv 2026
Benchmark figure for Agents-A1

Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent

Lei Bai, Zongsheng Cao, Yang Chen, Zhiyao Cui, Shangheng Du, Yue Fan, Shiyang Feng, Zijie Guo, Haonan He, Liang He, Xiaohan He, Shuyue Hu, Yusong Hu, Songtao Huang, Yichen Jiang, Hao Li, Xin Li, Dahua Lin, Weihao Lin, Fenghua Ling, Dongrui Liu, Zhuo Liu, Runmin Ma, Chunjiang Mu, Haoyang Peng, Tianshuo Peng, Jinxin Shi, Luohe Shi, Boyuan Sun, Zelin Tan, Shengji Tang, Qianyi Wang, Yiming Wu, Yi Xie, Xiangchao Yan, Jingqi Ye, Peng Ye, Fangchen Yu, Jiakang Yuan, Bihao Zhan, Bo Zhang, Chen Zhang, Shufei Zhang, Shuaiyu Zhang, Wenlong Zhang, Yiqun Zhang, Junpeng Zhao, Zhijie Zhong, Bowen Zhou, Yuhao Zhou.

Introduces Agents-A1, a 35B Mixture-of-Experts agentic model that scales long-horizon trajectories and heterogeneous agent abilities to reach or match trillion-parameter-level performance on long-horizon benchmarks.

ICML logoICML 2026
Figure for Beyond Gemini-3-Pro

Beyond Gemini-3-Pro: Revisiting LLM Routing and Aggregation at Scale

Shengji Tang, Weihao Lin, Peng Ye, Jingqi Ye, Hao Li, Yiqun Zhang, Xiaosong Wang, Bo Zhang, Shuyue Hu, Tao Chen, Lei Bai, Wanli Ouyang.

Large-scale routing and aggregation over open-source LLMs, showing how heterogeneous model collaboration can outperform a strong closed-source model with substantially lower cost.

ACL logoACL 2026
Figure for scalable multi-agent system

Open-Source LLMs Collaboration Beats Closed-Source LLMs: A Scalable Multi-Agent System

Shengji Tang, Jianjian Cao, Weihao Lin, Jiale Hong, Bo Zhang, Shuyue Hu, Lei Bai, Tao Chen, Wanli Ouyang, Peng Ye.

A scalable multi-agent reasoning framework that routes tasks to complementary open-source LLMs and aggregates answers for stronger complex-task performance.

ICLR logoICLR 2025
Figure for HiSplat

HiSplat: Hierarchical 3D Gaussian Splatting for Generalizable Sparse-view Reconstruction

Shengji Tang, Weicai Ye, Peng Ye, Weihao Lin, Yang Zhou, Tao Chen, Wanli Ouyang.

A hierarchical 3D Gaussian representation for sparse-view reconstruction, balancing cross-scene generalization, geometric consistency, and rendering quality.

AAAI logoAAAI 2024
Figure for Group Knowledge

Boosting Residual Networks with Group Knowledge

Shengji Tang, Peng Ye, Baopu Li, Weihao Lin, Tao Chen, Tong He, Chong Yu, Wanli Ouyang.

Introduces group knowledge transfer to strengthen residual networks by organizing subnet behaviors into coordinated groups. The method encourages richer interaction among residual branches and improves both training dynamics and final recognition performance.

ECCV logoECCV 2024
Figure for enhanced sparsification via stimulative training

Enhanced Sparsification via Stimulative Training

Shengji Tang, Weihao Lin, Hancheng Ye, Peng Ye, Chong Yu, Baopu Li, Tao Chen.

Extends stimulative training into a stronger sparsification framework that explicitly improves subnet quality during optimization. By making sparse sub-networks more competitive before pruning, it improves pruning stability and downstream lightweight-model accuracy.

NeurIPS logoNIPS 2024
Figure for S2HPruner

S2HPruner: Soft-to-Hard Distillation Bridges the Discretization Gap in Pruning

Weihao Lin, Shengji Tang, Chong Yu, Peng Ye, Tao Chen. Equal contribution.

Studies the discretization gap that appears when soft pruning decisions are converted into hard sparse networks. The proposed soft-to-hard distillation strategy transfers smoother optimization signals into deployable pruned models with stronger accuracy retention.

NeurIPS logoNIPS 2022
Figure for stimulative training

Stimulative Training of Residual Networks: A Social Psychology Perspective of Loafing

Peng Ye, Shengji Tang, Baopu Li, Tao Chen, Wanli Ouyang. Equal contribution.

Analyzes residual networks from a social psychology perspective, connecting performance degradation to inactive or under-contributing residual branches. The stimulative training strategy encourages broader subnet participation and improves the effectiveness of deep residual model training.

04

Academic Service

ICML logo Gold Reviewer
CVPR logo Reviewer
ECCV logo Reviewer
ICLR logo Reviewer
NeurIPS logo Reviewer
AAAI logo Reviewer
Reviewer