Reinforcement Learning

AgentCPM-Explore: Realizing Long-Horizon Deep Exploration for Edge-Scale Agents

ArXiv 2026. First author. Open-source 4B agent model achieving SOTA on GAIA & HLE, surpassing GPT-5 and Claude-4.5-Sonnet.

avatar
Haotian Chen

AgentCPM-Explore: Realizing Long-Horizon Deep Exploration for Edge-Scale Agents

ArXiv 2026. First author. Open-source 4B agent model achieving SOTA on GAIA & HLE, surpassing GPT-5 and Claude-4.5-Sonnet.

avatar
Haotian Chen

AgentCPM-Explore

🏆 **Project Lead** · Open-source 4B agent model achieving SOTA on GAIA & HLE benchmarks …

AgentCPM-Explore

🏆 **Project Lead** · Open-source 4B agent model achieving SOTA on GAIA & HLE benchmarks …

Reflective Reinforcement Tool Learning

Submitted to ACL 2026. Reflective reinforcement learning for tool learning.

avatar
Haotian Chen

Reflective Reinforcement Tool Learning

Submitted to ACL 2026. Reflective reinforcement learning for tool learning.

avatar
Haotian Chen

AgentRL

🏆 **Project Lead** · Fully asynchronous agent RL training infrastructure for the AgentCPM model family `100+ Tools` · `20+ Benchmarks` · `Full-cycle Visualization`

AgentRL

🏆 **Project Lead** · Fully asynchronous agent RL training infrastructure for the AgentCPM model family `100+ Tools` · `20+ Benchmarks` · `Full-cycle Visualization`

AgentCPM-GUI: Building Mobile-Use Agents with Reinforcement Fine-Tuning

EMNLP 2025 Demo. GUI agents with reinforcement fine-tuning. 1,200+ GitHub Stars.

zhong-zhang

AgentCPM-GUI: Building Mobile-Use Agents with Reinforcement Fine-Tuning

EMNLP 2025 Demo. GUI agents with reinforcement fine-tuning. 1,200+ GitHub Stars.

zhong-zhang