AFlow
Automating agent workflows
Zhang et al. · ICLR 2025 Oral
Represents workflows as code and automatically improves them through search and execution feedback.
Explore the projectResearch / Open work
Selected contributions by our team, spanning agent collaboration, generative evaluation, and large-scale training trajectories. This work informs how we design benchmarks and data for the next generation of agents.
Automating agent workflows
Zhang et al. · ICLR 2025 Oral
Represents workflows as code and automatically improves them through search and execution feedback.
Explore the projectMulti-agent collaboration
Hong et al. · ICLR 2024 Oral
Encodes standard operating procedures into multi-agent workflows to support complex task decomposition and verification of intermediate results.
Read the paperHuman-aligned video evaluation
Han et al. · CVPR 2025
Evaluates video generation quality with a multidimensional prompt suite and assessment strategies built around multimodal language models.
Read the paperChallenging terminal tasks
Merrill et al. · 2026
Measures long-horizon agent capabilities through real terminal tasks, executable environments, and independent verifiers.
Explore the benchmarkData recipes for reasoning
Guha et al. · 2025
Opens up reasoning-data recipes and training and evaluation workflows to study how data selection shapes model capabilities.
Explore the projectData recipes for agents
Raoof et al. · 2026
Uses open tasks, trajectories, and verification environments to investigate how to compose and scale training data for general-purpose agents.
Explore the projectA compact agentic benchmark
Shi et al. · 2026
Developed with contributions from Kevin Xiang Li, combining difficulty filtering, human audits, and an iterative process for improving task quality.
Explore the benchmarkExecution-free agentic trajectories
Li · 2026
Released by Kevin Xiang Li, a collection of 12.29 million agentic coding trajectories generated without executing repository-specific tests or commands.
Explore the datasetInterested in a research collaboration?