GEPA: The Prompt Optimizer That Beat Reinforcement Learning With 35x Fewer Rollouts
Reinforcement learning methods like GRPO have become the default way to adapt LLMs to new tasks, and they typically require tens of thousands of rollouts to do it. GEPA (Genetic-Pareto), accepted as an ICLR 2026 Oral, shows that reflecting on what went wrong in natural language and evolving the prompt accordingly beats GRPO by 6% on average, by up to 20% on individual tasks, while using up to 35x fewer rollouts.