News

Can Your AI Engineer a Robot?

Harvard and Georgia Tech’s RLE-Bench tests coding agents for physical systems

Key Takeaways

  • Harvard computer scientists have released RLE-Bench, an evaluation tool that tests how well AI coding agents can perform the task of engineering a physical robot.
  • The open-source benchmark allows researchers and developers to evaluate new and existing AI systems on a common set of robotic tasks. 

With the rapid rise of artificial intelligence in daily life, software coding has become increasingly automated, with powerful AI systems known as coding agents able to write and revise computer programs almost autonomously.

But what happens when an AI agent must contend not just with digital command lines, but with the physical world of robotics?

Researchers at the Harvard John A. Paulson School of Engineering and Applied Sciences (SEAS) and the Georgia Tech School of Computational Science and Engineering are taking a systematic approach to finding out.

A team led by Na Li, the Winokur Family Professor of Electrical Engineering and Applied Mathematics at Harvard SEAS, and Bo Dai, assistant professor at Georgia Tech, have developed an evaluation tool known as a benchmark that tests how well AI coding agents can perform the challenging task of engineering an actual, physical robot. The team includes Harvard graduate student Haitong Ma and Chenxiao Gao and Rushi Qiang at Georgia Tech. 

The benchmark, known as RLE-Bench, is like a standardized test for an AI robot learning engineer — a role that combines machine learning, robotics, control, software, and hardware.

An overview of the RLE-Bench framework. 

The new benchmark contains 48 tasks that test whether general-purpose coding agents can perform the kinds of engineering work required to build and operate robotic systems. Tasks range from developing control and perception algorithms to designing robot hardware and interfaces.

“Benchmarks are important because they give the field a common measuring stick, allowing us to compare systems under the same conditions and identify where important capability gaps remain,” Li said. “Robotics already has many valuable benchmarks, but many focus primarily on learning and evaluating control policies. RLE-Bench takes a broader view. Can an AI system perform the range of engineering tasks required to make a robotic system actually work?”

RLE-Bench organizes its tasks into four broad areas: interactive control, policy development, perception and estimation, and mechanical design. Together, they test not only an agent’s ability to generate code, but also its ability to reason about sensing, dynamics, stability, hardware constraints, and the interaction between software and the physical world.

The tasks run in simulated environments where agents can generate, execute, inspect, and refine their solutions. But the problems are grounded in physical constraints. A design that appears functional, for example, may become unstable once a payload is added or fail because the agent has not correctly accounted for mass, torque, or geometry.

In one example task, an agent must design a universal mobile base for multiple robotic arms that needs to reach objects at different heights on a shelf. A proposed design may allow the arm to reach every target while still failing certain stability tests where the complete robot would tip over.

“That is where physical reasoning becomes critical,” Ma said. “An agent can produce something that looks reasonable computationally, but once you consider the physics of the complete system, important failure modes can appear.”

A coding agent is shown engineering a mobile manipulator base that tips over, illustrating the need to benchmark different designs. 

RLE-Bench is open source and publicly available, allowing researchers to evaluate new AI systems on a common set of tasks.

The researchers view the current release as a foundation for a broader community benchmark. Robot learning engineers work on many problems that are not yet represented in RLE-Bench, including new forms of hardware design, simulation, system integration, debugging, safety, and deployment.

“RLE-Bench 1.0 is only a starting point,” Li said. “We hope researchers and robotics practitioners will contribute tasks based on the engineering problems they encounter in their own work. Our goal for RLE-Bench 2.0 is to capture a much broader picture of what an AI system needs to do to function as a robot learning engineer.”

Topics: AI / Machine Learning, Applied Computation, Applied Mathematics, Computational Science & Engineering, Computer Science, Electrical & Computer Engineering, Research, Robotics, Technology

Scientist Profiles

Na Li

Winokur Family Professor of Electrical Engineering and Applied Mathematics

Press Contact

Anne J. Manning | amanning@seas.harvard.edu