Exclusive: Harmonic launches new AI math benchmark
Add Axios as your preferred source to
see more of our stories on Google.

Illustration: Sarah Grillo/Axios
Nvidia-backed AI startup Harmonic is partnering with the American Institute of Mathematics on a new evaluation process designed by mathematicians, the startup told Axios exclusively.
Why it matters: AI companies often use benchmark scores to show that their models are improving. The partnership reflects a growing push to give domain experts more say in how that progress is measured.
Driving the news: The American Institute of Mathematics (AIM) and Harmonic will jointly develop an open benchmark based on mathematical problems selected by working mathematicians.
- The initial benchmark includes more than 50 number theory problems, ranging from generating examples to tackling open research questions.
- Unlike many existing AI evaluations, the framework will measure both whether models produce correct answers and whether they help mathematicians make progress on difficult research problems.
- Harmonic and AIM plan to release the benchmark materials and evaluation criteria publicly, allowing academic and industry researchers to contribute problems and results.
What they're saying: "Harmonic's goal is to amplify human creativity and discovery by designing AI tools that complement human insight," Harmonic CEO Tudor Achim said in a statement.
- AIM Executive Director Sergei Gukov called the partnership "an important step" because it centers the needs of mathematicians. He said AIM hopes to release additional benchmarks covering other research areas.
Zoom out: Companies say the next generation of evaluations should test how AI helps experts solve real-world problems, rather than how well it does on standardized exams.
- OpenAI recently challenged one of the field's most widely used evaluations, arguing that many of the benchmarks were broken.
- The company called for benchmarks "built by experienced software developers specifically to test model capabilities."
Zoom in: Harmonic has long argued that mathematics offers a particularly clear test of AI reasoning because many answers can be formally verified.
- The company, founded in 2023, is backed by Robinhood CEO Vlad Tenev.
The bottom line: Experts in specific fields are taking a larger role in defining what counts as progress for AI.
