James V. Roggeveen
Agent Evaluation | Scientific Reasoning | Applied Mathematics
I am a researcher and applied mathematician building agent evaluations for scientific and engineering domains. My work combines expert-designed tasks, programmatic graders, trajectory analysis, and differentiable scientific computing to test and extend the capabilities of frontier AI systems.
I am currently a Postdoctoral Fellow at Harvard University with Prof. Michael Brenner and a Visiting Faculty Researcher at Google DeepMind. At Google DeepMind, I contribute to agentic evaluations for scientific and engineering domains and was publicly credited for foundational support to the Gemini Deep Think team.
My recent work is focused on three connected areas:
- Scientific AI Evaluation: I build infrastructure for evaluating language-model agents on authentic scientific and engineering work. I co-led HARDMath2, a 211-problem benchmark for graduate-level applied mathematics, and built evaluation infrastructure for CMT-Benchmark, a condensed-matter theory benchmark built by expert researchers.
- Meshless PDE Solvers: I developed a meshless, differentiable spectral method for PDE inverse problems on irregular geometries. A useful pattern has been to require coding agents to pass independent numerical checks before reporting results, which helps catch invalid implementations while extending the solver to new applications.
- Automated Model Discovery: I build differentiable tools for discovering physical laws from partial measurements, including a framework for identifying tensor-valued constitutive laws from sparse flow data.
My Ph.D. in Mechanical Engineering at Princeton University, advised by Prof. Howard A. Stone, focused on particle dynamics in low-Reynolds-number flows and creating improved hydrodynamic models for characterizing the material properties of biological condensates using micropipette aspiration.
Across these projects, I am especially interested in building systems that let domain experts contribute high-quality scientific tasks without needing to own the full evaluation harness. My background gives me both sides of that problem: deep experience in physical modeling and scientific computing, and hands-on experience designing benchmark infrastructure for frontier language models.
Current focus: agent evaluations for scientific and engineering domains; expert-built benchmarks; programmatic grading; symbolic and numerical verification; independent checks for agent-generated scientific code; and failure analysis for frontier-model behavior.
news
| Feb 12, 2026 | CMT-Benchmark was accepted as an ICLR 2026 poster, and Google cited it in the Gemini 3 Deep Think release as an advanced theoretical-physics benchmark. |
|---|---|
| Feb 11, 2026 | Google DeepMind publicly credited me for foundational support to the Gemini Deep Think team. |
| Oct 29, 2025 | Check out our preprint on solving inverse problems over PDEs with a meshelss basis fitting method, out now on arXiv. |
| Oct 28, 2025 | Check out our preprint on fitting viscoelastic models to flow data using automatic differentiation, out now on arXiv. |
| Oct 06, 2025 | Check out our preprint on CMT-Bench, an LLM benchmark focused on condensed matter physics built by a panel of experts, out now on arXiv. |