Evaluation and training of scientific research agents

New Atlantic partners with university labs to test and improve the capacity of agents to support scientific discovery.

A curling Atlantic wave formed from fine charcoal dots.

Agents free up labs' time to focus more on the
questions at the core of their research, while improving
the pace and success rate of experimental work.

See example tasks

We develop research agents that can run
scientific software, check results, and help
researchers decide which hypotheses to test next.

Helping scientists turn evidence into discoveryLife sciences & materials science
01 /

Research agents

What could agents
do for a research lab?

After collecting data, researchers still need to check its quality, choose an analysis, run software, and determine whether the result supports their hypothesis. A different assumption or processing choice can change the conclusion.

Capable agents could do more of this analysis under a scientist’s direction: rerun a study, compare alternative methods, and identify which findings warrant another experiment.

A research agent is an AI system that can write and execute code, use scientific software, inspect outputs, and revise its next step. These are examples of the work we want such systems to be able to do.

I.

Analyze a treatment’s effect.

For a bulk RNA-sequencing study, an agent should be able to check sample labels and gene counts, fit a model comparing treated and control samples while accounting for batch differences where the study design allows, and report gene-expression changes with uncertainty and multiple-testing correction. This work gives a biologist a documented set of candidates for follow-up experiments.

II.

Test an analysis choice.

For a cryo-EM dataset, an agent should be able to compare particle selections and three-dimensional reconstructions, then check whether a protein feature remains visible across defensible processing choices. These checks help a structural biologist distinguish a supported feature from an artifact before interpreting its biological role.

III.

Prioritize an experiment.

For a battery study, an agent should be able to combine electrolyte recipes with charge–discharge records, compare capacity retention under matched test conditions, and rank formulations for further testing. Reporting uncertainty alongside each prediction helps a materials scientist choose between testing a likely improvement and resolving a gap in the data.

If these analyses become reliable and faster to run, scientists could test more explanations against existing data and choose new experiments with better evidence.

02 /

The source data

A paper rarely records
every decision.

A published figure may summarize weeks of processing. The files behind it can show which samples were excluded, which parameters changed, which analysis failed, and why the researchers accepted the final result.

We work with university labs to recover those records and explain the decisions. These records let us build tasks that test whether an agent can select data, diagnose an error, or justify a method using the evidence available at that step.

Records we work with
  • Raw measurements & sample metadata
  • Notebooks & experiment protocols
  • Analysis code & scientific software
  • Intermediate outputs & run logs
  • Failed analyses & negative results
  • Exclusion criteria & decision notes
03 /

Our work

We build tasks an agent
can run and be tested on.

Each environment combines a scientific question, the data and software needed to investigate it, and checks on the agent’s work. These environments can be used to evaluate and improve the performance of frontier AI models in supporting scientists' work.

01   /   Reconstruct

Reproduce the analysis.

We start with a completed, peer-reviewed study. With the lab, we trace raw data through code, intermediate outputs, and researcher decisions to the published result. The initial scope is computational work after data collection.

02   /   Build

Package a runnable task.

We specify the question, provide permitted inputs, and package the software so the analysis can be rerun. We reserve separate data for independent evaluation and keep reference answers out of task inputs, so agents are tested on evidence they have not learned from.

03   /   Evaluate

Check outputs and methods.

We build automated checks and lab-defined criteria: do the outputs reproduce, are comparisons valid, and does performance hold on withheld data? We compare with baseline methods and check for shortcuts such as reading an answer from a file.

Data rights & licensing

Universities retain ownership of their data. We coordinate permissions dataset by dataset, including restrictions from research sponsors. We keep researchers connected to downstream value by sharing revenue with the university and originating lab.

04 /

In practice

The inputs, the task,
and the checks.

These examples show how research workflows become tasks for evaluating and improving an agent’s ability to carry out scientific analyses.

Life sciences

01

Cryo-electron
microscopy

Reconstruct a protein density map from microscope movies.

An agent should be able to use RELION to correct motion, estimate microscope contrast effects, select particle images, separate them into classes, and refine a three-dimensional map. It should return the map, processing settings, and a record of which images it retained.

The scientific question: which structural features are supported by the images? A sharper-looking map alone is insufficient; the analysis must estimate resolution and check for overfitting.

Inputs & evaluation
Inputs
Recorded microscope movies, acquisition metadata, and processing software. The lab’s reconstruction and decision records provide references for designing the task.
Checks
Maintain independent particle half-sets during refinement. Measure agreement between their maps using Fourier shell correlation, account for masking effects, and inspect local resolution. Compare processing choices with the lab’s documented analysis.

Materials science

02

Battery electrolyte
selection

Rank formulations using recorded cycling experiments.

An agent should be able to join electrolyte recipes to cell test records, calculate capacity retention at a specified cycle count, and fit a model to rank formulations. It should compare cells under matched electrode chemistry, voltage limits, temperature, and cycling rate.

The scientific question: which formulation merits another test, and how uncertain is that choice? The task uses a fixed set of completed experiments whose outcomes can be withheld, then revealed as the agent selects candidates.

Inputs & evaluation
Inputs
Solvent, salt, and additive compositions; cell identifiers; cycling protocols; charge–discharge curves; and replicate measurements from completed experiments.
Checks
Keep replicate cells of the same formulation in the same data split. Evaluate prediction error, uncertainty calibration, and the measured performance of selected candidates against random selection and a simple baseline. Score only candidates with recorded outcomes; new formulations still need physical testing.

For universities & research labs

Research agents for
your lab’s workflows.

We work with your lab to reconstruct the computational workflow behind a published study and define criteria for evaluating an agent’s performance. Your lab’s data, code, and analysis records help us evaluate and improve how agents apply your methods and interpret evidence.

The aim is to develop agents that can reliably carry out these analyses under your team’s direction, reducing the time required to compare methods, test hypotheses, and assess evidence for further experiments.

Discuss a research workflow