dataqbs

OpenDiscoveryTrace: Process Traces for Evaluating AI Scientist Workflows

· Source: arXiv cs.AI

OpenDiscoveryTrace is a new dataset that records the full reasoning of AI agents dedicated to scientific research. Rather than limiting itself to final outputs—code, hypotheses, or papers—the resource captures every step of execution, including ideas, tool calls, observations, errors, reasons for revision, and self‑reported confidence levels. The archive contains 558 complete trajectories, covering 124 tasks in domains such as drug discovery, materials science, genomics, and scientific literature analysis. Seven different models were evaluated: three of the most advanced on the market and four open‑weight models, plus variants that use real‑time retrieval. Preliminary analyses, performed by assessors based on large language models, show that process traces reveal behavioral differences that are invisible when only final results are considered. For example, although the latest‑generation models achieve similar success rates (between 84 % and 89 %), one of them produces thirty times more errors than another, with distinct failure profiles between incorrect tool use and reasoning mistakes. The set also includes five benchmark tasks and baselines built with logistic regression, random forests, LSTM, and transformers, all under a CC BY 4.0 license. This initiative enables auditing of methodologies, identification of failures, and contribution to responsible governance of automated scientific agents, which is essential for ensuring reliability and reproducibility in AI‑assisted research.

Read the original article on arXiv cs.AI

This summary is an informational synthesis produced by dataqbs.com. All rights to the original content belong to its author and the cited media outlet. We act solely as curators of technology news and claim no authorship.

Read this in Español · Deutsch