ArcaScience reads the literature article by article and rebuilds a candidate's risk profile across sixteen categories of evidence. BRB-C put one of them to the test against five frontier AI models: recover the risk profile of a drug from 600 abstracts, each one cited to its source. ArcaScience found 104 of the 108 in the label. The strongest model found 70.
Adverse reactions recovered from 600 adalimumab abstracts and matched to a 108-concept key adjudicated by a pharmacovigilance expert. Each model read the same corpus and the same 24-article evidence package, with the same schema, parser and matcher.
Exact McNemar tests at concept level, Holm-adjusted over the fixed family of five. Largest adjusted p: 3.9 × 10⁻⁸. The four ArcaScience misses are named in the paper: caesarean section, gait disturbance, malignant disease, onycholysis.
Of the concepts where the two systems disagree, 37 were found by ArcaScience alone and 3 by Gemini alone. Against GPT-5.6 the gap widens to 53 concepts.
A reaction without a corpus citation does not count. ArcaScience processes each article on its own, so the passage behind every concept survives into the output.
Paid-model ceiling recorded on 3,819 safety articles in July 2026, repriced at catalog rates with prompt caching disabled.
Indications, interactions, contraindications, special populations, dose and route, mechanism, biomarkers, pharmacogenomics, comorbidities, causality, outcomes, complications, pathogens, procedures, formulation. The same per-article pipeline runs all of them. A frontier prompt has to do the same work inside a single context window, once per category or all at once.
ArcaScience processes each article on its own, keeps the passage behind every finding, and only then consolidates. Recall measured at 96.3% on 600 abstracts holds by construction at 3 million or 30 million, because each article is read the same way.
The frontier condition packs 30 abstracts into one generation. Recall lands between 47.2% and 64.8% on one category. Packing more categories or more context into the same window trades recall for cost, a trade the paper models from the published long-context literature.
Cited risk profiling concept recovery from a frozen literature corpus: matched-basis evaluation of the ArcaScience ensemble-AI pipeline and five frontier models.
Asked about adalimumab with no document at all, the frontier models recite between 12.9% and 44.8% of the 2002 label. Adalimumab has been on the market for over two decades, so that label sits in their training data. Even with that head start, they lost the benchmark.
Given the corpus and required to cite it, the same models score 13.4% to 18.5% against that label. Every BRB-C score of record comes from source-grounded calls with valid corpus citations. Your candidate has no label to recite. For it, only what a system can find and cite in the evidence exists.
Sixteen categories over 30 million articles, costed at catalog prices with prompt caching and negotiated discounts left out. This is what building a risk profile from everything published would cost each way.
Perplexity Sonar at the low end, Claude Opus 5 at the high end, each rereading every article for every target family.
From complete reuse of the instrumented July 2026 workflow to a proportional marginal workload for the families without direct telemetry.
At 3,000 articles the same scenario prices the pipeline at $0.0162 per article against $0.0362 for Perplexity, $0.0399 for GPT-5.6, $0.0620 for Gemini, $0.2011 for Kimi and $0.4944 for Claude. Deterministic retrieval and filtering remove most raw snippets before any paid call: 533,891 in, 6,513 out.
The corpus was frozen before scoring, the answer key was built by dated adjudication, and the misses of every system, ArcaScience included, are published by name.
The pipeline scored here is the one that builds ArcaScience benefit-risk assessments: the literature read in full, sixteen categories of evidence structured, every finding traceable to its source, delivered along FDA BRF, BRAT and CIOMS XII.