Validation Update: New Data, Biological Benchmarking, and LLM Stress Testing in Darwin

Author: CAIRN Institute

Published: August 24, 2026

Read time: 10–12 minutes

#Genomics #Bioinformatics #GeneRegulation #SystemsBiology #OpenScience #AgenticAI

Explore the repository on GitHub →

Darwin is powered by the CAIRN Institute GRN Atlas.

Quick Summary

Today we are publishing a validation-focused update for Darwin.

Over the past week, we expanded the atlas data surface, added and hardened a large set of new research skills, and ran three different validation layers:

The current repository state is materially stronger than the initial public release:

Most importantly, the new validation work was not limited to software contracts. We also reran direct biological benchmarks on the network itself and completed a milestone validation suite across import, TF activity, pathway activity, chromatin support, trajectory workflows, dsRNA, CRISPR, perturbation calibration, transferability, and packaged workflow generation.

Why This Update Matters

Many biological tools are evaluated only at the interface level.

That is not enough for a system like Darwin.

If a platform claims to support regulatory-network reasoning, intervention planning, and cross-species hypothesis transfer, then three things have to be tested separately:

This update is about method hardening. It is the difference between “the code runs” and “the scientific and workflow surfaces have been exercised directly.”

What New Data Was Added

The atlas now spans a broader species and data surface than the initial release branch.

Current species coverage is:

with dahlia onboarding prepared.

The atlas continues to combine curated regulation, inferred edges, promoter and motif context, orthology and transfer layers, pathway and trait annotation, perturbation surfaces, RNAi / dsRNA design surfaces, CRISPR-oriented heuristic surfaces, and importable omics, chromatin, and workflow packaging layers.

At the network level, the current build includes:

This is not just more data. It is a broader surface for validation, especially in plant and non-model workflows.

Direct Validation of Inferred and Projected Regulatory Edges

One of the key questions for Darwin is whether the inferred and projected edges are good enough to use as part of a defensible workflow.

We reran the direct gold-standard quality checks against the current build.

Gold-standard edge quality

For species where we have explicit positive and negative control sets, the refreshed reports remained strong:

SpeciesRecallSpecificityPrecision
Petunia93.75% (30/32)100.0%100.0%
Tomato84.21% (32/38)100.0%100.0%

Those numbers matter because petunia and tomato are exactly the kinds of species where a researcher is most likely to worry that projected or inferred layers are weaker than the canonical human or Arabidopsis surfaces.

Independent benchmark validation

We also reran the independent benchmark surfaces:

BenchmarkResult
Arabidopsis vs DAP-seq AUROC0.8801
Arabidopsis vs DAP-seq AUPRC0.6990
Arabidopsis vs DAP-seq precision@1000.9000
Human DoRothEA vs TRRUST AUROC1.0000
Human DoRothEA vs TRRUST AUPRC1.0000
Human DoRothEA vs TRRUST precision@1001.0000

These are important for two reasons:

Statistical Validation at the Network Level

We also regenerated the population-level network validation reports across the current species build.

Headline results:

SpeciesEdgesCoherence (σ)Multi-evidence zMotif enrichment
arabidopsis919,4491.710.28
tomato248,28835.172.3638.55x
petunia236,72730.252.0829.53x
human17,9466.82-2.24
mouse17,6920
rice16,933027.99x
potato11,40903.33x
pepper2,212028.5x

This layer is useful because it asks a different question from the gold-standard reports. Instead of asking “did this edge match a held-out truth set?”, it asks whether the network as a population shows the kinds of structural and motif-support patterns we expect from a plausible regulatory graph.

The New Validation Suite

Beyond the legacy validation scripts, we implemented and ran a full milestone benchmark suite under backend/scripts/.

The full suite completed successfully on Friday, August 21, 2026:

That means every planned validation script for the roadmap now exists and executed successfully.

Milestone benchmark coverage

MilestoneAreaStatus
M1omics import foundationpass
M2apathway activitypass
M2bTF activitypass
M3cell-type workflowspass
M4chromatin layerpass
M5trajectory workflowspass
M6RNAi / dsRNApass
M7CRISPR heuristicspass
M8perturbation calibrationpass
M9signaling → TFpass
M10living validation dashboardpass
M11transferability / onboardingpass
M12workflow packagingpass

Summary across the 12 milestone benchmark files after rerun:

What Had to Be Hardened

The validation work did not simply confirm that everything was already correct.

Several areas needed method hardening before the suite came back clean:

One important example was TF activity.

Earlier, the TP53-like signature benchmarks showed a failure mode where tiny perfect-overlap regulons could outrank TFs that explained more of the user’s signature. After hardening, literal TP53 recovery passed rank-1 for both ulm and wmean, while synthetic self-consistency cases still passed.

That is exactly the kind of fix that matters scientifically. It is not cosmetic. It changes whether an activity-scoring surface is trustworthy enough to be used in downstream reasoning.

New Skills and Expanded Workflow Surface

The skill layer has expanded substantially since the earlier release state.

The repository now contains 100 documented Darwin research skills:

The newer parts of the skill surface include major additions in these families:

Omics and cell-state workflows

Chromatin and promoter-support workflows

RNAi and CRISPR follow-up layers

Validation, calibration, and decision-support workflows

These additions matter because they push the system beyond static network browsing into import-first analysis, assay-oriented follow-up, and collaborator-facing handoff workflows.

A regulatory-network view in Darwin showing signed edges, confidence encoding, gene metadata, and expression context.
A regulatory-network view in Darwin showing signed edges, confidence encoding, gene metadata, and expression context. The same atlas surfaces validated in the benchmark suite are also exposed interactively in the UI and through the skill layer.

Single-Skill Testing

The single-skill story is now broader than the older baseline matrices.

There are two relevant views of single-skill testing in the repository:

1. Direct skill execution harness

This is the most literal check of whether each skill runs correctly through its own scripts/run.py entrypoint against the backend.

Current result:

2. Natural-language single-skill coverage inventory

This asks a different question: if a model receives a natural-language request, do we have routing cases that exercise the full skill surface?

Current inventory:

Historically, the clean baseline GPT-5.4 single-skill matrix remains:

The later expansion work brought the repository to the current 386-case inventory and full 100-skill surface coverage.

GPT-5.4 Orchestrator Testing

The strongest current orchestration result in the repository is the expanded GPT-5.4 matrix.

Current result:

This expanded orchestration layer includes import → contrast → upstream and trajectory chains, CRISPR design plus promoter and motif follow-up, inferred-edge validation against module structure, pathway activity plus phenotype or trait interpretation, conservation plus transferability plus family-rescue chains, and phenotype-targeting plus validation-plan handoff.

In practical terms, GPT-5.4 is now clean on the full expanded workflow surface currently represented in the repository.

Nemotron-3-Ultra Orchestrator Testing

We also ran the same expanded workflow surface through Nvidia Nemotron-3-Ultra via OpenRouter, with slower pacing to reduce avoidable provider noise.

Paced expanded run result on Saturday, August 22, 2026:

This is still useful. It tells us two things at once:

Nemotron fail families

The persistent fail set clustered into a few clear groups:

FamilyQuestionsPattern
Comparison under-chainingQ3, Q24partial retrieval without completing overlap, gene-info, or enrichment follow-up
Phenotype-first petunia planning and rankingQ36, Q50, Q54, Q81, Q83candidate discovery happened, but ranking, RNAi framing, or validation handoff was weak or incomplete
Intervention tradeoff / capability boundaryQ53, Q56model touched the right surface but did not complete the required planning synthesis
Cross-species transfer / family-rescue synthesisQ40, Q67weak transferability framing, sometimes compounded by provider overload
Motif / promoter / edit planningQ43, Q78started the motif side of the workflow but did not finish the full promoter/edit interpretation
Import-first returned-id workflowsQ65, Q66, Q71, Q77hardest current family for Nemotron
Decision-boundary / calibration / counterfactual synthesisQ60, Q72, Q76the right tools were partially called, but the final structure was not satisfied

The paced run also surfaced a few real chain-path issues that are not merely model weakness:

That is useful signal. It means the Nemotron stress pass did not just grade the model. It also helped expose hardening targets in the workflow layer itself.

Current external-orchestrator summary for Darwin.
Current external-orchestrator summary: GPT-5.4 is clean on the expanded orchestration matrix, while Nemotron remains useful as a portability and hardening probe that exposes weak families in phenotype-first planning, import-first chaining, and decision-boundary synthesis.

What This Update Says About Darwin Today

After the past week of work, the evidence is materially stronger in four different ways.

1. The atlas data surface is broader

The project now spans a larger multi-species network surface, including newer crop and ornamental-supporting builds beyond the original five-species emphasis.

2. The inferred and projected regulatory layers are benchmarked directly

We are not only asserting that the network is biologically useful. We refreshed the gold-standard quality reports, reran independent AUROC/AUPRC benchmarks, and regenerated the population-level network validation reports.

3. The workflow layer is substantially deeper

Import-first omics workflows, chromatin-aware follow-up, calibration surfaces, CRISPR comparison, onboarding readiness, and collaborator-facing packaging are now part of the validated surface.

4. The LLM testing is now a real stress surface

The system is no longer being tested only on isolated prompts. It is being tested on direct skill execution, natural-language skill routing coverage, full GPT-5.4 orchestration, and paced Nemotron orchestration on the harder chain families.

Practical Bottom Line

The most defensible current public statement is:

For academic and non-commercial use, the project is publicly available now.

For commercial use, productization, hosted deployment, or partnership discussions, contact CAIRN Institute.

Source Documents Behind This Update

This update is grounded in the current in-repo validation and testing artifacts, including:

Closing

The point of this week’s work was not just to add features.

It was to make the platform more defensible.

That means better data, more explicit validation, stronger assay-oriented workflows, broader direct testing, and clearer understanding of which LLM-driven workflow families are robust today and which still need hardening.

That is the standard we want Darwin to meet as a research platform.