Gene Regulation Skills: Now Supporting Claude, Codex, and Nemotron
Explore Darwin →Quick Summary
Darwin is a research system for working with gene regulatory networks: who regulates a gene, what evidence supports an edge, which regulators explain a gene set, what might change after perturbation, and which experiment is most useful next.
This release makes that system available as a structured skill layer for AI agents. Darwin now provides 100 documented skills: 99 callable research skills and one overview/router. The skills can be used through Claude, Codex, and Nemotron model workflows while exposing the atlas through defined tools rather than asking a model to improvise database queries or biological claims from memory.
| Evaluation | What it measures | GPT-5.6 Terra |
|---|---|---|
| Single-skill evaluation | Can the model choose and execute the right tool for one natural-language request? | 386/386 (100%) |
| Orchestration evaluation | Can the model complete a multi-step research workflow, use the required tool chain, and produce the requested conclusion? | 111/111 (100%) |
On the 111-task orchestration suite, Opus 4.6 reached 107/111 (96.4%) and Nemotron-3 Ultra reached 95/111 (85.6%).
Why a Skill Layer
Gene-regulation questions are rarely one lookup. A researcher may begin with a set of differentially expressed genes, ask which transcription factors explain it, check whether the conclusion transfers to a different species, assess evidence for a candidate edge, and then choose between RNAi, CRISPR, or a promoter edit.
Darwin gives the agent specific operations, arguments, and return values for each task. That makes the work inspectable and lets us test the workflow rather than relying on an answer that merely sounds plausible.
What Researchers Can Do
| Workflow area | What the skills support | Representative skills |
|---|---|---|
| Find and inspect | Resolve gene names; retrieve networks, regulons, paths, subgraphs, modules, and network motifs. | gene search, network, regulon, pathfinding |
| Interpret a gene set | Import and normalize lists; identify upstream regulators; run GO, pathway, and trait enrichment; score TF or pathway activity. | dataset import, upstream, enrichment, TF activity |
| Add biological context | Retrieve expression and coexpression; compare tissues or cell types; analyze trajectories; inspect signaling-to-TF relationships. | expression, cell-type regulation, trajectory drivers |
| Evaluate evidence | Audit network, motif, chromatin, enhancer, perturbation, and multi-omic support for a gene or edge. | evidence audit, cis-support audit, multiome audit |
| Work with regulatory DNA | Query promoter motifs, map peaks to genes, inspect genomic context, prioritize promoter edits, and assess variants. | motif query, peak-gene linkage, variant effect |
| Design interventions | Predict perturbation consequences; compare combinations and modalities; design dsRNA or CRISPR guides; assess off-target risk. | perturbation, dsRNA, CRISPR design |
| Compare species | Find orthologs, assess edge conservation, quantify transfer risk, rescue sparse evidence, and assess readiness. | orthology, conservation, transferability |
| Decide and hand off | Rank candidates, identify the smallest defensible validation step, build plans, and preserve provenance. | candidate triage, decision boundary, validation plan |
Single-Skill Testing
The single-skill matrix contains 386 natural-language cases spanning the full documented skill inventory. Each case checks whether the model selected the appropriate tool, supplied required arguments, successfully executed the tool against the atlas, and met task-specific output constraints where relevant.
GPT-5.6 Terra completed 386/386 cases. This is a validated composite result: 384 cases passed in the matrix run, and the final two were replayed after correcting a species-name boundary issue. The dsRNA API and grader now canonicalize standard scientific species synonyms; the dsRNA regression suite passed 10/10.
Orchestration Testing
The orchestration suite contains 111 multi-step questions. They require the agent to select a sequence of skills, carry information from one step to the next, and state the requested conclusion in its final answer.
- Gene discovery followed by network, regulon, evidence, or enrichment analysis.
- Imported omics data followed by differential, cell-state, trajectory, or TF activity analysis.
- Promoter, motif, chromatin, and variant or edit-consequence follow-up.
- RNAi or CRISPR design followed by off-target, perturbation, and pathway analysis.
- Inferred-edge comparison followed by curated-network or module validation.
- Orthology, conservation, transferability, and family-rescue analysis.
- Phenotype-first candidate discovery, literature grounding, prioritization, readiness assessment, and validation planning.
- Decision-boundary, counterfactual, calibration, and collaborator-handoff workflows.
Orchestration Results
| Model | Transport | Result | Interpretation |
|---|---|---|---|
| GPT-5.6 Terra | ChatGPT-authenticated Codex CLI with structured MCP tools | 111/111 (100%) | Validated composite result on the current routing and MCP configuration. |
| Opus 4.6 | Claude CLI with native MCP tools | 107/111 (96.4%) | Four misses were incomplete chains or semantically similar but incorrect tool choices. |
| Nemotron-3 Ultra | OpenRouter | 95/111 (85.6%) | More incomplete multi-step work and final-synthesis errors. |
These differences are actionable. They identify where routing instructions, tool interfaces, or evaluation cases need further hardening.
What These Results Mean
The results support a narrow but important claim: with structured access to Darwin's research tools, models can be evaluated for reliable tool use and workflow completion rather than assessed only by the fluency of their prose.
They do not mean that an LLM has independently proved a regulatory edge, predicted an in vivo phenotype, or replaced a biological experiment. Darwin's outputs should be interpreted as atlas-grounded hypotheses with visible evidence, uncertainty, and follow-up steps.
Practical Bottom Line
Darwin now offers a structured route from a research question to a regulatory hypothesis, evidence audit, intervention plan, or collaborator-ready report across the atlas's supported species and data layers.
The strongest current tool-use result is GPT-5.6 Terra at 386/386 single-skill cases and 111/111 orchestration cases. Opus 4.6 and Nemotron-3 Ultra provide additional evidence that the skill layer is portable across model families, while also showing where longer research chains remain challenging.