Gene Regulation Skills: Now Supporting Claude, Codex, and Nemotron

Author: CAIRN Institute

Published: September 29, 2026

Read time: 8–10 minutes

#Genomics #GeneRegulation #Bioinformatics #SystemsBiology #AIforScience #PlantScience

Explore Darwin →

Darwin is powered by the CAIRN Institute GRN Atlas.

Quick Summary

Darwin is a research system for working with gene regulatory networks: who regulates a gene, what evidence supports an edge, which regulators explain a gene set, what might change after perturbation, and which experiment is most useful next.

This release makes that system available as a structured skill layer for AI agents. Darwin now provides 100 documented skills: 99 callable research skills and one overview/router. The skills can be used through Claude, Codex, and Nemotron model workflows while exposing the atlas through defined tools rather than asking a model to improvise database queries or biological claims from memory.

EvaluationWhat it measuresGPT-5.6 Terra
Single-skill evaluationCan the model choose and execute the right tool for one natural-language request?386/386 (100%)
Orchestration evaluationCan the model complete a multi-step research workflow, use the required tool chain, and produce the requested conclusion?111/111 (100%)

On the 111-task orchestration suite, Opus 4.6 reached 107/111 (96.4%) and Nemotron-3 Ultra reached 95/111 (85.6%).

Scope: these are tool-use and workflow-reliability results. They do not establish that every network prediction is biologically correct. Biological validity requires the separate evidence, benchmark, and experimental layers in the atlas.

Why a Skill Layer

Gene-regulation questions are rarely one lookup. A researcher may begin with a set of differentially expressed genes, ask which transcription factors explain it, check whether the conclusion transfers to a different species, assess evidence for a candidate edge, and then choose between RNAi, CRISPR, or a promoter edit.

Darwin gives the agent specific operations, arguments, and return values for each task. That makes the work inspectable and lets us test the workflow rather than relying on an answer that merely sounds plausible.

What Researchers Can Do

Workflow areaWhat the skills supportRepresentative skills
Find and inspectResolve gene names; retrieve networks, regulons, paths, subgraphs, modules, and network motifs.gene search, network, regulon, pathfinding
Interpret a gene setImport and normalize lists; identify upstream regulators; run GO, pathway, and trait enrichment; score TF or pathway activity.dataset import, upstream, enrichment, TF activity
Add biological contextRetrieve expression and coexpression; compare tissues or cell types; analyze trajectories; inspect signaling-to-TF relationships.expression, cell-type regulation, trajectory drivers
Evaluate evidenceAudit network, motif, chromatin, enhancer, perturbation, and multi-omic support for a gene or edge.evidence audit, cis-support audit, multiome audit
Work with regulatory DNAQuery promoter motifs, map peaks to genes, inspect genomic context, prioritize promoter edits, and assess variants.motif query, peak-gene linkage, variant effect
Design interventionsPredict perturbation consequences; compare combinations and modalities; design dsRNA or CRISPR guides; assess off-target risk.perturbation, dsRNA, CRISPR design
Compare speciesFind orthologs, assess edge conservation, quantify transfer risk, rescue sparse evidence, and assess readiness.orthology, conservation, transferability
Decide and hand offRank candidates, identify the smallest defensible validation step, build plans, and preserve provenance.candidate triage, decision boundary, validation plan

Single-Skill Testing

The single-skill matrix contains 386 natural-language cases spanning the full documented skill inventory. Each case checks whether the model selected the appropriate tool, supplied required arguments, successfully executed the tool against the atlas, and met task-specific output constraints where relevant.

GPT-5.6 Terra completed 386/386 cases. This is a validated composite result: 384 cases passed in the matrix run, and the final two were replayed after correcting a species-name boundary issue. The dsRNA API and grader now canonicalize standard scientific species synonyms; the dsRNA regression suite passed 10/10.

Orchestration Testing

The orchestration suite contains 111 multi-step questions. They require the agent to select a sequence of skills, carry information from one step to the next, and state the requested conclusion in its final answer.

Orchestration Results

ModelTransportResultInterpretation
GPT-5.6 TerraChatGPT-authenticated Codex CLI with structured MCP tools111/111 (100%)Validated composite result on the current routing and MCP configuration.
Opus 4.6Claude CLI with native MCP tools107/111 (96.4%)Four misses were incomplete chains or semantically similar but incorrect tool choices.
Nemotron-3 UltraOpenRouter95/111 (85.6%)More incomplete multi-step work and final-synthesis errors.

These differences are actionable. They identify where routing instructions, tool interfaces, or evaluation cases need further hardening.

What These Results Mean

The results support a narrow but important claim: with structured access to Darwin's research tools, models can be evaluated for reliable tool use and workflow completion rather than assessed only by the fluency of their prose.

They do not mean that an LLM has independently proved a regulatory edge, predicted an in vivo phenotype, or replaced a biological experiment. Darwin's outputs should be interpreted as atlas-grounded hypotheses with visible evidence, uncertainty, and follow-up steps.

Practical Bottom Line

Darwin now offers a structured route from a research question to a regulatory hypothesis, evidence audit, intervention plan, or collaborator-ready report across the atlas's supported species and data layers.

The strongest current tool-use result is GPT-5.6 Terra at 386/386 single-skill cases and 111/111 orchestration cases. Opus 4.6 and Nemotron-3 Ultra provide additional evidence that the skill layer is portable across model families, while also showing where longer research chains remain challenging.