← Project DoLittle

Project DoLittle | Data update

A Careful First Pass at an Orca Acoustic Corpus

What it takes to turn public hydrophone archives into an auditable research asset: source-aware retrieval, exact 30-second locality, quality triage, and a conservative listening-cleanup path.

September 10, 2026Marine audioOrcaReproducibility
An orca swimming through deep blue water with subtle teal acoustic waves and data traces.

A usable marine-audio dataset is more than a pile of recordings. It needs to preserve where each sound came from, how it relates to neighboring audio, what its licensing terms permit, and how much uncertainty remains before it is used in a model. Our Orca work is building that foundation first.

26,335non-duplicate, provenance-backed 30-second Orca contexts retained
219.46 hof retained context across 13 source partitions
0contexts automatically tokenized: every selection still carries a quality gate

What We Have So Far

We have completed label-guided or annotation-led Orca context acquisition across DORI, SanctSound, OOI, and several DCLDE collections. Each retained item is an exact 30-second raw context with its parent recording order, time offset, source object or URL, and license partition. Multi-context passes also preserve neighbor relationships so future work can reconstruct chronology rather than treating every clip as an isolated sound.

Horizontal bar chart showing 26,335 retained 30-second Orca contexts by source partition.
Retained Orca contexts by source partition. These are source- and locality-backed candidates, not a claim that every context has passed human review or is ready for tokenization.

DORI-ONC is currently the largest partition, followed by DCLDE DFO-CRP, JASCO/VFPA, DFO-WDLP, and DORI-Orcasound. The differences are informative, but they are not a leaderboard for where Orcas are most common: coverage, annotation practice, recording length, and public access all shape these counts.

Why 30 seconds? It is long enough to preserve local acoustic context and neighboring-call structure, but small enough to retain, audit, and later process reproducibly without downloading full multi-gigabyte archive recordings.

Discovery Is Not Acceptance

Several additional continuous-audio sources have been scanned in bounded pilots. These pilots are deliberately separate from the retained Orca corpus. A detector-positive score window is a lead for review, not a biological conclusion or a dataset row.

Bar chart showing hours scanned in MBARI, Ogasawara, ONC, and Northern Norway pilots.
Continuous-audio discovery pilots are kept outside the Orca context inventory until their detector and human-review gates are resolved.

This distinction matters. MBARI Pacific Sound and Ogasawara pilots, for example, generated many detector score windows, but those windows remain pending source-specific review and threshold calibration. Likewise, inaccessible archive objects are recorded as access outcomes, not as evidence that an animal was absent.

Source-Aware Acoustic Quality Assessment

Hydrophone data is wonderfully varied: quiet water, distant calls, steady machinery, handling noise, tonal artifacts, and shifting ambient conditions can all occupy the same 30 seconds. We built an automated evaluator to measure raw-audio noise burden using temporal variation, band energy, spectral texture, low-frequency hum, persistent tones, impulsive changes, and frame-level noise-floor features.

The evaluator estimates a source-relative risk of moderate_or_worse noise. It does not identify species, prove a vocalization is present, or automatically decide whether a context is suitable for tokenization. Human listening still confirms the meaningful tiers: quiet/soft, moderate, or loud/repetitive noise.

Bar chart of within-source noise-ranking AUC results for four evaluated sources.
The evaluator can meaningfully prioritize review in some sources, but the quality bar for automated acceptance is deliberately higher. No source currently meets the 90% quiet/soft precision requirement for automatic selection.
Automated scores prioritize review. Human labels remain the confirmed evidence for acoustic quality and vocalization audibility.

A Conservative Cleanup Path for Listening

For confirmed quiet/soft candidates, and in a separate experimental partition for confirmed moderate-noise candidates, we now have a working listening-cleanup derivative. It uses the released Earth Species Project Biodenoising DNS48 checkpoint at its native 16 kHz output, applies a narrow 4 kHz Q35 notch to remove a persistent model residual, and peak-normalizes the result to 0.90.

Paired spectrograms of an audited Orca context before and after DNS48 cleanup with a 4 kilohertz notch and peak normalization.
An illustrative audited context. The cleaned derivative suppresses broad noise while retaining visible call structure. It is a listening derivative only: the raw context remains canonical.

We compared stationary spectral subtraction, a source-reference bounded spectral-gain method, DNS48 with different normalization strategies, an optional 4 kHz notch, and hybrid outputs that restored the original recording above 8 kHz. Listening showed that DNS48 could remove substantial noise while leaving audible whale vocalizations. The 4 kHz notch became necessary after peak normalization made a pre-existing narrow model residual noticeable.

The hybrid high-band experiment was useful precisely because it did not deliver a dramatic result. Across 16 paired clips, the restored band above 8.25 kHz was a median 1.10% of total low-plus-high energy and about 19.5 dB below the cleaned low band. We cannot yet tell whether that residual high-band structure is biological, environmental, or instrumental. The cleaner model-only notched output is therefore the working default; the hybrid remains an optional archival-fidelity derivative.

Scope boundary: DNS48 changes the acoustic representation substantially and emits 16 kHz audio. It is not a raw replacement, it is not valid input to our legacy 44.1 kHz Orca detector, and it is not yet a tokenization authorization. Its current role is to create an auditable, research-only listening and preparation derivative.

From Source Recording to a Prepared Derivative

Diagram showing the pipeline from public source and provenance, through 30-second raw context, noise risk triage, DNS48 cleanup, and a separate prepared partition.
The pipeline preserves raw data and provenance at every stage. The evaluator routes review; it does not erase or automatically accept recordings.
  1. Retain exact local context. Source provenance, license, parent recording order, and 30-second offsets travel with the audio.
  2. Measure noise burden. The evaluator scores raw-audio features and produces a source-relative review priority.
  3. Confirm the tier. Human listening distinguishes quiet/soft, moderate, loud/repetitive, no-audible-vocalization, and uncertain outcomes.
  4. Create a separate derivative. Quiet/soft clips move to the working DNS48 cleanup path; moderate clips stay in a separate cleanup-evaluation partition; loud or uncertain clips remain raw-only holds.
  5. Freeze an explicit future selection manifest. Any tokenization experiment must retain raw and derived paths, source/license partition, quality evidence, and the decision rule that selected it.

What Is Next

Next stepWhy it matters
Freeze a fresh source-stratified quality holdoutValidate a combined evaluator without tuning on the final test set.
Expand review in the largest retained partitionsConvert source-relative risk scores into confirmed quality tiers at scale.
Validate a detector for the 16 kHz cleanup representationOur legacy Orca detector expects features through 20 kHz and cannot evaluate DNS48 output fairly.
Keep sweeping accessible, well-mapped sourcesOrcasound archive access, full SanctSound terms, and OOI historical mapping remain real acquisition constraints.
Run a manifest-driven tokenization evaluationStart with confirmed quiet/soft contexts, keep moderate cleanup candidates separate, and preserve every decision.

Why Provenance and Uncertainty Matter

The primary outcome so far is an auditable workflow that preserves the route back to the original recording and records uncertainty around every derivative. This supports reproducibility, source-term compliance, and later independent review of model and data-selection decisions.

Method notes: counts are reconciled from the Orca candidate registry as of September 9, 2026. Denoising is based on a published research-only CC BY-NC DNS48 checkpoint. Charts summarize retained context and bounded discovery pilots; they do not estimate population abundance, calling rate, or full archive coverage.