Data / Custom datasets

The right audio. In the right structure.

Music, stems, sound effects and spoken word, selected and prepared for your model’s requirements.

Total music titles
13M+
Total stem assets
52M+
SFX + spoken word
5M
Languages
50+

Catalog-wide inventory. Each dataset is selected to match the agreed brief.

01 / Content

Sound takes different forms.

Complete recordings, individual parts, and sound selected for your model.

Music recordings

Instrumental + vocal

Genre · language · mood · instrumentation

Multitrack stems

The individual parts.

Availability varies by catalog.

Explore music data

Original multitracks in select catalogs; otherwise source-separated or derived. Not every title includes stems.

Explore stem data

Sound effects + spoken word

Sound libraries. Voice collections.

Language and technical coverage confirmed per dataset.Explore audio data

MIDI

Performance, in notes.

Original MIDI is available for select content only.

02 / Metadata

Sound, with context.

Human and source fields stay distinct from machine enrichment. A clear picture of the sound, with the origin of every annotation in view.

Inside the metadataDemo 1
Inspect in Data Lab

Master recording · 7 related stems

Stem origin is not disclosed.

Musical structure

Tempo
129 BPM
Key
D minor
Genres
Hip Hop/Rap / Rock

Human / source

Mood + meaning

Mood
confident / energetic
Energy
high
Keywords
intense / hard / hype / anthem

Human / source

Voice + language

Vocal
male vocal

Human / source

Production + sound

Instrumentation
vocals / bass / choir / drums / fx / guitar / strings / percussion / synth

Human / source

Actual fields from the supplied record. Coverage varies by asset and catalog.
Instrumentation describes the sound; it does not imply available stems.

Explore metadata fields
Across the catalog

Source-led.
Enriched to your brief.

File properties and asset IDs stay separate from annotations. Field coverage and enrichment follow your brief.

Human / source

~20

Source fields as a baseline.
Coverage varies by catalog.

Machine-derived

up to ~400

Enriched attributes, where available or added after ingestion.

Descriptor vocabulary
10,000+
Subgenre taxonomy
~3,000

Vocabulary size is not field coverage. Selected attributes and their origins are defined for each dataset.

03 / Preparation + delivery

Prepared for your workflow.

The audio, metadata and relationship records are prepared together, with a delivery method that fits your workflow.

  1. Define the corpus

    Content, languages, scale and intended model use.

  2. Agree the package

    Available assets, specifications, metadata and license.

  3. Receive the delivery

    Audio + CSV/JSON metadata. Updates scoped where available.

The dataset structure

WAV / FLAC

Audio

CSV / JSON

Metadata

Relationships

Asset IDs connect recordings, stems and records

Final contents and license scope are agreed for each dataset.

Typical source specifications
Source audio
WAV / FLAC
Sample rate
44.1 / 48 kHz
Bit depth
16 / 24-bit
Metadata
CSV / JSON

Generally stereo source audio. Final specifications and metadata coverage are agreed for each dataset.

Your delivery path.

Cloud bucket, customer portal, secure download or enterprise transfer — agreed with your team.

Applicable rights and provenance documentation can accompany the agreed dataset.

Understand the licensing

Work with ERISV

Tell us what you’re building.

Bring a model use and a rough dataset brief. We’ll help define the content, structure and delivery.

Contact Sales