Beyond the music recording.

Sound effects and spoken-word assets for custom audio datasets. Start with the sounds, language and descriptive detail your task requires.

5MCombined SFX + spoken-word assets

01 / Content

Different sounds need different briefs.

Confirm the available sounds, languages and source specifications for your selected dataset.

Sound effects

Define the sounds and descriptions your model needs, then confirm duration, channels and other file requirements.

Spoken word

Specify language and vocal characteristics alongside the intended use. Confirm transcript coverage; text is not available for every asset.

02 / Your requirements

Make the selection criteria concrete.

  1. Content

    Sound effects, spoken word or a combined selection. Define the specific sounds and language requirements.

  2. Annotations

    Descriptions, keywords, voice or language fields, and text where available. Specify the required source of each label.

  3. Files

    Confirm format, sample rate, bit depth, channels and duration requirements for the selection.

  4. Use + delivery

    Define the model use, applicable license and preferred delivery method alongside the content brief.

03 / Delivery

Audio and context, delivered together.

Source baseline
WAV / FLAC; typically 44.1 / 48 kHz, 16 / 24-bit
Metadata records
Master CSV and/or JSON
Transfer
Cloud bucket, secure download or agreed enterprise transfer
Coverage
Confirmed for the chosen content and fields

Confirm the text you need.

If the task requires transcripts, include that requirement in the brief. Confirm available text and its source for the selected content.

Explore the music example in Data Lab to see how audio and source metadata connect.

Inspect the music example

Work with ERISV

Describe the audio task.

Share the sound or speech content you need, the languages involved and the annotations your team can use.