Beyond the music recording.
Sound effects and spoken-word assets for custom audio datasets. Start with the sounds, language and descriptive detail your task requires.
01 / Content
Different sounds need different briefs.
Confirm the available sounds, languages and source specifications for your selected dataset.
Sound effects
Define the sounds and descriptions your model needs, then confirm duration, channels and other file requirements.
Spoken word
Specify language and vocal characteristics alongside the intended use. Confirm transcript coverage; text is not available for every asset.
02 / Your requirements
Make the selection criteria concrete.
Content
Sound effects, spoken word or a combined selection. Define the specific sounds and language requirements.
Annotations
Descriptions, keywords, voice or language fields, and text where available. Specify the required source of each label.
Files
Confirm format, sample rate, bit depth, channels and duration requirements for the selection.
Use + delivery
Define the model use, applicable license and preferred delivery method alongside the content brief.
03 / Delivery
Audio and context, delivered together.
- Source baseline
- WAV / FLAC; typically 44.1 / 48 kHz, 16 / 24-bit
- Metadata records
- Master CSV and/or JSON
- Transfer
- Cloud bucket, secure download or agreed enterprise transfer
- Coverage
- Confirmed for the chosen content and fields
Confirm the text you need.
If the task requires transcripts, include that requirement in the brief. Confirm available text and its source for the selected content.
Explore the music example in Data Lab to see how audio and source metadata connect.
Inspect the music exampleWork with ERISV
Describe the audio task.
Share the sound or speech content you need, the languages involved and the annotations your team can use.