Skip to main content
Speaker identification, 100+ languages, word-level timestamps. Perfect for depositions, hearings, and interviews.
Endpoint
Upload your audio to a vault, then transcribe with automatic result storage. The transcript is saved back to your vault when complete.
Response

Get Results (Vault Mode)

Response (completed)
Vault Mode Benefits:
  • Transcript automatically saved to your vault
  • No webhook setup required
  • Simpler polling with result_object_id
  • Audio stored securely in your vault

Direct URL Mode

For audio hosted elsewhere, provide a public URL directly.
Response

Parameters

Vault Mode

Direct URL Mode

Shared Options

Get Results (Direct URL Mode)

Response (completed)

Status Values

Processing Times

Examples

Deposition with Speaker Labels (Vault Mode)

Court Recording (Direct URL with Webhook)

Supported Formats

Audio: MP3, M4A, WAV, FLAC, OGG, OPUS, WebM
Video: MP4, WebM, MOV, AVI, MKV (audio track extracted)
Languages: 100+ including English, Spanish, French, German, Chinese, Japanese

Large Vault videos

Vault-backed videos at or above 5 GB are preprocessed automatically. Case.dev creates an internal MP3 derivative, transcribes that audio object, and keeps the original video as the transcript’s source object. The create response may initially return status: "preprocessing" and includes:
  • source_object_id: the original video used for playback
  • input_object_id: the internal audio derivative submitted for transcription
  • result_object_id: the transcript output object
The derivative preserves leading silence and source timing. Case.dev validates the derived audio duration before provider submission and fails the job when synchronization cannot be established. Transcript word and utterance timestamps therefore remain relative to the original video timeline. Subscribe to voice.transcription.* events to receive derivative and terminal status updates without polling.
Pricing: $0.01/minute. A 2-hour deposition costs $1.20.