DELTA VOICE
API REFERENCE · V2

Delta Voice API

Voice cloning, mastered speech, long-form jobs, local transcription and streamed AI conversations through one owner-protected HTTP API.

Authentication

Send Authorization: Bearer tts_owner_…. The raw token is never stored in SQLite. Every private query is scoped to its SHA-256 owner hash.

Unlimited usage

There are no per-user hourly or daily quotas for cloning, generation, transcription, calls, or API requests. Technical validation, maximum file sizes, bounded job sizes and the global GPU queue remain active to protect service stability.

Voices & versions

Every identity automatically receives the immutable built-in voice builtin_minecraft_villager. It is an unofficial original block-villager-style voice and does not bundle Mojang audio or third-party RVC weights. Use this ID anywhere a Voice ID is accepted.

curl -X POST https://tts.deltaservers.nl/api/v1/voices \
  -H "Authorization: Bearer $TOKEN" \
  -F "name=My Voice" -F "language=English" -F "consent=true" \
  -F "enhance=true" -F "reference_text=Exact transcript" \
  -F "reference_audio=@reference.wav"

GET /api/v1/voices
GET /api/v1/voices/{voice_id}
GET /api/v1/voices/{voice_id}/reference
GET /api/v1/studio/voices/{voice_id}/versions
POST /api/v1/studio/voices/{voice_id}/versions
POST /api/v1/studio/voices/{voice_id}/activate/{version_id}

Generate & master speech

curl -X POST https://tts.deltaservers.nl/api/v1/audio/speech \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"voice":"voice_xxx","text":"Hello world","format":"mp3","style":"broadcast"}' \
  --output speech.mp3

POST /api/v1/generate stores audio and returns JSON. POST /api/v1/audio/speech returns bytes directly. Optional fields include style, speed, pitch, volume, normalize, trim_silence, fade_ms, project_id, and subtitles. Supported styles are neutral, warm, calm, energetic, dramatic, whisper, and broadcast.

GET /api/v1/generations/{id}/subtitles?format=srt also accepts vtt. Timing is estimated per sentence from final audio duration.

Studio jobs

curl -X POST https://tts.deltaservers.nl/api/v1/studio/jobs \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"kind":"longform","voice":"voice_xxx","title":"Chapter 1","text":"Long document…","format":"mp3"}'

Job kinds are longform, dialogue, and batch. Long-form and dialogue output is merged into one file. Batch output is a ZIP. Poll GET /api/v1/studio/jobs/{id} and download from /api/v1/studio/jobs/{id}/download.

Additional owner-protected resources: /projects, /presets, /pronunciations, /shares, /webhooks, /overview, and /queue under /api/v1/studio.

Transcription

curl -X POST https://tts.deltaservers.nl/api/v1/transcriptions \
  -H "Authorization: Bearer $TOKEN" \
  -F "audio=@meeting.webm" -F "language=auto" -F "word_timestamps=true"

The response contains detected language, full text, segment timestamps and optional word timestamps. TXT, SRT, VTT and JSON export remain available in the web studio.

Voice calls

Create an owned call with POST /api/v1/calls, transcribe a microphone turn, then stream newline-delimited token events from POST /api/v1/chat/stream. Finished phrases are generated through the selected cloned voice and remain interruptible.

Signed webhooks

Create HTTPS-only webhooks through POST /api/v1/studio/webhooks. The secret is returned once. Verify X-Delta-Signature as an HMAC-SHA256 of the raw request body. Event types include generation.finished, generation.failed, job.finished, and job.failed.

Errors

JSON errors use {success:false,error:{code,message}}. Relevant codes include INVALID_TOKEN, INVALID_AUDIO, TEXT_TOO_LONG, QUEUE_FULL, MODEL_UNAVAILABLE, VOICE_NOT_FOUND, PROJECT_NOT_FOUND, and JOB_NOT_FOUND.

OpenAPI 3.1 JSON →