Delta Voice API
Voice cloning, mastered speech, long-form jobs, local transcription and streamed AI conversations through one owner-protected HTTP API.
Authentication
Send Authorization: Bearer tts_owner_…. The raw token is never stored in SQLite. Every private query is scoped to its SHA-256 owner hash.
Unlimited usage
There are no per-user hourly or daily quotas for cloning, generation, transcription, calls, or API requests. Technical validation, maximum file sizes, bounded job sizes and the global GPU queue remain active to protect service stability.
Voices & versions
Every identity automatically receives the immutable built-in voice builtin_minecraft_villager. It is an unofficial original block-villager-style voice and does not bundle Mojang audio or third-party RVC weights. Use this ID anywhere a Voice ID is accepted.
curl -X POST https://tts.deltaservers.nl/api/v1/voices \
-H "Authorization: Bearer $TOKEN" \
-F "name=My Voice" -F "language=English" -F "consent=true" \
-F "enhance=true" -F "reference_text=Exact transcript" \
-F "reference_audio=@reference.wav"GET /api/v1/voicesGET /api/v1/voices/{voice_id}GET /api/v1/voices/{voice_id}/referenceGET /api/v1/studio/voices/{voice_id}/versionsPOST /api/v1/studio/voices/{voice_id}/versionsPOST /api/v1/studio/voices/{voice_id}/activate/{version_id}
Generate & master speech
curl -X POST https://tts.deltaservers.nl/api/v1/audio/speech \
-H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/json" \
-d '{"voice":"voice_xxx","text":"Hello world","format":"mp3","style":"broadcast"}' \
--output speech.mp3POST /api/v1/generate stores audio and returns JSON. POST /api/v1/audio/speech returns bytes directly. Optional fields include style, speed, pitch, volume, normalize, trim_silence, fade_ms, project_id, and subtitles. Supported styles are neutral, warm, calm, energetic, dramatic, whisper, and broadcast.
GET /api/v1/generations/{id}/subtitles?format=srt also accepts vtt. Timing is estimated per sentence from final audio duration.
Studio jobs
curl -X POST https://tts.deltaservers.nl/api/v1/studio/jobs \
-H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/json" \
-d '{"kind":"longform","voice":"voice_xxx","title":"Chapter 1","text":"Long document…","format":"mp3"}'Job kinds are longform, dialogue, and batch. Long-form and dialogue output is merged into one file. Batch output is a ZIP. Poll GET /api/v1/studio/jobs/{id} and download from /api/v1/studio/jobs/{id}/download.
Additional owner-protected resources: /projects, /presets, /pronunciations, /shares, /webhooks, /overview, and /queue under /api/v1/studio.
Transcription
curl -X POST https://tts.deltaservers.nl/api/v1/transcriptions \
-H "Authorization: Bearer $TOKEN" \
-F "audio=@meeting.webm" -F "language=auto" -F "word_timestamps=true"The response contains detected language, full text, segment timestamps and optional word timestamps. TXT, SRT, VTT and JSON export remain available in the web studio.
Voice calls
Create an owned call with POST /api/v1/calls, transcribe a microphone turn, then stream newline-delimited token events from POST /api/v1/chat/stream. Finished phrases are generated through the selected cloned voice and remain interruptible.
Signed webhooks
Create HTTPS-only webhooks through POST /api/v1/studio/webhooks. The secret is returned once. Verify X-Delta-Signature as an HMAC-SHA256 of the raw request body. Event types include generation.finished, generation.failed, job.finished, and job.failed.
Errors
JSON errors use {success:false,error:{code,message}}. Relevant codes include INVALID_TOKEN, INVALID_AUDIO, TEXT_TOO_LONG, QUEUE_FULL, MODEL_UNAVAILABLE, VOICE_NOT_FOUND, PROJECT_NOT_FOUND, and JOB_NOT_FOUND.