Skip to content

Voice Cloning ​

Clone your own voice for TTS with a reference recording, consent, and live verification.

Authentication and safeguards ​

Use Authorization: Bearer <VOICELAB_API_KEY> on every endpoint below. Grant {"voices":"access"} for the complete flow: voices:read reads capabilities, sessions and metadata; voices:write uploads, approves, verifies, edits and deletes. Existing read-only keys do not gain write access. Keys belong on your server, never in browser code. Proxy microphone recordings through your trusted backend. These developer routes do not accept a user JWT. Dashboard /api/v1/account/voice-enrollments routes remain JWT-only.

All active paid plans qualify to issue a challenge and save a clone. Free users can prepare a reference and review its transcript, but cannot complete enrollment. Credit top-ups on a free plan do not unlock cloning. Reading or managing an existing clone does not require a currently active paid plan.

The voice owner must review the reference transcript, explicitly acknowledge consent, and record a fresh server-issued challenge. Both spoken content and speaker similarity must pass before the voice is enabled. The acoustic cutoff is 0.72. This is voice consistency and replay resistance, not proof of legal identity or guaranteed detection of synthetic voices. Never generate challenge audio with TTS, reuse the reference as the live recording, or bypass consent.

Check capabilities ​

GEThttps://api.voicelab.uz/v1/voice-enrollments/capabilities

Check availability, account eligibility and reference limits.

json
{
  "eligible": true,
  "available": true,
  "min_seconds": 10,
  "max_seconds": 30,
  "max_bytes": 10485760
}

available means verification is configured. Later requests can still fail if a provider is unavailable.

Upload a reference ​

POSThttps://api.voicelab.uz/v1/voice-enrollments

Normalize, optionally restore and transcribe one reference recording.

FieldRequiredContract
audioYesExactly one clear, single-speaker audio recording; 10–30 seconds; at most 10 MiB. Browser WebM/M4A and common audio formats are normalized server-side.
languageYesuz, en, or ru. This is fixed for the enrollment.
remove_noiseYesLiteral true or false. true runs Voice Isolator denoising and speech restoration on the reference and live challenge. Provider failure fails closed; no denoising-only fallback.
bash
curl --fail-with-body 'https://api.voicelab.uz/v1/voice-enrollments' \
  -H "Authorization: Bearer ${VOICELAB_API_KEY:?Set your server-side API key}" \
  -F 'audio=@reference.wav' -F 'language=en' -F 'remove_noise=true'

Returns 201 Created after transcribing the reference. Save voice_clone.id. Enrollment responses use this shape:

json
{
  "voice_clone": {
    "id": "vcs_EXAMPLE",
    "status": "transcript_review",
    "language": "en",
    "reference_transcript": "The words spoken in the reference recording.",
    "verification_status": "not_started",
    "verification_attempts": 0,
    "public_requested": false,
    "created_at": "2026-10-01T08:00:00Z",
    "updated_at": "2026-10-01T08:00:00Z"
  },
  "labels": {},
  "orb_palette": "cloud",
  "remove_noise": true,
  "request_id": "req_EXAMPLE"
}

Additional stage-dependent fields include name, description, transcript_approved_at, consent_accepted_at, challenge_text, challenge_language, challenge_expires_at, verification_transcript, verification_method, verified_at, failure_code, voice_id, and top-level score. Missing optional fields are not proof of successful verification. Private storage paths, provider speaker references and account IDs are never returned.

Approve the transcript ​

POSThttps://api.voicelab.uz/v1/voice-enrollments/{id}/transcript

Persist the owner's reviewed and corrected reference transcript.

Have the voice owner review and correct the transcript before submitting it. Send the full text, 1 to 20,000 Unicode characters. Returns 200 with status: "details_pending".

bash
curl --fail-with-body -X POST \
  'https://api.voicelab.uz/v1/voice-enrollments/vcs_REPLACE_ME/transcript' \
  -H "Authorization: Bearer ${VOICELAB_API_KEY:?Set your server-side API key}" \
  -H 'Content-Type: application/json' \
  -d '{"transcript":"The exact corrected words spoken in my reference recording."}'

Issue a live challenge ​

POSThttps://api.voicelab.uz/v1/voice-enrollments/{id}/challenge

Save metadata and consent, and issue a fresh microphone challenge.

json
{
  "name": "My narration voice",
  "description": "A calm voice for stories.",
  "labels": { "language": "en", "gender": "male", "age": "young", "accent": "Standard" },
  "orb_palette": "cloud",
  "consent_accepted": true
}

name is required, 1–100 Unicode characters; description can be empty and is limited to 1,000 characters. labels.language must match the reference language. Optional labels: gender (male, female, neutral), age (young, middle-aged, old), and a nonempty accent up to 80 UTF-8 bytes. Omit a label to remove it; do not send an empty label value. No other labels are accepted. orb_palette must be one of cloud, support, storytelling, narrator, energetic, calm, delicate, steady, deep, or bright. Consent must be explicitly true after the owner has acknowledged their rights. The API cannot change visibility or publish a voice.

Returns 200 with status: "verification_pending". Display voice_clone.challenge_text exactly as returned. New challenges contain reading text only, with no appended numbers or spoken verification code. Comet generates a neutral passage for roughly 12–15 seconds of reading. It receives only the language and a random seed, never the voice or transcript. Each request uses a fresh seed; repeated or nearly identical retry passages are rejected. A hidden server token identifies the single-use attempt. The challenge expires after ten minutes.

Each verification submission consumes its challenge. To retry after a failed attempt or expiry, call this endpoint again with the metadata and consent, then record the new text. The session allows at most five verification attempts. Concurrent submissions cannot share a challenge.

Verify and save ​

POSThttps://api.voicelab.uz/v1/voice-enrollments/{id}/verify

Check live speech content and speaker similarity before enabling a private voice.

Send exactly one audio part containing a fresh microphone recording of the issued text, 3–60 seconds and at most 10 MiB. Do not send expected_text, replacement reference audio, or a client-supplied similarity score: the persisted server challenge and reference are authoritative.

Voice matching runs first and requires a positive match with a finite score of at least 0.72. STT then checks the whole reading, allowing ordinary transcription differences. Both checks must pass; correct text cannot override a failed voice match. Existing unexpired challenges keep their original spoken-code rules. Request a new challenge to use reading text only.

bash
curl --fail-with-body \
  'https://api.voicelab.uz/v1/voice-enrollments/vcs_REPLACE_ME/verify' \
  -H "Authorization: Bearer ${VOICELAB_API_KEY:?Set your server-side API key}" \
  -F 'audio=@fresh-live-challenge.wav'

Success is 200. Require voice_clone.status == "ready", voice_clone.verification_status == "passed" and a nonempty voice_clone.voice_id. score is returned at the top level. Do not use the session ID (vcs_...) as a TTS voice ID. The clone is private and accessible only to the owning account.

Get enrollment state ​

GEThttps://api.voicelab.uz/v1/voice-enrollments/{id}

Read an account-owned enrollment and its persisted challenge/result.

Returns 200 with the shared envelope. Persist the session ID before advancing steps. After a lost response, inspect this state before resubmitting anything. Initialization does not implement Idempotency-Key; blindly repeating an upload may create a second session. Polling does not renew a challenge or restart processing. A timed-out verification can require a new challenge; never replay an old recording automatically.

Use the cloned voice ​

List voices with GET /v1/voices?kind=custom, or use the verified voice_id directly in TTS. TTS requires a separate tts:write permission and normal usage billing applies. A voice ID does not grant another account access. Use these HTTP routes if your SDK has no cloning helper.

Edit a cloned voice ​

GEThttps://api.voicelab.uz/v1/voices/{id}

Read editable metadata for your own enabled cloned voice.

Returns the metadata object directly, without a data wrapper:

json
{ "name": "My narration voice", "description": "Calm", "labels": { "language": "en" }, "orb_palette": "cloud" }

Only your own enabled clones are accessible here. System voices, other users' clones, missing clones and deleted clones return 404. The catalog's can_manage indicates ownership; changes still require write permission.

PATCHhttps://api.voicelab.uz/v1/voices/{id}

Replace editable metadata for your own cloned voice.

Send the complete metadata object above. Success is 204 No Content. Name, description, labels and orb follow the same validation rules as enrollment. Reference audio, language, ownership, provider mapping, consent, verification and visibility cannot be edited. Unknown JSON fields are rejected. Omit optional labels from the complete labels map to remove them.

Delete a cloned voice ​

DELETEhttps://api.voicelab.uz/v1/voices/{id}

Permanently revoke and remove your cloned voice and its private recordings.

Success is 204 No Content; no body is needed. This is destructive: obtain the owner's explicit confirmation first. Deletion immediately revokes synthesis and removes saved links, then removes the private reference/live recordings, generated preview and inference registration. The audit record remains marked as deleted. System voices and other users' clones cannot be deleted. Repeating deletion for the same owned clone is safe. If cleanup fails, the clone stays revoked; retry the same DELETE to finish cleanup rather than uploading a replacement.

Limits and errors ​

Enrollment requests have a three-minute server deadline. Allow slightly longer on your client. Keep API keys out of URLs, audio filenames and logs. Enrollment abuse limits are separate from the general developer API's 70 RPS: initialization permits five/hour with burst two; transcript approval, challenge issuance and verification each permit ten/hour with burst three, per account per API replica. They are shared with dashboard attempts, not multiplied per key. At most five active custom voices are permitted per account.

HTTPerror.codeAction
401 / 403invalid_api_key / insufficient_scopeCheck key, expiry, revocation, IP restrictions and read/write permission.
402voice_clone_subscription_requiredActivate a paid plan before challenge/verification.
400 / 413invalid_multipart, invalid_json, audio_too_largeFix the encoding/body or file size.
404voice_clone_not_foundMissing or inaccessible session/voice; never probe other accounts.
409voice_clone_invalid_state, voice_limit_reachedInspect session state; complete prior steps or remove an owned clone.
410voice_clone_challenge_expiredIssue a fresh challenge and record its new text.
422validation_failed, invalid_audioCorrect metadata or use valid, clear audio within the duration limits.
422challenge_mismatch, speaker_mismatchObtain a new challenge and fresh recording; do not reuse audio.
429voice_clone_rate_limitedRespect Retry-After; do not rotate keys to bypass the account limit.
503voice_cloning_unavailable, speaker_verification_unavailableNo bypass; inspect state and retry with bounded backoff.

See errors and security for the common error envelope. Keep request_id for support, never API keys, private recordings or signed URLs.