Theme
MCP tools reference
The Node MCP server defines 22 tools. This reference follows the tool definitions. Use your connection's tools/list response for the schema deployed on that server.
MCP inputs are a subset of the REST API. Do not pass extra REST parameters unless the tool advertises them. In particular, chat_completions does not advertise streaming or JSON response-format options.
LLM
list_models
List available LLM models with pricing and limits
No input parameters.
chat_completions
Create an LLM chat completion. Returns completion ID and text. Omitted idempotency keys are generated automatically.
| Input | Type | Required | Description |
|---|---|---|---|
model | string | Yes | Model ID from list_models |
messages | array | Yes | Conversation messages |
max_tokens | number | No | Maximum output tokens (1-4096) |
temperature | number | No | Sampling temperature (0-2) |
idempotency_key | string | No | Key saved before the first call; generated if omitted |
Each message requires role and string content. Roles are system, user, assistant, or tool; tool_call_id is optional.
get_llm_request
Get LLM request status, token usage, and credit details
| Input | Type | Required | Description |
|---|---|---|---|
request_id | string | Yes | LLM request ID from chat_completions |
Text-to-Speech
list_tts_languages
List supported TTS languages and model capabilities
No input parameters.
list_voices
List available voices, optionally filtered by language
| Input | Type | Required | Description |
|---|---|---|---|
language | string | No | Optional language code filter (e.g., "en", "uz", "ru") |
text_to_speech
Generate speech as base64-encoded WAV. Omitted idempotency keys are generated automatically.
The linked Node version reads the response as bytes without checking for 202 JSON. Use the REST TTS flow when pending jobs are possible; this tool does not poll them.
| Input | Type | Required | Description |
|---|---|---|---|
text | string | Yes | Text to synthesize (1-5000 characters) |
language | string | Yes | Language code (e.g., "en", "uz", "ru") |
voice_id | string | Yes | Voice ID from list_voices |
speed | number | No | Speech speed (0.5-2.0, default 1) |
idempotency_key | string | No | Key saved before the first call; generated if omitted |
list_tts_generations
List TTS generation history
| Input | Type | Required | Description |
|---|---|---|---|
limit | number | No | Results per page (1-100, default 30) |
cursor | string | No | Pagination cursor from previous response |
get_tts_generation
Get TTS generation details with audio URL
| Input | Type | Required | Description |
|---|---|---|---|
generation_id | string | Yes | Generation ID |
delete_tts_generation
Hide a completed TTS generation from history. This does not cancel pending work or immediately erase retained audio.
| Input | Type | Required | Description |
|---|---|---|---|
generation_id | string | Yes | Generation ID to delete |
Speech-to-Text
speech_to_text
Transcribe audio file. Accepts audio as base64-encoded data. Returns transcription ID for polling. Omitted idempotency keys are generated automatically.
| Input | Type | Required | Description |
|---|---|---|---|
audio_base64 | string | Yes | Base64-encoded audio file (MP3, WAV, M4A, OGG, WebM, FLAC) |
language | string | Yes | Language code (uz, en, ru) |
include_speakers | boolean | No | Enable speaker diarization (default false) |
idempotency_key | string | No | UUID saved before the first call; generated if omitted |
get_transcription
Poll after speech_to_text until completed or failed. Read the transcript only after completion.
| Input | Type | Required | Description |
|---|---|---|---|
transcription_id | string | Yes | Transcription ID from speech_to_text |
list_transcriptions
List STT transcription history
| Input | Type | Required | Description |
|---|---|---|---|
limit | number | No | Use 10, the current REST page size; the tool schema accepts 1 to 100 |
cursor | string | No | Pagination cursor |
update_transcription
Update transcription title
| Input | Type | Required | Description |
|---|---|---|---|
transcription_id | string | Yes | Transcription ID |
title | string | Yes | New title |
delete_transcription
Delete a transcription and its audio (destructive)
| Input | Type | Required | Description |
|---|---|---|---|
transcription_id | string | Yes | Transcription ID to delete |
export_transcription
Export transcription as TXT, JSON, SRT, or VTT. Returns content as base64-encoded data.
| Input | Type | Required | Description |
|---|---|---|---|
transcription_id | string | Yes | Transcription ID |
format | string | Yes | Format: txt, json, srt, vtt. |
Voice Isolator
isolate_voice
Remove background noise from audio with optional speech restoration. Returns job ID for polling. Omitted idempotency keys are generated automatically.
| Input | Type | Required | Description |
|---|---|---|---|
audio_base64 | string | Yes | Base64-encoded audio file (max 300 MiB) |
title | string | No | Optional job title |
speech_restoration | boolean | No | Enable Sidon speech restoration (default false) |
restoration_model | string | No | Restoration model (sidon) |
idempotency_key | string | No | UUID saved before the first call; generated if omitted |
get_isolation
Poll after isolate_voice until processing and MP3 conversion complete. Download when audio_status is completed and audio_url is present. Stop on either failure.
| Input | Type | Required | Description |
|---|---|---|---|
isolation_id | string | Yes | Isolation job ID |
list_isolations
List voice isolation history
| Input | Type | Required | Description |
|---|---|---|---|
limit | number | No | Results per page (1-50, default 20) |
cursor | string | No | Pagination cursor |
create_isolation_export
Create an export of isolated audio in specified format
| Input | Type | Required | Description |
|---|---|---|---|
isolation_id | string | Yes | Isolation job ID |
format | string | Yes | Format: mp3, wav, flac, ogg, original. |
get_isolation_export
Get export status and download URL
| Input | Type | Required | Description |
|---|---|---|---|
isolation_id | string | Yes | Isolation job ID |
format | string | Yes | Export format |
hide_isolation
Hide isolation job from history (does not delete audio)
| Input | Type | Required | Description |
|---|---|---|---|
isolation_id | string | Yes | Isolation job ID to hide |
Realtime
create_realtime_ticket
Create a short-lived WebSocket ticket for realtime TTS or STT
| Input | Type | Required | Description |
|---|---|---|---|
service | string | Yes | Service: tts, stt. |
Results and errors
Most successful tools return JSON in an MCP text content block. text_to_speech returns audio_base64, format, sample_rate, and idempotency_key. export_transcription returns export_base64 and format.
The Node server marks tool failures with isError: true and includes error details in a text content block. Check this flag even if the MCP HTTP request succeeds. Gateway authentication and transport errors can occur before a tool runs.
Creation tools generate an idempotency key when one is omitted. For reliable retries, supply a key before the first request and reuse it with the same input. Do not create another billable request just to poll status.
Limits and realtime
The Node HTTP gateway allows 10 MiB per JSON request by default, including base64 overhead. This is smaller than the REST upload limits. Check the operator's limit before sending large recordings.
Use the websocket_url and expires_at returned by create_realtime_ticket; do not construct a URL or assume a fixed lifetime. Configure the voice or language through the subsequent service protocol. See Realtime TTS and Realtime STT.