MCP server

The Model Context Protocol lets an assistant call tools directly. Our MCP server turns this API into five of them, so someone can say “transcribe this recording and tell me who said what” and have it happen — no code, no copying job ids around.

It ships inside the JavaScript SDK, so there is nothing separate to install.

Setup

Add this to your MCP client's configuration. In Claude Desktop that is claude_desktop_config.json; other clients use the same shape.

json
{
  "mcpServers": {
    "speechrevolutions": {
      "command": "npx",
      "args": ["-y", "speechrevolutions", "mcp"],
      "env": {
        "SPEECHREVOLUTIONS_API_KEY": "stt_..."
      }
    }
  }
}

Restart the client. Create a key at the console if you do not have one — new accounts start with $10 of credit, which is about 55 hours of audio.

The key is read from the environment and never leaves your machine: the server runs locally as a subprocess of your MCP client, and talks to our API the same way the SDK does.

Tools

ToolWhat it does
transcribe_audioTranscribe a local file or public URL and return the text. Waits for the result, so it suits recordings up to about half an hour.
submit_transcription_jobStart a long recording and return a job id immediately. Use for anything longer, and for batches.
check_jobWhether a job is processing, complete or failed.
get_transcriptRead a completed transcript as text, JSON, SRT or VTT.
list_jobsRecent jobs on the account, newest first.

What comes back

Transcripts are returned as readable text with speaker labels and timestamps, not as the raw JSON payload:

text
Job: f1c3ae59-00bc-4e39-aa9e-5afed8c3974a · Detected language: en · Speakers: 3

[0:00] SPEAKER_1: You are already ahead of 90% of the people your age if you
can just communicate competently.
[0:12] SPEAKER_2: That tracks with what we saw in the hiring data.

That is deliberate. The caller is a language model, and word-level JSON for a long recording spends the context it needs to answer the question. Ask get_transcript for format: "json" when you actually want per-word timings and confidences.

Long transcripts are truncated with an explicit notice naming the job id and how much was cut, so nothing is silently summarised as if it were complete. Raise max_characters to get more.

Cost

The same as any other call: $0.003 per minute of audio, with diarization and timestamps included. Tool calls that only check status or list jobs are free.

Troubleshooting

If the server does not appear, run it by hand — it prints its errors to stderr, which some clients hide:

bash
SPEECHREVOLUTIONS_API_KEY=stt_... npx -y speechrevolutions mcp

It should print a readiness line and then wait for input. A message about the key being unset means the env block above did not reach the process.