VoiceStudioDocs

Call VoiceStudio from another app

Discover VoiceStudio 0.5.1 on loopback, transcribe a file, or connect a local speech client.

VoiceStudio 0.5.1 can serve speech tools to an editor, TUI, desktop app, or agent running on the same machine. Start VoiceStudio first, then choose the smallest interface that fits the job.

Pinned to the 0.5.1 release

This guide describes the speech platform published with VoiceStudio 0.5.1. The tagged source remains the authority for this release.

Use native control on port 3902 to start, stop, or toggle VoiceStudio's microphone. Use the speech service on port 3900 to transcribe a file, stream partial and final text, or connect an MCP client.

Discover the running service

Ask the desktop control sidecar which speech interfaces are available:

curl http://127.0.0.1:3902/.well-known/voicestudio-speech

The response should name protocol voicestudio.speech.v1 and include absolute URLs for control, batch transcription, streaming, output sessions, and MCP. The backend has a second discovery document:

curl http://127.0.0.1:3900/.well-known/voicestudio-speech

Use discovery instead of hard-coding every path. It also tells an integration which service version is running.

Transcribe a file

The batch route accepts the same multipart shape as an OpenAI-compatible audio client:

curl -sS http://127.0.0.1:3900/v1/audio/transcriptions \
  -F "file=@recording.wav" \
  -F "model=whisper-1"

This path is useful for scripts, editor commands, and existing SDK clients that already know the OpenAI transcription request shape.

Trigger VoiceStudio's own capture

If VoiceStudio should own the microphone, model choice, dictation pill, and final delivery, call the native control sidecar:

curl -X POST http://127.0.0.1:3902/v1/dictation/start
curl -X POST http://127.0.0.1:3902/v1/dictation/stop
curl -X POST http://127.0.0.1:3902/v1/dictation/toggle

The installed executable also accepts --dictate-start, --dictate-stop, and --dictate-toggle. Its single-instance bridge forwards the command to the running app without opening another Studio window.

For hooks and TUIs, the dependency-free Python client exposes the same work:

python -m backend.speech_client status
python -m backend.speech_client toggle
python -m backend.speech_client transcribe recording.wav

Bring your own microphone interface

Connect an editor extension or custom GUI to:

ws://127.0.0.1:3900/v1/audio/transcriptions/stream

Send binary WebM/Opus frames. Raw signed 16-bit mono PCM is available with ?pcm=1&sr=16000. Send the following control frame to finish without closing the socket:

{"type":"input_audio.end"}

Responses carry the protocol and a session ID. Streaming models may emit partial text and utterance finals before the whole-session summary.

Connect an agent

Modern MCP clients can use Streamable HTTP at http://127.0.0.1:3900/mcp. Stdio-only clients can launch:

python -m backend.mcp_shim

Use native dictation control when the goal is to speak into a focused agent prompt. Add MCP only when the agent also needs file transcription or other speech tools.

Keep the boundary local

The native control sidecar binds only to IPv4 loopback and rejects untrusted browser origins. Ordinary web pages cannot silently turn on the microphone. Loopback clients need no credential.

Remote audio stays behind VoiceStudio's API-key boundary. Restrict remote access to a trusted network and use HTTPS or WSS outside a fully trusted LAN. An API key authenticates the client but does not isolate the network.

For output-session insertion, JSON-RPC, Herdr bindings, and the full transport map, read the tagged speech-platform guide and integration examples.

On this page