NanoASR API
Offline speech recognition. This page describes the endpoints this server has actually mounted, and nothing it does not serve.
- version v1.1.1
- auth apikey
- upload limit 1024 MiB
- dialect openai
- dialect native
- dialect era
- dialect realtime
Authentication
Every request needs an API key, as a bearer token:
curl -H "Authorization: Bearer $NANOASR_KEY" \
-F file=@call.wav -F model=gigaam-v3-ctc-punct-ru \
http://localhost:8080/v1/audio/transcriptions
A browser cannot set a header when it opens a WebSocket, so
/v1/realtime also accepts the key as a subprotocol:
openai-insecure-api-key.<key>. Any script on the page can read a
key the page holds, which is what "insecure" refers to; it is still a real key and
is checked like one.
Served without a credential:
/healthz, /readyz, /docs, /, /ui.
Streaming model
Realtime recognition is served by one model, loaded for the life of the server. A session cannot choose another one.
- model t-one-ctc-ru@2025-09-08
- rate 8000 Hz
- language ru
- punctuation no
- word timestamps yes
- sessions 0 of 32
This model writes neither punctuation nor capitals — no streaming Russian model does. The offline endpoints still do.
OpenAI audio API
Drop-in for /v1/audio/transcriptions, so an OpenAI SDK works by changing base_url and nothing else.
Transcribe an uploaded file.
multipart/form-data with file, and optionally model, language, response_format (json, verbose_json, text, srt, vtt) and repeated timestamp_granularities[] (word, segment). prompt is applied as a comma-separated hotword list rather than as an LM prompt, and temperature is accepted and ignored because the decoders do not sample; both are reported back as warnings. Asking for word timings from a model that cannot produce them yields segment timings and a warning, or 422 with X-NanoASR-Strict: 1.
Not implemented: this build ships transcription models only.
Answers 501. Translation needs a model that writes a language it did not hear.
List the models that can be passed as `model`.
Transcription models only: the supporting models (VAD, punctuation, diarization) and the streaming models are left out, because passing one here would fail. Each entry carries the NanoASR state, languages and whether it produces word timestamps.
Describe one model.
NanoASR API
Everything the server can do, in its own shapes: queued jobs with progress, word timings, diarization, and model administration. Errors are problem+json (RFC 9457).
Transcribe an uploaded file and wait for the result.
multipart/form-data with file, plus model, language, channel_mode, diarize, num_speakers, punctuate, itn, hotwords, decoding_method and word_timestamps. Cancelling the request stops the decode between batches rather than at the end of the file.
Queue the same work and return a job.
Takes the same fields as transcribe, plus webhook_url for a signed delivery on completion. The upload is held on disk until the job reaches a terminal state and is then deleted.
List this key's jobs.
Scoped to the calling key unless the key is administrative. Paged by cursor.
Fetch one job and its result.
Follow a job as server-sent events.
text/event-stream of numbered job states. Last-Event-ID resumes rather than replays. A job that has already finished yields one catch-up event and closes.
Cancel a queued or running job.
A diarization pass already under way cannot be interrupted and runs to its end.
List installed models and what is resident.
List the models available for download.
Download a catalog model, streaming progress as SSE.
Load a model into memory now.
Release a model.
Keep a model resident, or stop keeping it.
Swap in another revision without dropping requests.
The new instance is loaded and warmed before the pointer moves.
The effective configuration, with secrets redacted.
whisper-asr-webservice contract
For clients written against ahmetoner/whisper-asr-webservice. It mounts at the root, so it is enabled deliberately rather than by default.
Transcribe and return the transcript as a file.
multipart with audio_file, plus output (txt, json, srt, vtt, tsv), task, language, word_timestamps and encode. task=translate is refused rather than answered with a transcription.
Report the language of the audio.
Answered from the model's declared languages, not from an acoustic language identifier: these are monolingual models, so the honest answer is what the model was trained on.
Queue the same work and return a task id.
Poll a queued task.
OpenAI realtime transcription (websocket)
Recognition while the audio is still arriving, over a websocket, in the event protocol of OpenAI's realtime API with intent=transcription. Needs a streaming model: realtime.model in the configuration.
Open a streaming recognition session (websocket upgrade).
Client events: session.update (or transcription_session.update), input_audio_buffer.append with base64 audio, .commit and .clear. A binary frame is accepted as raw audio in the declared format, which saves base64's third of overhead. Server events: transcription_session.created and .updated, input_audio_buffer.speech_started, .speech_stopped, .committed, .cleared, conversation.item.created, conversation.item.input_audio_transcription.delta and .completed, error, and nanoasr.warning for a parameter that was accepted and ignored. Every delta also carries the full hypothesis in nanoasr.text, because a streaming decoder can retract a word it already sent and no sequence of deltas expresses that. input_audio_format takes pcm16 (24 kHz by default), g711_ulaw, g711_alaw, or an object {"type":"audio/pcm","rate":16000}. The model, the decoding method and the endpoint timings belong to the server and cannot be chosen per session; prompt is not applied. Authenticate with Authorization: Bearer, or from a browser with the subprotocol openai-insecure-api-key.<key>.
Not implemented: this server mints no ephemeral tokens.
Answers 501 and says how to connect with an ordinary API key instead.
Server
Endpoints that belong to no dialect.
Liveness, with the native library versions.
Readiness, and the queue depth.
503 when the queue is full, which is when a load balancer should stop sending work here.
This page.
The same endpoint list, as OpenAPI.
Redirects to /docs.
The test UI.
Served without a credential because a browser does not send a bearer token when it loads a script tag. The SPA discovers whether the API needs a key from the first 401 it gets.