API Documentation

Contract 1.0 · base URL https://api.agileascolto.agile.software · changes are listed in What's new and in GET /health.

Authentication

Every request carries the key in the Authorization: Bearer aga_… header (or X-API-Key). In the WebSocket, when the client cannot send headers (browsers), use the agileascolto, key.aga_… subprotocol. Never put the key in the URL.

Every key carries one or more scopes — transcribe (files), stream (WebSocket), batch (asynchronous jobs), metrics (measures and overview), test (extra scope, never alone: marks a trial key, covers no route) — and can have an optional expiry: beyond it, 403 KEY_EXPIRED. To try the API without onboarding there is a passe-partout trial key, time-limited and with a low quota.

Models

NameWhat it is
agileascolto-turboGeneral Italian model, the default. Streaming and files.
agileascolto-discoveryModel under evaluation for the telephone (names, numbers, codes). Enabled per key, on request.
whisper-1Accepted only as a compatibility alias for the default model, for clients that send it out of habit.

GET /v1/models lists the models your key can use.

GET /v1/codici/blocchi?tipo=cf&sentito=<code heard before> gives the blocks of that type with the phrase to say and, for each one, posizionabile: a block can be spliced only if the heard code is long enough, and asking for one that cannot be spliced throws a turn away. Without sentito you get just the table of blocks.

File transcription — POST /v1/audio/transcriptions

Compatible with the OpenAI API. multipart/form-data body: file (wav, mp3, m4a, ogg, webm, flac… up to 200 MB and 60 minutes), model, language (default it), response_format = json | text | verbose_json | srt | vtt (subtitles: at most two lines of 42 characters and 7 seconds per caption; long sentences are split at word boundaries). words=1 (or timestamp_granularities[]=word, as in OpenAI): words with start and end in verbose_json, and subtitles cut at the pauses between words. prompt and temperature are accepted and ignored. codici=1: dictated codes (IBAN, POD, PDR, tax code, case numbers read letter by letter) come out packed into a single token: «i ti otto uno di zero uno sette zero…» → IT81D0170…, «bi esse trattino uno cinque…» → BS-15767. The rest of the sentence is unchanged. The response also carries codici: the tokens found with their type, valido (mod 97 checksum for the IBAN, check character for the tax code, shape for the others) and motivo; if the code does not hold and one single repair would make it hold, also riparato and regola — valido stays false: it is a proposal to be confirmed, not an accepted code. atteso=cf (with codici=1): the type of code you asked for (iban|cf|pod|pdr|telefono|pratica) — the candidate that answers your question carries atteso: true, and if the ear dropped a character and no candidate has that shape any more, the closest one is adopted and re-checked as that type (adottato: true), with a reason you can say out loud. A valid candidate is never adopted and adoption makes nothing valid. ricomponi=<code heard before> + blocco=cognome (with codici=1 and atteso): when the code does not hold and there is no repair, ask for one block instead of the whole code (codice fiscale: cognome, nome, nascita, comune, controllo; IBAN: paese, banca, sportello, conto; POD and PDR: testa, coda) and send the answering turn with those two fields: you get ricomposto back, with the code put together again and re-checked. giudicabile says whether that type's check has a checksum and can therefore reject a wrong recomposition (IBAN and codice fiscale yes, POD and PDR no: there valido: true only means «the shape holds» and the code must be read back to the person). You pick the block: the frontend does not know where the error is.

from openai import OpenAI
c = OpenAI(base_url="https://api.agileascolto.agile.software/v1", api_key="aga_…")
r = c.audio.transcriptions.create(model="agileascolto-turbo", file=open("chiamata.wav", "rb"), language="it")
print(r.text)
curl https://api.agileascolto.agile.software/v1/audio/transcriptions \
  -H "Authorization: Bearer aga_…" -F file=@chiamata.wav -F model=agileascolto-turbo -F language=it
{"text": "…"}

With verbose_json: {"task","language","duration","text","segments":[{"id","start","end","text"}],"model","processing_ms"}. Segments are the sentences as the engine closed them, with start and end in seconds. With words=1 there is also "words":[{"word","start","end"}], at the top (as in OpenAI) and inside each segment. For long jobs (scope batch), POST /v1/audio/jobs accepts the file and GET /v1/audio/jobs/{id}?format=srt (or vtt, text) delivers the finished file.

Streaming — WS /v1/audio/stream?model=…

The client opens the WebSocket, sends a start message, then 16-bit mono PCM audio in 20 ms frames, and receives partials and finals while speaking.

DirectionMessageMeaning
client → server{"type":"start","language":"it","sample_rate":16000}Start. sample_rate 8000 (telephone) or 16000. Optional "codici":true (or ?codici=1 in the URL): final messages come out with dictated codes packed, partial ones do not.
server → client{"type":"started"}Ready to receive audio.
client → serverbinary PCM s16le mono frames20 ms per frame (640 bytes at 16 kHz, 320 at 8 kHz), at real-time rate.
server → client{"type":"partial","text":"…"}Provisional text of the sentence in progress.
server → client{"type":"final","text":"…","speech_final":true}Closed sentence. speech_final true = the person has finished speaking.
client → server{"type":"flush"}Close the current sentence immediately (turn end decided by the client); the session stays open.
client → server{"type":"stop"}End of session: the last final arrives, then {"type":"closed"}.
server → client{"type":"error","code":"…","status":429,"message":"…"}Frontend error, then closed with code 1008.
server → client{"type":"closing","code":"SHUTTING_DOWN","retry_after_s":10}The frontend is shutting down (a redeploy): the final of what you have already said follows immediately, then the close with code 1012 ("service restart"). Reopen the session after retry_after_s: WebSocket libraries retry on 1012 by themselves.
import asyncio, json, websockets
async def main():
    async with websockets.connect("wss://api.agileascolto.agile.software/v1/audio/stream?model=agileascolto-turbo",
                                  additional_headers={"Authorization": "Bearer aga_…"}) as ws:
        await ws.send(json.dumps({"type": "start", "language": "it", "sample_rate": 16000}))
        # … await ws.send(frame_pcm) ogni 20 ms …
        await ws.send(json.dumps({"type": "stop"}))
        async for m in ws:
            print(m)
asyncio.run(main())

Usage — GET /v1/usage

A key sees only its own usage: by month, operation (transcribe, stream, batch) and model, with requests, seconds of audio and milliseconds of processing. The used_s_this_month field is compared with quota_s_month. Without writing code: the usage page shows the same numbers when you paste the key (which travels only in the header, never in the address).

Overview — GET /v1/quadro

For those who monitor (a key with the metrics scope): the keys with requests, seconds and the month's quota percentage (?mese=AAAA-MM, default the current month), totals, and engine status. Numbers only: never secrets, never texts. The same overview is read on the server with agileascolto-keys.py quadro and, with a metrics key, on the usage page. The perimeter is the key's and it is written in the reply (perimetro): a key with a tenant sees its own company, a control-room key (no tenant) the whole service. The /metrics counters are the service's and cannot be narrowed to one company: to a key with a tenant they answer 403 METRICS_TENANT_FORBIDDEN.

Status — GET /health

No key required. Reports whether the frontend and the engines respond, the contract version, active streams, and the link to what's new.

Self-service console — /v1/console

Routes apart from the transcription contract: they are not authenticated with an aga_… key but with a user token agt_… of your tenant (Authorization: Bearer agt_…). They let a company's administrator (or an individual person) manage their own keys and their own people alone, without asking us for anything. You see and touch only your own tenant (403 TENANT_FORBIDDEN); every operation lands in the log, never with a key's secret nor with an invite token.

RouteWhat it does
POST /v1/console/keysIssues a key of your tenant. The secret is shown once only, here. 403 PLAN_EXCEEDED if the key alone would promise more than the company plan. Self-service scopes are transcribe, stream and batch: metrics and test are issued by the control room (403 SCOPE_NOT_SELF_SERVICE).
GET /v1/console/keysYour keys with status and the month's usage, plus the company plan and its seconds. Each one says whose it is: owner_email and owner_stato next to owner (an active key of a suspended person is visible at a glance); ?owner= narrows to a person (id or email) or to a product. Never the secret, never a person's token.
POST /v1/console/keys/{key_id}/rotateA new key with the same rights; grazia_ore keeps the old one alive for the grace period.
POST /v1/console/keys/{key_id}/revokeImmediate revocation.
POST /v1/console/keys/revoke-by-ownerRemoves in one move every active key of a person or of a product (owner: id, email or product name) — the gesture for whoever left the company. No grace period: a revocation hands out no new secret, so whoever wants the integration to keep running rotates the key instead. 404 OWNER_NOT_FOUND if that name has no key.
POST /v1/console/usersInvites a person into your tenant (email, ruolo, scadenza_giorni). The invite token is shown once only, here, and it expires: 7 days by default, 90 at most (expires_at in the reply). 409 USER_EXISTS if they are already one of your people.
POST /v1/console/users/acceptWhoever received an invite accepts it with their own token: you cannot accept on someone else's behalf.
GET /v1/console/usersYour people with role (amministratore, utente) and status (invitato, attivo, sospeso).
POST /v1/console/users/{utente_id}/suspendSuspends a person: their token stops entering the console. Their keys are never touched on their own: chiavi in the body is the choice — "lascia" (default: they stay, and the log says which) or "revoca" (their active keys go in the same move). 409 LAST_ADMIN on the last active administrator: your company cannot lock itself out.
POST /v1/console/users/{utente_id}/reactivateReactivates whoever comes back.
PUT /v1/console/users/{utente_id}/roleChanges the role. Inviting, suspending and assigning roles belongs to the administrator (403 ADMIN_REQUIRED); the list is read by anyone in the tenant.
POST /v1/console/users/{utente_id}/rotate-tokenA new token for a person, identity unchanged (same id, same role, same keys, same history in the log): shown once only, here. You change your own, even as a plain user; your people's is changed by an administrator. By default the old one dies at once (grazia_ore: 0) and then gives 403 TOKEN_ROTATED. 409 USER_NOT_ACTIVE if the person is suspended.
GET /v1/console/auditOperations log of your tenant: who, what, on which key or on which person, with what outcome. A row about one of your keys is there even when the control room wrote it, and it says whose the actor is. Filters key_id, persona (id or email) and operazione; limite from 1 to 1000 (100 by default), beyond that it is 400 AUDIT_LIMIT_INVALID — no silent truncation — and the reply says quante rows match and whether it is troncato. Never a token, never a fingerprint, never an email.
GET /v1/console/tenants/{tenant}/pianoThe company plan (agreed requests per minute and seconds per month) and its seconds for the month.

The company plan — the tenant's ceiling

rate_per_min and quota_s_month live on the key, and keys are issued by the administrator from the console: they are your own choice, not the contract. The contract is the company plan (the tenant plan): requests per minute and seconds per month for the whole company, summed over all its keys — revoked ones included, for the month's seconds.

Only the control room writes the plan (PUT /v1/console/tenants/{tenant}/piano, header X-Regia-Superadmin): from the console a company cannot widen its own ceiling, because that is the agreed part. It reads it whenever it wants: the administrator in GET /v1/console/keys (piano, tenant_used_s_mese) and in GET /v1/console/tenants/{tenant}/piano, whoever holds a key in GET /v1/usage (piano_tenant, tenant_used_s_this_month).

It bites in two places. When a key is created that on its own would promise more than the plan: 403 PLAN_EXCEEDED, and that holds for an unlimited key too. And on every request, over the sum of all the company's keys: 429 TENANT_RATE_LIMIT and 429 TENANT_QUOTA_EXCEEDED. These two say something different from their namesakes on the key: RATE_LIMIT and QUOTA_EXCEEDED are about that key (the administrator can issue another one or widen it), the TENANT_* ones are about the company, and one more key changes nothing. The per-minute rate is a sliding window (the last 60 seconds), not the wall-clock minute.

0 means "no ceiling", as for keys, and a tenant with no written plan gets the deployment's starting one: 0 and 0 by default, so anyone not using plans sees no change in behaviour.

Why it exists, measured: with the plan "6 requests per minute, 6 seconds per month" and an administrator issuing himself 3 keys in good faith from the console, without the plan the company passed 18 requests per minute against a ceiling of 6 (3.0x) and 18 seconds in the month against a quota of 6 (3.0x); with the plan, 6 and 6 (1.0x). It is the measure from our own test bench, not a capacity promise.

Requesting a licence — POST /v1/richieste-licenza

Public route, no key needed: it is asked by whoever does not have one yet. JSON body with nome, email and uso required, azienda, piano (prova, standard or enterprise; absent means prova) and volume_stimato optional. The sito_web field must be left empty: it is the anti-robot trap.

201 answer with the number to quote when writing to us: {"numero":7,"id":"req_1a2b3c4d","stato":"registrata","messaggio":"richiesta registrata n. 7"}.

No request is swallowed: the row is written before any delivery is attempted, and once it exists the answer is 201 whatever happens next. A run on our box takes it to whoever must read it and marks it done only on successful delivery: if the channel is down, the request stays and the next run retries.

The refusals have names: 400 LICENSE_REQUEST_INVALID (it says which field), 400 LICENSE_REQUEST_REJECTED (the trap was filled: no row written), 413 LICENSE_REQUEST_TOO_LARGE (body over 4096 bytes, checked while reading), 429 RATE_LIMIT with Retry-After (at most 5 per minute per address).

Privacy: we keep only the SHA-256 of the address, and our log holds number, id and plan — never the name, never the email, never the text of what you wrote. The list of requests (GET /v1/richieste-licenza) is ours only, and without the control room token configured it is closed even to us (403 SUPERADMIN_REQUIRED).

Errors

StatusCodeWhen
401KEY_MISSING, KEY_INVALIDKey missing or unknown.
403KEY_REVOKED, KEY_EXPIRED, SCOPE_NOT_ALLOWED, MODEL_NOT_ALLOWEDKey revoked, expired, or it does not cover the operation or the model.
403INVITE_EXPIRED, TOKEN_ROTATEDIn the console: an unaccepted invite past its expiry, or a token withdrawn by a rotation. The secret was recognised — ask for a new one: this is not a 401.
404MODEL_UNKNOWNNonexistent model name.
400AUDIO_INVALID, FORMAT_UNSUPPORTEDFile not decodable or too short; response_format (or job format) not supported.
501TIMESTAMPS_UNAVAILABLESubtitles requested from a model whose engine does not provide segment times: they are not made up.
413AUDIO_TOO_LONG, SESSION_TOO_LONG, AUDIO_TOO_LARGEOver 60 minutes per file or 4 hours per session; AUDIO_TOO_LARGE = body above 200 MB.
429RATE_LIMIT, QUOTA_EXCEEDED, TENANT_RATE_LIMIT, TENANT_QUOTA_EXCEEDEDToo many requests per minute, or the month's seconds are used up (even in the middle of a stream). The TENANT_* ones are the whole company's ceiling (the tenant plan), not this key's: issuing another key does not help.
503ENGINE_DOWN, CAPACITY, SHUTTING_DOWNSHUTTING_DOWN = the frontend is shutting down (a redeploy): not a fault and not a full frontend, it is back in a few seconds, and whoever was already inside receives their final first. Engine unreachable, or too many parallel streams, or the jobs queue is full — by count or by seconds of audio in RAM (the message says how many are free), or too many requests holding audio in flight at once (a request waits its turn, then 503). There is never a silent fallback.
502ENGINE_ERRORThe engine rejected the request.

Error body: {"error":{"code":"…","message":"…"}}.

The quota (and the company plan) also counts the audio that has come in and is not booked yet: an open call, a file being processed, a queued job. Three lines together do not get three times the quota, and GET /v1/usage reports those seconds in in_corso_s and tenant_in_corso_s. The audio block that is refused never reaches the engine and is not billed.

When retrying makes sense, the refusal also says when: the 429s and the 503 CAPACITY carry the Retry-After header (whole seconds) and the same number in error.retry_after_s — on the WebSocket, where headers do not exist, only the field. On the per-minute ceiling it is the instant a slot frees up, on the quota it is the start of next month, on a full frontend it is a hint of a few seconds, on a redeploy (SHUTTING_DOWN) it is the time to shut down and come back. The other refusals do not carry it: retrying would change nothing.

Privacy

Audio travels in memory and is not written to disk; text is not recorded. Of the traffic, only seconds, requests and times remain, per key. No data leaves the European Union.