What's new
One entry per date, the most recent at the top. Every entry says what changed, who it can be useful to and how to adopt it. The full contract is in the API documentation.
Contract version: 1.0. A model name never changes its behaviour: if the model changes, the name changes.
Full log for integrators: docs/NOVITA-EN.md in the repository, aligned with this page.
30/9/2026The three plans with the price on request, and a licence request that does not get lost.
What changed
- The landing page said «price per minute, being defined» and to ask for a key there was only a
mailto:. Measured: 0 of the three plan names and 0 «on request» on either face, 0 requests recorded — and the landing page was served to nobody (from the seat of whoever opens a browser: 0 out of 1). - Three plans, with the same content in Italian and in English:
prova(trial: scopestranscribeandstream,quota_s_month7 200 s per month — two hours,rate_per_min10 per minute, 30 days),standard(alsobatch, quota and rate tailored, written on the licence at signature) andenterprise(tailored scopes includingmetrics, dedicated installation). Price on request on all three: no figures. - The numbers come from a measurement: the declared capacity of one engine is 1 440 hours of audio per month and the trial plan is 0.14 % of it. The real numbers live on the licence and are set when it is issued: a plan is the template of default values, not a promise of capacity.
POST /v1/richieste-licenza, a public route:201with the number to quote. The row is written before any delivery is attempted, and whoever must read it gets it from a run that marks it done only on successful delivery: no request is lost because a channel was down. Refusals with names:LICENSE_REQUEST_INVALIDsays which field,413over 4096 bytes,429RATE_LIMITwithRetry-After(at most 5 per minute per address).- Privacy of the request: only the SHA-256 of the address; the service log holds number, id and plan — never the name, never the email, never the text of what you wrote.
- The form on both faces: if the route answers badly the message is readable, without codes, and what you wrote stays in the fields. The
mailto:remains an alternative route, not the only one. - The landing page is now served with the repository's real configuration, and is measured from the user's seat: 7 elements out of 8. What is missing is the comparison with the market leader with real numbers, for one reason only: their key is not in the vault yet.
Who it is for
Whoever has no key yet and wants to understand what it costs and how to start.
How to adopt it
Nothing to do if you already have a key. The contract stays 1.0: additions only, no previous answer changes.
30/9/2026Every key has a name: you know which one you are revoking.
What changed
- Three keys the same administrator issues for three different uses (production, test bench, laptop) were indistinguishable: same
owner, same scopes, same quota. Measured: 0 rows out of 3 said what the key is for, and the only field telling them apart was theid. A name sent at creation time was silently thrown away, with a201. - The
nomefield onPOST /v1/console/keys, in the list and in the reply: it is notowner(whose the key is), it is what it is for. Unique among your company's active keys, case- and space-insensitive: a name already in use is409 KEY_NAME_TAKENand says which key carries it; empty, over 60 characters or with control characters is400 KEY_NAME_INVALID, and the key is not created. The name of a revoked key is freed. - The rotation carries the name over, like the expiry: the
key_idchanges, the name does not. - It is searchable:
GET /v1/console/keys?nome=collaudo(the whole name, not its beginning); a name matching nothing returns empty, not every key. Measured: from 3 keys out of 3 to 1, rows to read by hand from 3 to 0. PUT /v1/console/keys/{key_id}/nome: renames a key that already exists (only if active), and the reply says what it was before.- The log says what the key is called (
nome_chiave): «revoke ab12cd34» told nobody anything. The name is not copied into the log but resolved at read time, and another company's key name never appears — not even on the row of a refused attempt.
Who needs it
Whoever has more than one integration: revocation takes the key id as input, and whoever cannot tell them apart revokes at random.
How to adopt it
Nothing is forced: a key without a name works as before and the contract stays 1.0. The name is read by your company and by us: it is not a place for a person's data.
30/9/2026The log also says what WE did to your keys, and it is searchable.
- When the control room steps in on a company's keys (a revocation, a rotation, a key issued for it) the row is now in the company's log, and it says whose the actor is (
attore_tenant,attore_e_della_mia_azienda). Measured: from 0 operations out of 4 seen to 4 out of 4. - Three filters on
GET /v1/console/audit:?key_id=,?persona=(id or email, looking both at who underwent the operation and at who performed it) and?operazione=. A filter matching nothing returns empty, not everything. Measured: from 0 questions out of 3 to 3 out of 3, and the rows to read by hand from 100 to 0. - The ceiling is declared:
limitefrom 1 to 1000, 100 by default. Before,limite=1000000returned the whole log (20,124 rows, 2,604 kB) and a negative limit did too; now it is400 AUDIT_LIMIT_INVALID, and the reply saysquanterows match and whether it istroncato: no silent truncation. - Never a token, never a fingerprint, never an email in the rows: it was already so, now it is measured on every run.
30/9/2026The service's numbers are not a customer's numbers.
- From the self-service console a company only gives itself the scopes it buys —
transcribe,stream,batch: asking formetrics(the monitoring overview) ortestis now403 SCOPE_NOT_SELF_SERVICE, and those keys are issued by the control room. Before, asking was enough: 4 requests out of 4 accepted. GET /v1/quadrostops at the key's perimeter and writes it down (perimetro): your company if the key has a tenant, the whole service only for the control room. Measured: from 3 companies out of 3 and 4 keys not your own (with other customers' names, products and seconds of audio) to 1 company and 0 keys not your own.GET /metricsare the service's counters and have no per-company label: to a key with a tenant they answer403 METRICS_TENANT_FORBIDDENand say where its numbers are.- Nothing changes for those who use the API: your company's numbers are where they were —
GET /v1/console/keys(keys, seconds per key, company total, plan, quota percentage) andGET /v1/usageper key. A dashboard of yours using ametricskey keeps working and sees your company.
30/9/2026The keys of whoever left the company: now you can see them, and remove them in one move (frontend).
- A person who had issued keys from the console and then leaves the company used to leave behind credentials that were alive and invisible: suspending her removed the console, not her keys. Measured: 3 keys out of 3 still active and transcribing after the suspension, and from the list the administrator recognised 0 out of 3 as hers (
ownerwas an opaqueu_…, the email appeared nowhere). - The list says whose each key is:
GET /v1/console/keyscarriesowner_email,owner_statoandowner_e_personanext toowner— an active key of a suspended person is visible at a glance. Never a person's token, never its fingerprint. Measured afterwards: 3 out of 3 attributed from the list response alone. - The per-person filter:
?owner=accepts the id, the email (case-insensitive) or a product name; anownermatching nothing returns an empty list, not every key. - Her keys are removed in one move:
POST /v1/console/keys/revoke-by-ownerinstead of one revocation per key knowing which ones they are (3 calls → 1). An administrator calls it, and everyone on their own. A name with no key is 404OWNER_NOT_FOUND. - No suspension revokes anything on its own: the key belongs to the company, not to the person, and inside it there is the integration that person had set up. The choice is yours and it is explicit (
chiavi: "lascia"by default,"revoca"on request) and it ends up in the log either way. - Revocation by owner has no grace period, unlike rotation: rotation hands you a new secret, a revocation does not — a window would only keep valid the secret of whoever left (measured: it still transcribed for all 24 hours). If you need the integration alive, rotate the key. The transcription contract is unchanged.
30/9/2026The token you enter the console with can be changed, and an invite does not last forever (frontend).
- A person's
agt_…token is the self-service credential. We measured what it is worth if it ends up in a chat: a stolen person token performs 9 console operations out of 9 — including the tenant's keys, which its holder manufactures at will and which transcribe audio; a stolen key, 0 out of 9. - The remedies were 0 out of 3: suspending the person stopped her too, and reactivating her handed back the same secret. Now
POST /v1/console/users/{id}/rotate-tokengives a new secret without changing identity — same id, same role, same history, same keys — and whoever is left with the old one reads 403TOKEN_ROTATED. Measured after: the thief from 3 out of 3 to 0, the person with the new one 3 out of 3. - You change your own, even as a plain user: whoever saw their token leak does not wait for the administrator. On another company's people it is 403
TENANT_FORBIDDEN. - By default the old one dies at once (
grazia_ore: 0), unlike a key rotation which keeps the old key alive for a day: a key is rotated for hygiene, a person token because it leaked. The grace can be asked for, for the token inside a script. - An unaccepted invite expires: 7 days by default (
scadenza_giornito shorten it, 90 at most), then 403INVITE_EXPIRED. Before, an invite sent to an old address still entered the console 400 days later. The expiry belongs to the invite, not to the person: when the invite is accepted it is cleared. The transcription contract is unchanged.
30/9/2026You invite your company's people yourself (frontend).
- The self-service console already knew how to hand out your tenant's keys; its people, no. Roles, statuses and isolation between companies were already inside the service, but with no door: letting a colleague in took a request to us. Measured: of the six things a company's administrator must be able to do alone, 0 out of 6 could be done from where they sit; now 6 out of 6.
- The routes:
POST /v1/console/users(invite),POST /v1/console/users/accept(whoever was invited accepts with their own token),GET /v1/console/users(your people with role and status),POST /v1/console/users/{id}/suspend,.../reactivate,PUT /v1/console/users/{id}/role. - The invite token is shown once only, like a key's secret: afterwards only its fingerprint is kept, and there is no way to read it again — not from the list, not from the log.
- Inviting, suspending and assigning roles belongs to the administrator: a plain user gets 403
ADMIN_REQUIRED(the list they do read). Another company's people are 403TENANT_FORBIDDEN, not an empty list. - Your company cannot lock itself out: the last active administrator can be neither suspended nor demoted (409
LAST_ADMIN). Before, one click left it with zero administrators and nobody able to invite. - Two clicks on the invite do not make two people: the same email invited again is 409
USER_EXISTS. Every operation leaves a row in the log with who did it and on whom: the person's id, never their token and never their email. The contract is unchanged.
30/9/2026A burst of audio no longer makes anyone wait for their final (engine v0.3.3).
- The engine computes a partial every so often while you speak: it is a preview of the text and it costs a whole ASR run on the pool shared by all the calls, hence the rule «one partial at a time». The rule held as long as the audio arrived at clock pace, the way a phone sends it; it did not hold on a burst — a client catching up after a network gap, a file pushed in streaming — because the guard was raised when the computation began, not when the partial entered the queue.
- Measured on the box (test engine,
basemodel on CPU, 2 ASR workers, 4 s of speech per call): 7 partials asked per call in a burst against 1–3 at pace; and the final you ask for withstopwas queued behind all of them — with 8 calls at once the latest one after 33.0 s instead of 12.8 s. - Now the guard is raised when the partial enters the queue: on a burst one starts, and the final no longer waits for work that it was itself about to replace.
- At phone pace nothing changes: the partials you receive are the same as before, with the same timings, and the contract is unchanged.
- In the engine's diary the closing line carries
chiesti=N, the partials put in the queue: neither the ones sent (partials=) nor the ones discarded empty (scartati=). The log saidpartials=1while seven were queued.
30/9/2026When the frontend restarts, whoever is talking does not hear it from the silence (frontend).
- The frontend restarts on every update, and a restart arrives while people are talking. Until yesterday that was a socket dying: 0 calls out of 3 received the final of what they had already said — a text the service had in hand — and 0 out of 3 a warning.
- Now whoever is talking receives
{"type":"closing","code":"SHUTTING_DOWN","retry_after_s":10}, then the final of what they had already said, then the close with code 1012 ("service restart"), which WebSocket libraries retry by themselves. Measured: from 0 out of 3 to 3 out of 3. - Long jobs queued or running go to
failedwithSHUTTING_DOWNimmediately and are not billed: they live in RAM only, a restart loses them, and knowing it is worth more than a404arriving later. - Whoever arrives while the frontend is shutting down gets 503
SHUTTING_DOWNwithRetry-After, andGET /healthanswers 503 withdraining: true: a load balancer takes it out of rotation before the address goes silent. Measured: address silent from 1.06 s to 0.08 s. - Nothing to change in the products:
closingis one more message, 1012 is the standard code, the contract stays 1.0. Window 15 s (AGILEASCOLTO_CHIUSURA_S,0= as before): beyond it the frontend exits anyway.
29/9/2026The quota also counts the audio that is coming in right now (frontend).
- A call books its seconds when it ends, and a file once the engine has answered: in the meantime, until yesterday, those seconds existed for nobody. The monthly quota (and the company plan) was therefore not a cap on the audio listened to, but a cap on one call at a time.
- Measured: three lines open together got 18 s against a cap of 6 (3.0x, on both the key and the tenant plan) and a file sent while speaking saw nothing (2.0x). After the fix, same bench: 1.0x in all three cases.
GET /v1/usagecarries two new fields,in_corso_sandtenant_in_corso_s: the seconds that came in and are not booked yet (open calls, files being processed, queued jobs). Added fields: the contract stays 1.0.- A long session re-reads the month's usage every 15 s (
AGILEASCOLTO_RINFRESCO_QUOTA_S), and the audio block refused for an exhausted quota is no longer billed (before: 6.5 s billed against a quota of 6). - Nothing to change in the products: whoever stays inside the quota sees no difference, and without quota or plan nothing changes at all.
29/9/2026When the service says no, it also says when to retry (frontend).
- Refusals that are worth retrying now carry
Retry-After(whole seconds) and the same number inerror.retry_after_s: the 429s (RATE_LIMIT,QUOTA_EXCEEDED,TENANT_RATE_LIMIT,TENANT_QUOTA_EXCEEDED) and the 503CAPACITY. On the WebSocket, where headers do not exist, the field is in the error message. Added fields: the contract stays 1.0. - The three numbers say different things: on the per-minute ceiling it is the exact instant a slot frees up (the window is sliding); on the monthly quota it is the start of next month, because before then nothing gets through anyway; on a full frontend it is a hint of a few seconds, deliberately spread (
AGILEASCOLTO_RIPROVA_PIENO_S,AGILEASCOLTO_RIPROVA_SPREAD_S) so that everyone refused together does not come back together. - The other refusals — 401, 403, 404, 413, 501, 502 — do not carry it: they are refusals on the merits, and retrying would change nothing.
- Nothing to change in the products:
Retry-Afteris read by HTTP libraries and OpenAI clients on their own. - Why: measured with a client that retries the way libraries retry. Without the number, getting 6 requests through a ceiling of 3 per minute meant knocking in vain 7 times instead of 1 and taking 75.8 s instead of the 60 s minimum; with the quota already used up it sent 5 useless requests in 15 s. With the number: 1 knock, 62.6 s, and a single request (0.03 s) once the quota is over.
29/9/2026With the second ear the seats really double (frontend).
- The frontend can sit in front of several engines, and adding one is there to serve more clients. As long as the capacity ceilings belonged to the frontend, though, the seats were shared: one engine's clients filled up the ceiling of the clients of another engine that was idle, and the extra machine gave nothing to anyone.
- Now the ceilings sit where the capacity they protect sits. Per engine: the streams (
AGILEASCOLTO_MAX_STREAM), the batch requests (AGILEASCOLTO_MAX_BATCH) and the long jobs waiting (AGILEASCOLTO_MAX_JOBS_CODA), with one queue and one worker per engine. The frontend keeps the audio in RAM (AGILEASCOLTO_MAX_CODA_S) and the requests in flight (AGILEASCOLTO_MAX_IN_VOLO), because the RAM is one for however many engines there are. - With a single engine nothing changes: the same numbers as before (8 streams, 8 jobs waiting, one job at a time). That is the case for all of today's deployments.
- With two engines the seats are twice as many and the full engine no longer causes a rejection for whoever asks for the other one. The 503
CAPACITYnow names the full engine, and a queued job'spositionis the position in its own engine's line. GET /healthcarries new fields inside each engine:streams_active,streams_max,jobs_running,jobs_queued,jobs_queued_max;jobs.runningsays how many jobs are running now on the whole frontend (with N ears they can be N). The same numbers are in/metrics, one line per configured engine. Fields are added: the contract stays 1.0.- Why: measured with two fake ears and real clients. Before, the second ear added 0 streams, whoever asked for the idle engine got 503
CAPACITY, the free ear's batch waited 2.63 s instead of 0.13 s and the free ear's long job waited 2.53 s instead of 0.06 s. After: 4 seats on two ears, batch 0.12 s, long job 0.13 s.
29/9/2026A ceiling on the requests with audio in flight (frontend).
- The previous day's ceiling limits what the queue holds; this one limits the peak: how many requests may hold audio in RAM at the same time, on
POST /v1/audio/jobsand onPOST /v1/audio/transcriptions. They are nowAGILEASCOLTO_MAX_IN_VOLO, by default 4. - It is needed because a request that is coming in already holds the upload body and the decoded PCM in memory, and the file's duration is known only after decoding it: admission cannot stop it earlier. How many arrived together was the client's choice.
- For integrators nothing changes up to 4 requests at a time. Above the ceiling the request waits its turn (up to 30 s) and only then gets 503
CAPACITY: in a burst the file does not have to be sent again. Waiting costs 64 KiB, against the 45 MB of a request that got in. GET /healthcarriesinflightandinflight_max: the requests holding audio in RAM now and their ceiling. The same numbers are in/metricsand in/v1/quadro.- Why: measured with the client in a separate process, 12 requests at once of 600 s each are worth 736 MB of peak (45.1 MB per request) that no ceiling was counting. Under a RAM ceiling those same 12 kill the frontend with 0 jobs admitted and no 503 to anyone. With the semaphore, same burst and same ceiling: 12 admitted, 0 refused, frontend still standing.
29/9/2026The job queue has a ceiling on the audio held in RAM (frontend).
POST /v1/audio/jobskeeps the whole audio of the file in RAM until the job is over, and the queue was measured in jobs (8 waiting): those same eight places are worth 1.7 MB if the files last 6 seconds and 3110 MB if they last the 3 hours the contract promises.- There is now a ceiling on the seconds of audio in RAM as well (
AGILEASCOLTO_MAX_CODA_S, by default 21,600 s = 6 h = 691 MB): twice the maximum duration of one job, so that a file as long as the contract says always fits into an empty queue. - Above the ceiling comes a 503
CAPACITYthat says how many seconds are free; a file that on its own would not fit into any queue, not even an empty one, gets 413AUDIO_TOO_LONGand not a "try again" that would never come. GET /healthcarriesqueue_audio_sandqueue_audio_max_sin the job queue: the seconds of audio in RAM and their ceiling. Once the jobs are donequeue_audio_sis 0 — the audio dies with the job.- Why: without a ceiling on the sum, when the RAM runs out it is not one request that fails — the frontend dies, and the jobs, which live in RAM only, all disappear, for every client. With the previous limits and a 160 MB frontend the seventh job had the process killed without a single 503; with the ceiling, the sixth gets a 503 and the frontend stays up. Whoever sends short files notices nothing.
29/9/2026Have one block of the code dictated again, not the whole of it: ricomponi + blocco (frontend).
- When a code does not hold up and there is no repair, there is no need to have the whole of it read again: ask for one block — of the tax code
cognome,nome,nascita,comune,controllo; of the IBANpaese,banca,sportello,conto; of POD and PDRtestaandcoda(phone numbers and case numbers are short: they are read back whole). - Send the answer's turn with
ricomponi=<the previous code>andblocco=cognome— on the stream the message{"type":"ricomponi","sentito":"…","blocco":"cognome"}, which arms the next final. Back comesricompostowith the code put together again and re-checked. - Look at
giudicabile: it says whether the check for that type has a checksum and can act as a judge of the recomposition. Out of 33 wrong recompositions of IBANs and tax codes, those that pass the check are 0; on PODs, which only have a shape, 12 out of 12 pass — withgiudicabile: falsethe code must be read back to the person. - You choose the block, because you are the one talking to them: the frontend does not know where the error is, and the rule that made it guess got 33 blocks right out of 77. Measured on 107 wrong codes: knowing which block is broken, a single block puts 57 of them right.
- The ladder: if one block is not enough, ask for the next one, in the order of the code, and let the check tell you when to stop. It works because the first graft brings the code back to its canonical length: 30 tax codes out of 36 and 29 IBANs out of 29 make it to the end without dictating anything whole again (median 4 questions of 3–5 characters), PODs and PDRs all in 2 questions, 0 wrong recompositions accepted.
- It is worth it because a short piece can be heard and a whole code cannot: on the same reports a tax code dictated whole comes out right 9 times out of 90 and an IBAN 23 out of 90, while a code of 7–10 characters is between 74 % and 93 %.
- New route
GET /v1/codici/blocchi?tipo=cf&sentito=…: the blocks of that type with the sentence to say and, for each one,posizionabile— a block can only be grafted if the code that was heard is long enough, and asking for one that does not graft throws a turn away (it would be 44 questions out of 180 on the archive's 36 wrong tax codes). Blocks that do not graft are skipped and come back into play on the next round: that is what takes the IBAN from 27/29 to 29/29. - Whoever does not send
ricomponisees exactly what they saw before.
29/9/2026You can tell the frontend which code you asked for: atteso (frontend).
- The frontend recognises codes by their shape: if the ear drops a character the shape no longer holds and the code comes out with the wrong type or with no type at all (a 14-character tax code comes out as
sconosciuto). Whoever was looking for their own type found nothing to say to the person. - With
atteso=cf(batch, jobs,?atteso=cfor"atteso":"cf"in the stream'sstart, or the message{"type":"atteso","tipo":"cf"}mid-session) the candidate that answers the question carriesatteso: true; if none of them has that shape any more, the closest one is adopted and re-checked as that type (adottato: true), with the reason to say out loud: "length: 14 characters, a tax code has 16". - A valid candidate is never adopted (in that sentence it is somebody else's code) and adoption makes nothing valid: the check stays the real one for the type. Measured on the bench's reports: 186 wrong codes with no candidate of the expected type, 48 adopted, 47 right, 0 adoptions on plain speech.
- Whoever does not send
attesosees exactly what they saw before.
29/9/2026A code that does not hold up also says how it can be repaired (frontend).
- With
codici=1, a code that fails the check carries two extra fields when there is one and only one repair that makes it hold:"riparato": "IT881E26716388", "regola": "pod: E mancante". Two rules and only two: the POD'sEthat the ear loses among the digits, and the IBAN country heard as "ID" instead of "IT". validostaysfalse: the repair is a proposal to be read back and confirmed ("I have IT881E26716388, is that right?"), not an accepted code. Over the phone this saves dictating 27 characters again without risking handing over somebody else's code.- Measured on the bench's reports (12 proposals, 12 right, 0 wrong, 0 on plain speech) and live with the default model (12 audio files, 2 repairs, 2 right). Two looser rules — recomputing the tax code's check character, guessing the IBAN's CIN — were measured and discarded: they invented formally perfect codes.
- Whoever does not read the new fields sees exactly what they saw before.
28/9/2026The stream says which ear is listening to you (frontend).
- The
{"type":"started"}message with which the WebSocket stream answers thestartnow carries themodelfield: the name of the ear that will transcribe this session —{"type":"started","model":"agileascolto-turbo"}. It is our own name, the same one that appears in/v1/usage, in/v1/modelsand in the log, even if you asked for thewhisper-1alias or did not ask for any model at all. - It is for those who must be able to say who listened to a call: before, only the batch (
verbose_json.model) and a job's status said it, while on the stream the alias and the default model hid the ear. The field is always there, it is not requested, and with it the session can be found again in the key's usage. - Nothing else changes: whoever does not read it sees the previous stream, identical.
23/9/2026A log of the discarded partials (engine v0.3.2).
- The engine writes into its log every partial that the ASR discards for nothing (
empty,ghost(...),degen[...]): a linepartial scartato: <reason> (<ms> of speech, asr <ms>)and ascartati=counter in the closing line. - No text in the log: the contract is unchanged. It is there to understand why, at real-time pace, the first partial was coming out late — the first attempt comes back empty and before this could not be seen.
23/9/2026Words with their times (engine v0.3.1).
words=1(ortimestamp_granularities[]=word, as in OpenAI) on/v1/audio/transcriptionsand on jobs: inverbose_jsonevery word has a start and an end in seconds, at the top and inside every segment. On the WebSocket stream:{"type":"start","words":true}and every final carrieswords.- The
srt/vttsubtitles asked for withwords=1take the real times of the words: the cuts fall on the pauses between words and every caption runs from its first word to its last. Without it, the time is split in proportion to the characters, as before. - It is the foundation for "who spoke when" on classroom recordings.
22/9/2026SRT and WebVTT subtitles, jobs on long files.
response_format=srtorvttin the batch,?format=srt|vtt|texton jobs: captions of at most two lines of 42 characters and 7 seconds, long sentences broken at word boundaries.verbose_jsonnow carries the start and the end of every segment.POST /v1/audio/jobs: files up to 3 hours in a queue, one at a time, with progress (engine v0.2.1);keep=0= the result is delivered once only and never stored.
22/9/2026Checking dictated codes, and an engine that talks to nobody.
What changed
- With
codici=1every final also carriescodici: the codes found in the text with their type (IBAN, tax code, POD, PDR, phone number, case number),valido(mod 97 checksum for the IBAN, check character for the tax code, shape for the others) and amotivoin plain words ("length: 25 characters, IT has 27"). The product asks for a repeat only when it is needed, and only for the block that is needed. - The engine contacts no external service: the inference library's telemetry is off, the model is loaded from the local cache. Verified with a repeatable audit (strace + sockets), zero connections.
Who it is for
For those who collect IBANs, tax codes, PODs and PDRs over the phone; for those who must be able to say where the data ends up.
How to adopt it
No forced change: codici appears only with codici=1. The rules for the confirmation dialogue are in the documentation.
22/9/2026Watching over the service.
What changed
- The service watches itself: the frontend reads
/healthzevery minute and only raises its voice when it has to —AGILEASCOLTO ALLARMEwhen something falls (the frontend or a single engine),ANCORA GIU'every 30 minutes,RIENTROwith the minutes it was away. When everything is fine it says nothing. /healthzgainsengine_down_total(the503 ENGINE_DOWNsince startup) anduptime_s: if the counter grows between two checks it is an alarm even when the engines are back up, so that a few seconds of outage inside a phone call does not disappear.- No automatic remedy: the old one stays switched on.
Who it is for
For those who operate the service and want to be warned when something falls, without noise when everything is fine.
How to adopt it
Nothing to do: the watch is internal to the service, it requires no configuration on the clients' part.
21/9/2026A usage page per key, and post-processing of dictated codes.
What changed
GET /uso: you paste the key and see the quota and the usage by month, operation and model; with themetricsscope, the overview of all the keys and the state of the engines. The key never travels in the address.codici=1(batch form,?codici=1orstart{"codici":true}on the stream): IBANs, tax codes, PODs and case numbers dictated letter by letter come back compact ("bi esse trattino 1 5 7 6 7"→BS-15767). On the finals only, on request.GET /v1/quadrofor those who watch over the service (metricsscope).
Who it is for
For those who keep an eye on the usage of a product or of a customer; for those who transcribe codes dictated over the phone.
How to adopt it
Nothing to change in the clients: the page opens in a browser; codici is one more option.
16/9/2026Contract 1.0: OpenAI-compatible file API, WebSocket streaming, keys with a quota in seconds.
What changed
POST /v1/audio/transcriptionswith the OpenAI format (json,text,verbose_json).WS /v1/audio/stream: partials and finals in streaming, audio at 8 or 16 kHz.- Keys per product and per customer, with allowed models, a per-minute limit, a monthly quota in seconds and rotation;
GET /v1/usage. - No silent fallback: if the engine does not answer,
503 ENGINE_DOWN.
Who it is for
For those who already use an OpenAI- or Deepgram-compatible transcription API and want to keep the audio in Europe.
How to adopt it
Ask for a key, change the base_url, use model="agileascolto-turbo".