Audio Notes

Architecture

Source code: https://github.com/Coderctrldev/Audio_Notes

Overview

A user uploads an audio file. The API stores it and returns immediately. A background worker validates and splits the audio, transcribes the pieces with Gnani's speech-to-text API, and asks an LLM (Google Gemini) for a structured summary. The browser polls a status endpoint, so the user sees a live percentage, the current stage, an estimate of the time left, and a plain-language reason whenever something fails or a summary can't be produced.

Never blocks the requestAll slow work happens in the worker, not in the API call.
ResumableFinished chunks are saved, so retry and restart skip them.
Always explains itselfEvery failure has a stored reason shown in the UI.
User stays in controlCancel, delete, retry and regenerate the summary.

System diagram

Loading diagram…

Request flow: upload to summary

Loading diagram…

  1. The browser checks the extension and that the file isn't empty, then uploads with XMLHttpRequest so the user sees real upload progress. There is no client-side size limit.
  2. POST /uploads validates type, language and size, creates an uploads row, streams the file to object storage, enqueues a job and returns 202 with the id. No transcription happens inside this request.
  3. The browser opens the upload's page and polls GET /uploads/{id} every two seconds until the status is final.
  4. The worker writes progress to Postgres after every chunk, which is what the poll reads.
  5. The home page lists past uploads (searchable and filterable) and each one reopens with its stored transcript and summary.

Upload lifecycle

Loading diagram…

Status changes use guarded updates (UPDATE … WHERE status = …). That is what lets exactly one worker claim a job, and what stops a cancelled job from being overwritten by a worker that was already running.

Where files and data live

Loading diagram…

Handling long audio

Gnani's REST endpoint accepts a short clip per request (30 seconds recommended, 60 at most). The worker decodes the file with ffmpeg to mono 16 kHz WAV and cuts it into 30-second chunks, so a 2-minute file becomes 4 or 5 chunks and a 2-hour file about 240. Up to TRANSCRIBE_CONCURRENCY chunks (default 3) are sent in parallel and the results are joined in chunk order. Memory use doesn't grow with file length: the upload streams to storage and the worker keeps only a few chunks in flight.

Progress is reported as a percentage: 5% once the audio is prepared, up to 90% across the chunks, 95% while summarizing and 100% when done. The UI also shows parts done, elapsed time and an estimated time remaining based on the chunk throughput actually observed.

Languages

The language selector offers the Gnani STT languages: English (India), Hindi, Bengali, Gujarati, Kannada, Malayalam, Marathi, Punjabi, Tamil, Telugu, Assamese and Odia. The server rejects any other code. The summary step also checks the transcript, so audio in a language that wasn't selected, or isn't supported, produces a clear message instead of a misleading summary.

Summaries and when there isn't one

Loading diagram…

A successful summary has a title, overview, key points, decisions, action items and topic tags, and the model is told not to invent anything that isn't in the transcript. Cheap checks run first (no speech, too short) so no LLM call is wasted. If the LLM itself fails, the upload still completes with the transcript, and Regenerate re-runs only the summary without redoing transcription.

Synchronous vs background

Synchronous (API request)Background (RQ worker)
Validate extension / language / non-empty, stream to storage, create the row, enqueue, return 202. Reads, cancel, delete, retry.Download, ffprobe validation, splitting, per-chunk ASR with retries, joining, LLM summary, all status and progress updates. Jobs have a 2-hour timeout.

Keeping transcription out of the request avoids HTTP timeouts and keeps the API responsive while audio is processed.

API

EndpointPurpose
POST /uploadsUpload a file with a language; returns 202 and an id
GET /uploads, GET /uploads/{id}List, or fetch status, percent, transcript and summary
POST /uploads/{id}/cancelStop a queued or running job (worker stops at the next checkpoint)
DELETE /uploads/{id}Remove the row, chunks and stored audio (running jobs must be cancelled first)
POST /uploads/{id}/retryResume a failed, cancelled or stalled upload; finished chunks are skipped
POST /uploads/{id}/resummarizeRegenerate only the summary from the stored transcript

Failure handling

SituationWhat happens
Wrong type, empty or corrupt fileRejected up front, or ffprobe/ffmpeg fails in the worker and the upload becomes failed with a readable message.
ASR 429 / 500 / 503 / networkRetried with exponential backoff (1 s, 2 s, 4 s).
ASR 400 / 403Not retried (bad input or key); fails immediately with a clear message.
A chunk failsIts row is marked failed, the upload fails, and retry resumes from that chunk.
Summary too short, vague, wrong language, unclearUpload completes; the UI explains why there is no summary and keeps the transcript.
LLM service errorUpload completes with the transcript; the summary can be regenerated.
User cancels or deletesGuarded status updates make the worker stop at its next checkpoint without overwriting the result.
Dead worker / duplicate deliveryAn atomic claim lets one worker run a job; a job idle for 15 minutes can be retried from the UI.
Storage or queue failureAn error is returned (and logged with a traceback), or the row is marked failed so it never sits in queued unexplained.
Browser offlinePolling continues through temporary errors and says so.

Trade-offs and limitations

With more time

Stack

Next.js · FastAPI · PostgreSQL · Redis + RQ · S3-compatible object storage · ffmpeg/ffprobe · Gnani speech-to-text · Google Gemini (summary) · Mermaid (diagrams).