Architecture
Source code: https://github.com/Coderctrldev/Audio_Notes
Overview
A user uploads an audio file. The API stores it and returns immediately. A background worker validates and splits the audio, transcribes the pieces with Gnani's speech-to-text API, and asks an LLM (Google Gemini) for a structured summary. The browser polls a status endpoint, so the user sees a live percentage, the current stage, an estimate of the time left, and a plain-language reason whenever something fails or a summary can't be produced.
System diagram
Loading diagram…
Request flow: upload to summary
Loading diagram…
- The browser checks the extension and that the file isn't empty, then uploads with XMLHttpRequest so the user sees real upload progress. There is no client-side size limit.
POST /uploadsvalidates type, language and size, creates anuploadsrow, streams the file to object storage, enqueues a job and returns202with the id. No transcription happens inside this request.- The browser opens the upload's page and polls
GET /uploads/{id}every two seconds until the status is final. - The worker writes progress to Postgres after every chunk, which is what the poll reads.
- The home page lists past uploads (searchable and filterable) and each one reopens with its stored transcript and summary.
Upload lifecycle
Loading diagram…
Status changes use guarded updates (UPDATE … WHERE status = …). That is what lets exactly one worker claim a job, and what stops a cancelled job from being overwritten by a worker that was already running.
Where files and data live
- Original audio: S3-compatible object storage, key
uploads/<id>/original.<ext>, accessed only through a smallstorage.py. Local development uses a self-hosted S3-compatible server; a deployment uses a managed bucket (Cloudflare R2, S3, …) with only environment changes. - Chunks: temporary files in the worker, deleted when the job ends. Only their transcripts are persisted.
- Metadata, transcript, summary, per-chunk state: Postgres. The summary is stored as JSON text, so its shape can evolve without a migration.
- Job queue: Redis (RQ).
Loading diagram…
Handling long audio
Gnani's REST endpoint accepts a short clip per request (30 seconds recommended, 60 at most). The worker decodes the file with ffmpeg to mono 16 kHz WAV and cuts it into 30-second chunks, so a 2-minute file becomes 4 or 5 chunks and a 2-hour file about 240. Up to TRANSCRIBE_CONCURRENCY chunks (default 3) are sent in parallel and the results are joined in chunk order. Memory use doesn't grow with file length: the upload streams to storage and the worker keeps only a few chunks in flight.
Progress is reported as a percentage: 5% once the audio is prepared, up to 90% across the chunks, 95% while summarizing and 100% when done. The UI also shows parts done, elapsed time and an estimated time remaining based on the chunk throughput actually observed.
Languages
The language selector offers the Gnani STT languages: English (India), Hindi, Bengali, Gujarati, Kannada, Malayalam, Marathi, Punjabi, Tamil, Telugu, Assamese and Odia. The server rejects any other code. The summary step also checks the transcript, so audio in a language that wasn't selected, or isn't supported, produces a clear message instead of a misleading summary.
Summaries and when there isn't one
Loading diagram…
A successful summary has a title, overview, key points, decisions, action items and topic tags, and the model is told not to invent anything that isn't in the transcript. Cheap checks run first (no speech, too short) so no LLM call is wasted. If the LLM itself fails, the upload still completes with the transcript, and Regenerate re-runs only the summary without redoing transcription.
Synchronous vs background
| Synchronous (API request) | Background (RQ worker) |
|---|---|
| Validate extension / language / non-empty, stream to storage, create the row, enqueue, return 202. Reads, cancel, delete, retry. | Download, ffprobe validation, splitting, per-chunk ASR with retries, joining, LLM summary, all status and progress updates. Jobs have a 2-hour timeout. |
Keeping transcription out of the request avoids HTTP timeouts and keeps the API responsive while audio is processed.
API
| Endpoint | Purpose |
|---|---|
POST /uploads | Upload a file with a language; returns 202 and an id |
GET /uploads, GET /uploads/{id} | List, or fetch status, percent, transcript and summary |
POST /uploads/{id}/cancel | Stop a queued or running job (worker stops at the next checkpoint) |
DELETE /uploads/{id} | Remove the row, chunks and stored audio (running jobs must be cancelled first) |
POST /uploads/{id}/retry | Resume a failed, cancelled or stalled upload; finished chunks are skipped |
POST /uploads/{id}/resummarize | Regenerate only the summary from the stored transcript |
Failure handling
| Situation | What happens |
|---|---|
| Wrong type, empty or corrupt file | Rejected up front, or ffprobe/ffmpeg fails in the worker and the upload becomes failed with a readable message. |
| ASR 429 / 500 / 503 / network | Retried with exponential backoff (1 s, 2 s, 4 s). |
| ASR 400 / 403 | Not retried (bad input or key); fails immediately with a clear message. |
| A chunk fails | Its row is marked failed, the upload fails, and retry resumes from that chunk. |
| Summary too short, vague, wrong language, unclear | Upload completes; the UI explains why there is no summary and keeps the transcript. |
| LLM service error | Upload completes with the transcript; the summary can be regenerated. |
| User cancels or deletes | Guarded status updates make the worker stop at its next checkpoint without overwriting the result. |
| Dead worker / duplicate delivery | An atomic claim lets one worker run a job; a job idle for 15 minutes can be retried from the UI. |
| Storage or queue failure | An error is returned (and logged with a traceback), or the row is marked failed so it never sits in queued unexplained. |
| Browser offline | Polling continues through temporary errors and says so. |
Trade-offs and limitations
- Fixed 30-second cuts can land mid-word, which may slightly hurt accuracy at chunk boundaries.
- Execution is at-least-once with an atomic claim, not exactly-once: a worker that crashes after receiving an ASR result but before saving it redoes that chunk on retry.
- Polling every 2 seconds is simple and robust, but less efficient than push updates.
- There is no authentication: anyone with the URL can see all uploads. Tables are created at startup rather than through migrations.
- A very long transcript is condensed in segments before the final summary, which can lose detail.
- A cancel is cooperative: an ASR call already in flight finishes first, so it can take a few seconds to take effect.
- The language check on the summary is an LLM judgement, so unusual code-mixed speech can occasionally be flagged incorrectly (Regenerate is available).
With more time
- Per-user accounts and private uploads; signed URLs to play the original audio.
- Silence-aware chunking with a small overlap, to avoid cutting words.
- Gnani's batch API for very large files; adaptive concurrency from rate-limit headers.
- Server-sent events instead of polling; a worker heartbeat instead of the 15-minute rule.
- A transactional outbox (or sweeper) for lost enqueues; Alembic migrations; worker tests in CI.
- An expiry policy for stored audio; speaker labels and timestamps in the transcript.
Stack
Next.js · FastAPI · PostgreSQL · Redis + RQ · S3-compatible object storage · ffmpeg/ffprobe · Gnani speech-to-text · Google Gemini (summary) · Mermaid (diagrams).