ABN
FR

Sep 2025 – present

Podcast Transcript AI — transcription at scale

Took over and rebuilt Podcast Transcript AI: self-hosted Whisper on a GPU server and a public library of 52,700+ transcripts in 73 languages.

Role
Lead developer (took over in 2025)
Period
Sep 2025 – present
Visit the live sitePodcast Transcript AI homepage

Results

Public figures and figures from the codebase, checked October 2026.

  • 52,700+public transcripts in the library (October 2026)source
  • 41,000+hours of audio transcribed into the librarysource
  • 73languages representedsource
  • 3,200+backend testsfrom the codebase

The product

Paste a podcast link from Apple Podcasts, Spotify or RSS and get a transcript, summary, chapters and FAQs. Transcripts are published in a public, searchable library, and developers can use the same engine through a paid API.

My role

Another developer built the first version. I took over in September 2025 and have written nearly all of the code since — about 576 of 593 backend commits — including the move to self-hosted transcription.

What I built

  • Self-hosted transcription. Replaced a paid speech-to-text API with whisper.cpp running on our own GPU server, with three model sizes.
  • A queue across two servers. Jobs live in MongoDB with paid, free and crawler tiers; workers on two machines claim jobs, recover from crashes and never run more than the GPU’s memory allows.
  • Long audio. Files are split into chunks with their own time budgets, timestamps are stitched back together, and progress is real.
  • Cost-controlled AI. Summaries route between a local model and a hosted one, with a daily spending cap — added after an outage in which an empty API balance published 531 transcripts without summaries.
  • A library that grows on its own. A crawler reads the charts, finds each show’s feed and transcribes back catalogues when the GPU is idle, so paying users always go first.
  • Finding the audio. When a platform blocks downloads, a fallback chain searches other public sources and matches the episode by its duration.
  • More features. Speaker labels, search across every transcript, chat with an episode, PDF/SRT/VTT exports, Notion and Obsidian export, and a public API with credit plans.
  • A second brand on the same backend. A sister transcription product runs on this backend with isolated accounts, sessions and payments.

Hard problems

  • A clean-up step that froze the server. Removing Whisper’s repeated lines was quadratic and blocked the server for minutes on long episodes; I rewrote it to run in bounded time.
  • Billing that’s right every time. API refunds are claimed exactly once, and an append-only ledger records every credit change.

Stack

  • Next.js
  • Node.js
  • Express
  • MongoDB
  • whisper.cpp
  • Meilisearch
  • DeepSeek
All work

Have a product to build or fix?

Tell me what you're working on. I'll reply with questions or a first plan.

Start a project