Podcast Transcription Services Compared: Human vs AI vs DIY Whisper (Real Cost per Hour, 2026)

podcast transcription service cost per hour compared: human, AI SaaS and DIY Whisper
Built by me · 180 min freePodskriftPodcast to Text TranscriptionSee how it works →
On this page

A podcast transcription service turns an episode's audio into text, and in 2026 you choose between three kinds: human transcriptionists at about $120 per audio hour, AI services that cost roughly $1 to $12 per hour, and DIY Whisper at $0.36 per hour through OpenAI's API. Price is the easy part to compare. Whether the transcript is complete, and whether the service can even get the audio, is harder to see on a pricing page.

I built a podcast transcription service, Podskrift, so I've seen the failure modes from the inside. This guide compares the three types with prices taken from each vendor's pricing page in October 2026, and shows the checks I now run on any transcript. Disclosure: Podskrift is my own product, and it's in the comparison.

What a podcast transcription service actually does (human vs AI vs DIY)

Every service does the same three things: get the audio, convert speech to text, and give you a file. They differ in who does the second step.

  • Human transcription. A person types the audio, sometimes on top of an AI draft. Rev guarantees 99%+ accuracy for its human service. You pay per minute, and you wait hours rather than minutes.
  • AI transcription. A speech model writes the transcript, and the service wraps it in an editor, speaker labels and exports. Most "podcast transcript AI" tools you find in search are this type, built on a model like OpenAI's Whisper.
  • DIY Whisper. You run the model yourself, either the open-source Whisper on your own machine or OpenAI's API with your own key. You handle downloads, file limits and formatting.

The type you need depends on what happens to the text afterwards. A court filing needs a human. Notes from an episode you listened to on Tuesday don't.

Real cost per hour of audio, compared

Pricing pages mix per-minute rates, monthly plans and "media hours", so I converted everything to dollars per hour of audio. Prices were checked on each vendor's pricing page on October 3, 2026. Monthly-billed prices are shown; most vendors are cheaper billed annually.

ServiceTypePublished priceCost per audio hour
RevHuman$1.99/min$119.40
Happy ScribeHuman proofreadingfrom $2.00/minfrom $120.00
Happy ScribeAI top-up credits$0.20/min$12.00
SonixAI, pay as you go$10/hr$10.00
Descript HobbyistAI editor, subscription$24/month for 10 media hours$2.40 if you use all 10 hours
Podskrift packAI$5 for 300 minutes$1.00
OpenAI whisper-1 with your own keyDIY API$0.006/min$0.36
Open-source Whisper on your computerDIY localNo per-minute feeYour hardware and time

Otter isn't in the table because it isn't priced per hour. Its Pro plan is $16.99 per user per month and allows 10 imported audio or video files per month; the free plan allows 3 imported files in total. It's built for meetings, so an imported podcast episode counts against a file limit rather than a minute rate.

Three things stand out:

  1. Human transcription costs about 100 times more than AI. It's worth it when every word carries legal or editorial weight. For notes and research it isn't.
  2. Subscriptions are only cheap if you use them. Descript's $2.40 per hour assumes you use all 10 hours every month, and Descript is a video and audio editor first.
  3. The model itself is cheap. A one-hour episode through whisper-1 costs OpenAI $0.36. Most of what you pay a podcast transcription service for is everything around the model.

For per-episode costs at Whisper's rate, see the table in my free podcast transcription tool post. Podskrift's own pricing: the first 180 minutes are free with no card, then a one-time $5 pack covers 300 minutes, or you add your own OpenAI key and pay OpenAI directly.

Accuracy: where AI transcription breaks

On a clean studio recording, a modern AI transcript is usually good enough to read and search without edits. It breaks in predictable places:

  • Crosstalk. Two people talking at once often come out as one merged sentence, and speaker labels get swapped.
  • Names and jargon. Guest names, product names and acronyms are the most common errors. Check every one you plan to quote.
  • The language setting. This one cost me money. An early version of Podskrift defaulted the language to Norwegian. A listener transcribed a Japanese episode without changing it and got phonetic nonsense back, and we paid for it. Whisper doesn't treat the language you pass as a hint. It treats it as a constraint and forces the audio into that language.

The fix for the last one is simple: leave the language on auto-detect unless the clip is short, heavily accented or mixes languages, and never set a language the episode isn't in. Whichever podcast transcription service you use, find out what its language setting actually does before you run a batch.

If you need guaranteed accuracy, that's the case for human transcription or a human proofreading pass over an AI draft. Happy Scribe sells exactly that combination.

The completeness problem nobody mentions

Accuracy gets all the attention, but the worse failure is a transcript that is accurate and incomplete. Nothing looks wrong. The text reads fine. A chunk of the episode just isn't there.

I hit two causes while building Podskrift's audio pipeline:

  1. File size isn't proportional to time. OpenAI's speech-to-text API accepts files up to 25 MB, so long episodes have to be split. Podcast audio is often variable bitrate, so splitting by bytes doesn't give equal stretches of time. One episode with a dense first half produced a 27 MB piece against a 24 MB target, over the API limit.
  2. The reported length can be wrong. Many podcasts insert ads dynamically, which stitches the file together from separate segments. On one such file, the duration the file reported was shorter than the real audio. An early version of Podskrift trusted that number, and 16 minutes were never sent for transcription. No error was raised.

The fix for both was to stop trusting file metadata: Podskrift now converts audio to 16 kHz mono and splits it into 15-minute pieces by walking the actual audio stream. You can't see any of this from outside a service, so test the output instead.

The last-timestamp test

Download the .srt version of the transcript and compare its final timestamp with the episode length shown in your podcast app (not the length the audio file reports, which is the number that can be wrong):

# Print the start and end time of the last subtitle in an .srt file
grep -E '^[0-9]{2}:[0-9]{2}:[0-9]{2},[0-9]{3} --> ' episode.srt | tail -n 1

If the episode runs 1:12:40 and the last subtitle ends at 0:56:10, part of the audio was never transcribed. A gap of a few seconds at the end is normal, since outros are often music. A gap of minutes isn't. It takes ten seconds, and it works on any podcast transcription service that exports .srt.

The input problem: upload vs RSS vs search by name

Most comparison pages skip the first step: getting the audio into the service. For podcasts, that's often the hardest part.

InputWho it suitsThe catch
Upload a filePodcasters with their own recordings, interviewsPodcast apps don't give you the episode file
Paste an RSS feedAnyone who can find the show's feedListening apps don't show the feed URL
Search by name or paste a linkListeners who heard an episode in an appSpotify-exclusive shows have no public audio

Upload-first services (Rev, Sonix, Descript, Happy Scribe) are built for people who already have the audio. If you produce the show, that's you. If you heard the episode in Spotify or Apple Podcasts, you first have to find the feed or the file. I built Podskrift's search because of that gap: you search the show or paste an episode link, and it fetches the audio itself. The step-by-step version is in how to transcribe a podcast to text, and the platform-specific cases are covered in the Apple Podcasts transcript guide and the Spotify transcript export guide.

Which podcast transcription service fits you

  • You produce the podcast. Use an upload-first AI service with speaker labels and an editor (Descript, Sonix, Happy Scribe). If you publish full transcripts for accessibility, budget for a human proofreading pass.
  • You listen and want the text. Use a service that finds episodes by name or RSS so you skip the download. For a single episode, a free tier is usually enough. More on the options in the podcast transcript generator guide.
  • You're a journalist quoting a published interview. An AI transcript to find the passage, then check the quote against the audio. The best transcription software for interviews is whatever gets you to the exact timestamp fastest, which is why the .srt matters.
  • You're researching across many episodes. Cost per hour dominates, so use pay-as-you-go AI or DIY Whisper, then keep the files somewhere searchable. I wrote up that workflow in podcast transcript search.
  • You need legal or verbatim accuracy. Pay for human transcription. At $119.40 per hour from Rev, it's expensive and the right call.

The best transcription services aren't the ones with the longest feature list. They're the ones that get the right audio, transcribe all of it, and charge a rate that fits how often you use them.

FAQ: podcast transcription service

What is the best podcast transcription service?

It depends on the job. Human services (Rev, Happy Scribe proofreading) are best when a transcript is legal or published verbatim. AI services are best for notes, research and show notes. DIY Whisper is best when you transcribe a lot and are comfortable with scripts. Price per audio hour and how the service gets the audio matter more than brand.

Is there a free podcast transcription option?

Yes, with limits. Sonix gives 30 free minutes, Otter's free plan allows three file imports in total, and Podskrift gives 180 free minutes with no card. Running open-source Whisper on your own computer has no per-minute fee at all, but you supply the audio file and the hardware.

How much does podcast transcription cost per hour?

As of October 2026: about $119 to $120 per hour for human transcription (Rev, Happy Scribe), $10 to $12 per hour for pay-as-you-go AI (Sonix, Happy Scribe top-ups), $1 per hour for a Podskrift pack, and $0.36 per hour if you call OpenAI's whisper-1 model with your own key.

Is AI podcast transcription accurate enough to quote?

For clear conversational audio, usually yes as a draft. Check names, numbers and anything you publish as a direct quote against the audio. Crosstalk, heavy accents and a wrongly set language are where AI transcripts fail most.

What is the best transcription software for interviews?

For your own recorded interviews, use a tool that accepts file uploads and labels speakers, such as Sonix, Happy Scribe or Descript. For interviews that were published as podcast episodes, a tool that finds the episode by name or RSS saves the download step.

Built by me · 180 min freePodskriftPodcast to Text TranscriptionSee how it works →

Stay in the loop

Productivity tips, tool reviews and early access to new products, straight to your inbox.