# Podcast Transcription Services Compared: Human vs AI vs DIY Whisper (Real Cost per Hour, 2026)

Compare every podcast transcription service type: human, AI and DIY Whisper. Real cost per hour from current pricing pages, plus the failure modes to test.

Topic: Podcasts
Published: 2026-10-03
URL: https://productivitytech.io/podcast-transcription-service/

A **podcast transcription service** turns an episode's audio into text, and in 2026 you choose between three kinds: human transcriptionists at about $120 per audio hour, AI services that cost roughly $1 to $12 per hour, and DIY Whisper at $0.36 per hour through OpenAI's API. Price is the easy part to compare. Whether the transcript is complete, and whether the service can even get the audio, is harder to see on a pricing page.

I built a podcast transcription service, [Podskrift](https://podskrift.com/?utm_source=productivitytech&utm_medium=site&utm_campaign=podcast-transcription-service), so I've seen the failure modes from the inside. This guide compares the three types with prices taken from each vendor's pricing page in October 2026, and shows the checks I now run on any transcript. *Disclosure: Podskrift is my own product, and it's in the comparison.*

## What a podcast transcription service actually does (human vs AI vs DIY)

Every service does the same three things: get the audio, convert speech to text, and give you a file. They differ in who does the second step.

-   **Human transcription.** A person types the audio, sometimes on top of an AI draft. Rev [guarantees 99%+ accuracy](https://www.rev.com/pricing) for its human service. You pay per minute, and you wait hours rather than minutes.
-   **AI transcription.** A speech model writes the transcript, and the service wraps it in an editor, speaker labels and exports. Most "podcast transcript AI" tools you find in search are this type, built on a model like OpenAI's Whisper.
-   **DIY Whisper.** You run the model yourself, either the [open-source Whisper](https://github.com/openai/whisper) on your own machine or OpenAI's API with your own key. You handle downloads, file limits and formatting.

The type you need depends on what happens to the text afterwards. A court filing needs a human. Notes from an episode you listened to on Tuesday don't.

## Real cost per hour of audio, compared

Pricing pages mix per-minute rates, monthly plans and "media hours", so I converted everything to dollars per hour of audio. Prices were checked on each vendor's pricing page on October 3, 2026. Monthly-billed prices are shown; most vendors are cheaper billed annually.

| Service | Type | Published price | Cost per audio hour |
| --- | --- | --- | --- |
| [Rev](https://www.rev.com/pricing) | Human | $1.99/min | $119.40 |
| [Happy Scribe](https://www.happyscribe.com/pricing) | Human proofreading | from $2.00/min | from $120.00 |
| [Happy Scribe](https://www.happyscribe.com/pricing) | AI top-up credits | $0.20/min | $12.00 |
| [Sonix](https://sonix.ai/pricing) | AI, pay as you go | $10/hr | $10.00 |
| [Descript](https://www.descript.com/pricing) Hobbyist | AI editor, subscription | $24/month for 10 media hours | $2.40 if you use all 10 hours |
| [Podskrift](https://podskrift.com/?utm_source=productivitytech&utm_medium=site&utm_campaign=podcast-transcription-service) pack | AI | $5 for 300 minutes | $1.00 |
| OpenAI [`whisper-1`](https://developers.openai.com/api/docs/models/whisper-1) with your own key | DIY API | $0.006/min | $0.36 |
| Open-source Whisper on your computer | DIY local | No per-minute fee | Your hardware and time |

Otter isn't in the table because it isn't priced per hour. Its [Pro plan](https://otter.ai/pricing) is $16.99 per user per month and allows 10 imported audio or video files per month; the free plan allows 3 imported files in total. It's built for meetings, so an imported podcast episode counts against a file limit rather than a minute rate.

Three things stand out:

1.  **Human transcription costs about 100 times more than AI.** It's worth it when every word carries legal or editorial weight. For notes and research it isn't.
2.  **Subscriptions are only cheap if you use them.** Descript's $2.40 per hour assumes you use all 10 hours every month, and Descript is a video and audio editor first.
3.  **The model itself is cheap.** A one-hour episode through `whisper-1` costs OpenAI $0.36. Most of what you pay a podcast transcription service for is everything around the model.

For per-episode costs at Whisper's rate, see the table in my [free podcast transcription tool](/free-podcast-transcription-tool/) post. Podskrift's own pricing: the first 180 minutes are free with no card, then a one-time $5 pack covers 300 minutes, or you add your own OpenAI key and pay OpenAI directly.

## Accuracy: where AI transcription breaks

On a clean studio recording, a modern AI transcript is usually good enough to read and search without edits. It breaks in predictable places:

-   **Crosstalk.** Two people talking at once often come out as one merged sentence, and speaker labels get swapped.
-   **Names and jargon.** Guest names, product names and acronyms are the most common errors. Check every one you plan to quote.
-   **The language setting.** This one cost me money. An early version of Podskrift defaulted the language to Norwegian. A listener transcribed a Japanese episode without changing it and got phonetic nonsense back, and we paid for it. Whisper doesn't treat the language you pass as a hint. It treats it as a constraint and forces the audio into that language.

The fix for the last one is simple: leave the language on auto-detect unless the clip is short, heavily accented or mixes languages, and never set a language the episode isn't in. Whichever podcast transcription service you use, find out what its language setting actually does before you run a batch.

If you need guaranteed accuracy, that's the case for human transcription or a human proofreading pass over an AI draft. Happy Scribe sells exactly that combination.

## The completeness problem nobody mentions

Accuracy gets all the attention, but the worse failure is a transcript that is accurate and incomplete. Nothing looks wrong. The text reads fine. A chunk of the episode just isn't there.

I hit two causes while building Podskrift's audio pipeline:

1.  **File size isn't proportional to time.** OpenAI's [speech-to-text API](https://developers.openai.com/api/docs/guides/speech-to-text) accepts files up to 25 MB, so long episodes have to be split. Podcast audio is often variable bitrate, so splitting by bytes doesn't give equal stretches of time. One episode with a dense first half produced a 27 MB piece against a 24 MB target, over the API limit.
2.  **The reported length can be wrong.** Many podcasts insert ads dynamically, which stitches the file together from separate segments. On one such file, the duration the file reported was shorter than the real audio. An early version of Podskrift trusted that number, and 16 minutes were never sent for transcription. No error was raised.

The fix for both was to stop trusting file metadata: Podskrift now converts audio to 16 kHz mono and splits it into 15-minute pieces by walking the actual audio stream. You can't see any of this from outside a service, so test the output instead.

### The last-timestamp test

Download the `.srt` version of the transcript and compare its final timestamp with the episode length shown in your podcast app (not the length the audio file reports, which is the number that can be wrong):

```bash
# Print the start and end time of the last subtitle in an .srt file
grep -E '^[0-9]{2}:[0-9]{2}:[0-9]{2},[0-9]{3} --> ' episode.srt | tail -n 1
```

If the episode runs 1:12:40 and the last subtitle ends at 0:56:10, part of the audio was never transcribed. A gap of a few seconds at the end is normal, since outros are often music. A gap of minutes isn't. It takes ten seconds, and it works on any podcast transcription service that exports `.srt`.

## The input problem: upload vs RSS vs search by name

Most comparison pages skip the first step: getting the audio into the service. For podcasts, that's often the hardest part.

| Input | Who it suits | The catch |
| --- | --- | --- |
| Upload a file | Podcasters with their own recordings, interviews | Podcast apps don't give you the episode file |
| Paste an RSS feed | Anyone who can find the show's feed | Listening apps don't show the feed URL |
| Search by name or paste a link | Listeners who heard an episode in an app | Spotify-exclusive shows have no public audio |

Upload-first services (Rev, Sonix, Descript, Happy Scribe) are built for people who already have the audio. If you produce the show, that's you. If you heard the episode in Spotify or Apple Podcasts, you first have to find the feed or the file. I built Podskrift's search because of that gap: you search the show or paste an episode link, and it fetches the audio itself. The step-by-step version is in [how to transcribe a podcast to text](/transcribe-podcast-to-text/), and the platform-specific cases are covered in the [Apple Podcasts transcript guide](/apple-podcast-transcript-generator/) and the [Spotify transcript export guide](/spotify-podcast-transcript-export/).

## Which podcast transcription service fits you

-   **You produce the podcast.** Use an upload-first AI service with speaker labels and an editor (Descript, Sonix, Happy Scribe). If you publish full transcripts for accessibility, budget for a human proofreading pass.
-   **You listen and want the text.** Use a service that finds episodes by name or RSS so you skip the download. For a single episode, a free tier is usually enough. More on the options in the [podcast transcript generator](/podcast-transcript-generator/) guide.
-   **You're a journalist quoting a published interview.** An AI transcript to find the passage, then check the quote against the audio. The best transcription software for interviews is whatever gets you to the exact timestamp fastest, which is why the `.srt` matters.
-   **You're researching across many episodes.** Cost per hour dominates, so use pay-as-you-go AI or DIY Whisper, then keep the files somewhere searchable. I wrote up that workflow in [podcast transcript search](/podcast-transcript-search/).
-   **You need legal or verbatim accuracy.** Pay for human transcription. At $119.40 per hour from Rev, it's expensive and the right call.

The best transcription services aren't the ones with the longest feature list. They're the ones that get the right audio, transcribe all of it, and charge a rate that fits how often you use them.

## FAQ: podcast transcription service

<!-- faq -->
