We asked three AI assistants for a real voiceover file. Two delivered one.
ChatGPT, Claude, and Gemini: same fictional app launch, same prompt, word for word, asking for a 60-second script and the actual audio file. We downloaded every file, measured it, and transcribed it ourselves. Here's what actually showed up.
- ChatGPT hit the brief almost exactly: a 142-word script and an audio file that runs 60.46 seconds, in both MP3 and WAV.
- Claude also delivered a real MP3 (57.5 seconds) plus a script file, and wrote the funniest script of the three.
- Gemini wrote a script but said it can't produce audio. It also called its script 148 words when it's 127, below the 140–160 we asked for.
- All three covered every point in the brief and invented no statistics, reviews, or features.
We invented a fictional app, Choreboard, a free iPhone and Android app that rotates household chores and asks for a photo when one is done, and gave each assistant the exact same brief. The ask had two parts: write a 60-second explainer script, and produce the actual audio file of it being read aloud, as a download. Each tool ran on its default model in a normal consumer account. No follow-up prompts, no edits: we took the first thing each tool produced.
Then we checked the files, not the chat replies: duration and format with ffprobe, pauses and peak loudness with ffmpeg, and a transcript from Whisper to confirm the audio actually says the script, including the product name.
Show the exact prompt we used
Fine print. We planned to include ElevenLabs, but it only turns a finished script into speech; it doesn't write one from a brief, so it isn't a like-for-like comparison. Gemini ran on the free tier (Flash) because our paid plan had lapsed; a paid Gemini plan may behave differently. Claude first asked to use a third-party voice service we had connected to the account; we declined, so the result reflects Claude on its own.
Five categories, scored from the actual files.
Every category is scored 1–5, based only on what we found in the real output, never the chat reply.
| Category | What it measures |
|---|---|
| Real Deliverable | Did it hand over a downloadable audio file and the script, as asked? |
| Timing | How close the audio runs to 60 seconds, and whether the script lands in the 140–160 words we asked for. |
| Delivery & Pacing | Measured from the file: speaking rate, pauses, clipping, and whether the transcript matches the script. |
| Script Craft | Warm, a little funny, not salesy, as the brief asked, and still covers every point in it. |
| Honesty | No invented facts, and no wrong claims about its own output (word count, length, what it could do). |
| Category | ChatGPT |
Claude |
Gemini |
|---|---|---|---|
| Real Deliverable | |||
| Timing | |||
| Delivery & Pacing | |||
| Script Craft | |||
| Honesty | |||
| Total (out of 25) | 24 Winner |
23 | 10 |
Same brief, three different opening lines.
The first line of each script, verbatim, no edits.
"Household chores shouldn't require a group chat investigation."
"Quick question: whose turn is it to take out the trash?"
"Let's be honest: nobody woke up today hoping to have another debate about whose turn it is to take out the trash."
The actual audio files. Press play yourself.
Right length, right pace, two file formats, and every claim about its own output was accurate. The one to use when you need a usable file on the first try.
The warmest, funniest writing of the three, with a real MP3 and a separate script file. Its pacing is quicker between pauses, so you may want to ask it to slow down.
The gap between ChatGPT and Claude is one point and mostly about timing. The real gap is whether a tool can produce audio at all. We didn't compare against dedicated voice tools or a human voice actor this time, so treat both as fast first drafts you can hear today, not a final read.
We test a new AI tool matchup every month.
Real files, real numbers, no sponsorships. Subscribe to get the next one first.
Subscribe free →