FOWL AI Tests · Monthly · Test #1

We gave four AI tools the same messy resume. Only one checked its own work.

Claude, ChatGPT, Codex, and Claude Code — same fictional candidate, same prompt, word for word. We downloaded every file they actually produced and counted the pages ourselves. Here's the real scoreboard.

Tools testedClaude.ai, ChatGPT, Codex, Claude Code
Test case7-year B2B marketing manager, targeting a senior role
What we checkedReal files, not chat transcripts
Key Findings — 15-Second Version
  • Claude.ai was the only tool that verified its own page length before showing us the result.
  • Codex produced the most detailed, thorough resume of the four.
  • None of the four tools fabricated a credential, certification, or team size.
  • ChatGPT and Codex both quietly ran to 2 pages without flagging it.

We invented a realistic candidate — Jordan Reyes, 7+ years in B2B SaaS marketing, applying for a Senior Marketing Manager, Lifecycle & Retention role — and gave all four tools the exact same unformatted background and the exact same instruction: turn this into a polished, ATS-friendly resume. No follow-up prompts, no edits. We took the first thing each tool produced, downloaded the actual file, and opened it.

PersonaFictional, invented for this test
PromptIdentical across all four, word for word
Output checkedThe real downloaded .docx / .pdf, not the chat reply
VerificationPage counts confirmed in Microsoft Word
Show the exact prompt we used
I need help writing a resume for this target role: Senior Marketing Manager, Lifecycle & Retention at a mid-size B2B SaaS company (50-200 employees). Here's my background, unformatted — please turn this into a polished, ATS-friendly resume: Jordan Reyes jordan.reyes@email.com | (555) 123-4567 | Austin, TX | linkedin.com/in/jordanreyes Work history: - Marketing Manager at BrightPath Software (SaaS, project management tools), Jan 2022–present. Ran email lifecycle campaigns, grew trial-to-paid conversion from 11% to 17% over 18 months. Managed a $400K/year paid marketing budget across Google and LinkedIn ads. Built and launched a customer re-engagement flow in HubSpot that recovered about $180K in at-risk annual recurring revenue. Managed one direct report (Marketing Coordinator). - Senior Marketing Associate at Lumen Analytics (data analytics startup, ~40 people), Jun 2019–Dec 2021. Owned content marketing calendar, grew organic blog traffic from 8K to 47K monthly visitors in 2 years. Ran A/B tests on landing pages that improved demo request conversion by 22%. Coordinated with sales to build lead scoring criteria. - Marketing Coordinator at Retail Insights Co (retail SaaS), Aug 2017–May 2019. Supported event marketing (ran 3 regional conferences, 200-400 attendees each), managed social media accounts, wrote weekly newsletter to 15K subscriber list. Education: BA in Communications, University of Texas at Austin, 2017. Skills I use regularly: HubSpot, Google Analytics, Google Ads, LinkedIn Campaign Manager, Figma (basic), SQL (basic, can pull simple queries), ChatGPT/Claude for campaign copywriting and content briefs, Asana, Looker (dashboards, not building them). I don't have a formal certification in anything. I've never managed a team bigger than 1 person. I'm applying because I want to move from a smaller company into a role with more strategic ownership over retention/lifecycle at a bigger SaaS company. Please write the full resume.

Five categories, scored from the actual files.

Every category is scored 1–5, dot for dot, based only on what we found in the downloaded file — never the chat reply.

5 = Excellent 4 = Good 3 = Acceptable 2 = Weak 1 = Poor
CategoryWhat it measures
Content DepthHow much of the candidate's real work and impact made it into the resume, without padding.
Structure & ATS-SafetyReal selectable text, no tables/text boxes/columns, contact info in the body — the things that actually break applicant tracking systems.
Page-Fit VerificationWhether the tool checked that its own output actually matched the stated length before handing it over.
Visual PolishTypography, spacing, and design choices in the real file.
HonestyWhether the tool invented any credential, certification, or scope of experience not present in the source material.
Category
Claude.ai
ChatGPT
Codex
Claude Code
Content Depth & Coverage
Structure & ATS-Safety
Page-Fit Verification
Visual Polish
Honesty (no fabricated credentials)
Total (out of 25) 23
Winner
16 18
Depth Winner
16

Same person, four completely different opening paragraphs.

The "Professional Summary" each tool wrote from the identical source material — verbatim, no edits.

Claude.ai

"Lifecycle and retention marketer with 8+ years across B2B SaaS, focused on turning customer journeys into measurable revenue. Owns email lifecycle programs, trial-to-paid conversion, and churn recovery end to end..."

ChatGPT

"Results-driven SaaS marketing professional with 8+ years of experience leading lifecycle marketing, demand generation, content strategy, and paid acquisition initiatives for B2B software companies..."

Codex

"Lifecycle and growth marketing manager with 7+ years of B2B SaaS experience across retention, trial conversion, marketing automation, paid acquisition, content, and funnel optimization..."

Claude Code

"Marketing Manager with 7+ years in B2B SaaS, specializing in lifecycle marketing and retention. Grew trial-to-paid conversion 6 points (11% → 17%) and recovered $180K in at-risk ARR..."

The actual files, not screenshots of a chat window.

Only one tool checked its own work
Claude.ai was the only one of the four that rendered its own file, noticed a problem (Jordan's resume was spilling onto a second page), and fixed it before handing it over. ChatGPT and Codex both quietly produced 2-page resumes and said nothing. That's not a writing-quality gap — it's a verification gap, and it's the kind of thing you'd only catch by opening the actual file.
Codex went the deepest — and that's a real, legitimate win
Codex's resume covers more of Jordan's actual work than any of the other three: 7 bullets at the current role versus 4, a dedicated Tools section separate from soft skills, and a subtitle that mirrors the exact target job title word for word. If you value thoroughness over brevity — and plenty of good career advice says two pages is fine past ~7 years of experience — Codex is the strongest single output here.
Nobody fabricated a credential
This is the genuinely reassuring result: all four tools were told explicitly that Jordan has no certifications and has never managed more than one direct report, and all four respected that. The only shading we found was ChatGPT listing "Team Leadership" as a core competency — arguable, not a fabrication. No tool invented a degree, a cert, or a team that doesn't exist.
🏆 The Verdict
Best for depth
Codex

Most thorough coverage of the actual work, sharpest role targeting. Check the page count yourself before you send it.

Best for polish & self-checking
Claude.ai

The only one that verified its own output. Best visual design of the four. Trimmed some detail to make it fit.

There isn't one universal winner here — the honest finding is that these four tools optimized for different things without telling us. If you want the fuller picture, tell your tool explicitly to check the page count and keep the detail — none of them will do both by default yet.

Should you switch tools?
Use Claude.ai if you want polished, verified output with fewer surprises.
Use Codex if you prefer maximum detail and don't mind reviewing the formatting yourself.
Use ChatGPT if you want a solid, balanced draft but plan to proofread the final document anyway.
Use Claude Code if your workflow is already code-first and you don't need advanced document design.

We test a new AI tool matchup every month.

Real files, real numbers, no sponsorships. Subscribe to get the next one first.

Subscribe free →