Hey Reader, Anthropic shipped a new model on June 9. Three days later, the US government export-controlled it and it vanished for everyone, everywhere. It came back July 1. It's called Claude Fable 5, the first model in the Claude 5 family and the first Anthropic has put in a tier above Opus. Anthropic's claim: it beats every model they've ever released to the public, it's state of the art on nearly every benchmark they tested, and its lead grows as tasks get longer. Stripe reported it migrated a 50-million-line codebase in a day, work they estimated at two months for a full team. Those are their numbers, not mine. I run a 12-function marketing operation on Claude for under $300 a month, and I've sat through every "best model ever" launch since GPT-4. The benchmark table never tells you what a model will do for your marketing. One thing does: your own hardest job. So instead of telling you Fable 5 is the best AI yet, I'll give you the 45-minute test that answers it for your work, this week. Two things before we start:
That second one is how this newsletter grows, and I appreciate every forward. 1. What Fable 5 actually is Mythos-class: Anthropic's new model tier above Opus. Fable 5 and Mythos 5 are the same underlying model. Fable is the version with safeguards added for general release; Mythos goes only to vetted cybersecurity and research partners. Four facts worth having before you test:
The claim that matters for marketers is the first one. Most of the marketing work that eats your week is long, not hard: audits, campaign builds, list work, reporting. That's the exact territory where Anthropic says this model pulls away. Which makes it testable. 2. Get access (5 minutes)
Screenshot: the model picker with Fable 5 selected. Pro tip: $50 per million output tokens sounds steep until you do the math on chat work. A 2,000-word deliverable is roughly 3,000 output tokens, about 15 cents. Long autonomous runs cost more because the model burns tokens thinking, so budget a few dollars for the test, not a few cents. 3. Pick the job (this decides everything) Long-horizon task: work that takes many dependent steps over hours, where step 12 only comes out right if steps 1 through 11 did. The wrong test is a clever question. Every frontier model answers clever questions well now, so you learn nothing. The right test is a job you'd book a full day for. It qualifies if it meets all four:
From a marketing operation, jobs that qualify: a full content audit (pull every post from last quarter, score each against your positioning, deliver a kill-keep-redo list), a campaign built from a one-page brief through emails, landing copy, and ad variants that hold one argument, or a competitor repositioning teardown across five sites with quotes as receipts. Pro tip: pick a job you've already done manually. Your past output is the ruler. Without it you're grading on vibes. 4. Run the same brief on both models One rule: identical brief, your current model and Fable 5, and you don't steer either one midway. Here's the brief template. Fill the brackets and paste the same text into both: You're my [role: content lead / campaign manager / marketing analyst]. The job: [the whole job, not the first step. "Audit all 34 blog posts we published in Q2 against the attached positioning doc."] The material: [paste or attach everything: positioning doc, the posts, the campaign brief, last quarter's numbers. Nothing to invent.] The standard: [what a pass looks like. "A senior contractor would hand me a ranked list with one reason per item, and would flag anything that contradicts our positioning."] Do not: [your banned moves. "No generic advice. No inventing data. If a number isn't in the material, don't use one."] The format: [deliverable, length, structure. "One table, then a 300-word recommendation."] Work through the whole job before showing me anything. If information is missing, make the call, then list your assumptions at the end instead of stopping to ask. If you read my June email on prompts versus briefs, this is that structure with one addition: the last paragraph. Long-horizon models do their best work when you let them finish. Run it on your current model first. Then a fresh chat on Fable 5. Then walk away and let them work. Screenshot: the two outputs side by side. 5. Score it like a manager, not a fan When both are done, score each output with this: Score each output 1-5 on: 1. Completion: how much of the job got done without me? 2. Judgment: were the midway decisions ones I'd have made? 3. Edits: how few minutes to make this shippable? (5 = ship as is) 4. Trust: would I run this monthly without checking every line? Decision rule: if Fable 5 beats your current model by two or more points total, move that job, and every job shaped like it, onto Fable 5 and pay the credits. If it doesn't, you just saved yourself the upgrade and a month of second-guessing. Either result is worth the 45 minutes. That's the point of testing instead of reading takes. Your first 45 minutes with Fable 5 Minutes 0-10: turn on usage credits, open a Fable 5 chat. By minute 45 you know something no launch post could tell you: whether the most capable model Anthropic has ever shipped is the best model for your marketing. Where this falls short Honesty, as always:
One more thing. Everything above is the same test I run on my own operation every time a model ships. Anthropic doesn't pay me and I earn nothing if you choose Fable 5. I run marketing on these models every day and report what holds up. If you run the test and want a second pair of eyes on which of your marketing functions to move onto it, grab a free call with me here: https://calendar.app.google/BCm78pVrBnPL3pWf8. It's the working-session version of this guide, on your operation instead of a hypothetical. And if this saved you a research afternoon, forward it to one marketer who needs the test. — Louis P.S. The single first step: pick the job before you touch the model. The test is only as good as the job you choose. |
Practical AI marketing for marketers who want real systems, not theory. I run a 12-function marketing operation solo for under $300/month. Weekly: workflows, prompts, and tools you can implement.
Hi Reader, I'll say what I always say: Claude is the better tool for marketing. Better brand voice consistency. Better at following constraints. Better at holding a style guide across a long conversation. I use it every day. I recommend it to every client. But here's the reality: I talk to marketing leaders every week, and many are on ChatGPT. Budget decisions, IT approvals, enterprise contracts — the tool is already locked in. So I stopped arguing about which AI is best and started asking a...
Hey Reader, Most people think AI tools are neutral. Put a prompt in, get an answer out. The tool is just a tool. That's not how it works. Every major AI company was built by someone with a specific worldview — about what intelligence should do, what it should refuse, what it should prioritize when it has to choose. Those beliefs don't live in a mission statement. They get encoded into the training data, the guardrails, the reward functions, the things the model will and won't say. The output...
Hey Reader, Quick question: when you open Claude, which model do you pick? Most marketers I talk to pick Sonnet or Opus by default and leave it there. A few go straight to Opus for everything because it feels like the "safe" choice. Both are costing them — in tokens, in speed, and in output quality on tasks the model wasn't built for. The thing is that picking the wrong tool is like using a forklift to move a coffee cup or hiring a lawyer to write a grocery list. I run a 12-function AI...