Fable 5.1 leads, but its top setting costs more

Its highest score came at maximum effort, while a lower setting narrowed the cost trade-off and AWS adds a test route.

In partnership with

SPD-BEEHIIV:3b548401-4e22-4279-9290-4d38fa699f7c:R9:8d417a90da9198d05397291b
Its highest score came at maximum effort, while a lower setting narrowed the cost trade-off and AWS adds a test route.
Superpower DailyRead online/Account
Daily issue / Second-Order EffectsThursday, September 3, 2026
Our toolsSuperpower ChatGPT/WFH.team/Snipman

Today's briefing

What matters today

AI’s gains are arriving with attached work: Fable’s new benchmark lead costs more at its top setting, while repair, oversight, and safeguards surface elsewhere. From freelancer cleanup to robotaxi explanations and bot-to-bot hiring, the day’s updates put downstream details in focus.

Inside today's briefing
01Freelancer.com says listings for AI-output corrections rose 87% to 10,760, but listings do not show completed work or earnings.
02MIT and Motional’s planning tool exposed a robotaxi path that would have hit a cyclist before emergency braking intervened.
03One job seeker had ChatGPT interview an AI recruiter after five screenings produced no follow-up; the company did not comment.
04Meta’s glasses now stop recording when the capture light is covered after filming begins, though other workarounds remain unclear.
â–¶
Listen to this newsletterAudio edition / About 2 min↗
Anthropic’s Fable 5.1 Takes Benchmark Lead, but Its Top Setting Costs More Per Task

Lead story / benchmark

Fable 5.1 leads a benchmark at a higher cost

Anthropic’s Fable 5.1 has taken the top measured position on Artificial Analysis’ Intelligence Index, with a score of 66 at its maximum effort setting. Artificial Analysis said that was the highest score it had recorded, but estimated the run cost $3.76 per benchmark task, compared with $3.14 for Fable 5 at maximum effort.

The result varies sharply by effort level. Across five settings, Fable 5.1 scored from 58 to 66, while output-token use ranged from 13.1 million to 143.7 million tokens per Intelligence Index task. The top score is therefore a result from the most token-intensive setting, not a default measure of cost or behavior across applications.

Anthropic kept Fable 5’s listed input, output, and cache-write prices, while cutting cache-read pricing to $0.25 per million tokens. Artificial Analysis attributed the higher maximum-effort task cost chiefly to roughly 1.7 times as many output tokens as Fable 5, despite an estimated $1.40 per-task cache saving. At xhigh effort, Fable 5.1 scored 65 at an estimated $2.72 per task.

The benchmark lead is not decisive on every subtest: Artificial Analysis said Fable 5.1 and Opus 5 had overlapping GDPval-AA v2 confidence intervals and were effectively tied on AA-Briefcase. Safeguards and routing are also part of the measured result, with fallback to other Claude models accounting for about 4% of output tokens in the evaluation. Fable 5.1 is now generally available on AWS, giving enterprises a route to test token use, safeguards, and fallback behavior on their own workloads.

Read full story  ↗
 

AI made PMs faster. Multiplayer mode is still broken.

A PM can summarize research, draft a PRD, and mock up a prototype before lunch. The hard part starts when the team has to decide what actually gets built.

Jira Product Discovery gives product teams one place to capture insights, prioritize ideas with consistent frameworks, and build living roadmaps stakeholders can rally around.

And because it’s connected to Jira, the context behind every decision stays with the work—so developers and their agents know not just what to build, but why.

AI helps PMs move faster. Jira Product Discovery helps the whole team build with confidence.

 
WFH.team logoA tool for your workflowWFH.teamA focused feed of carefully selected remote roles and practical work-from-home resources.Remote work, without the noisy job-board scroll
 
Browse remote roles  ↗
Freelancer.com Sees AI-Cleanup Listings Rise 87% as Creatives Shift to Repair Work

research

Freelancers are being hired to repair AI output

Freelancer.com says listings seeking AI-output corrections rose 87% to 10,760 worldwide between August 2025 and June 2026. The work spans rebuilding images, repairing voiceovers and 3D models, and revising prose. Upwork and Fiverr reported growth too, but those measures track listings, gigs, and searches rather than completed projects or earnings.

Continue reading  ↗
Job Seeker Sends ChatGPT to AI Recruiter After Five Interviews Without Follow-Up

culture

A job seeker sent ChatGPT to an AI recruiter

Christopher says he used ChatGPT Voice to impersonate him in a 10-minute call with Riley, an AI recruiter used by Everforth Apex Systems, after five interviews drew no response. A fictional résumé-matched applicant also completed a 23-minute interview without follow-up. The reported tests do not establish Everforth’s criteria or how widespread AI-generated applications are; the company did not respond to WIRED.

Continue reading  ↗
MIT and Motional Build AI That Exposes Robotaxi Planning Errors

research

MIT and Motional built a tool to expose robotaxi errors

CW-Net places readable concepts inside an autonomous-driving planner’s final trajectory choice, generating an explanation in real time rather than after the fact. In a private-track test, it revealed a robotaxi had selected a path that would have hit a cyclist; emergency braking prevented impact. MIT and Motional report less than a 1% change in driving capability after adding the system.

Continue reading  ↗
Drexel Univercity secondary sponsor media

Drexel’s MS in Artificial Intelligence & Machine Learning

Learn More  ↗
Second-Order Effects themed section header

A benchmark lead raises a practical question: what does a higher top-setting cost change in production?

Effect 01launch
Abliteration Removes Refusal Mechanisms for Offensive Cyber WorkRead story ↗
Effect 02pricing change
Fable Cache Becomes Cheaper After 24 Reuses When Gemini MissesRead story ↗
 

Daily tool drop

5 AI tools worth knowing today

Selected for fit, not rank
Monid logoMonidOne key connects AI agents to more than 1,800 APIs across data and content services.Best for / Builders wiring agents to external toolsOpen â†—
Dial logoDialGives agents phone numbers for calls, SMS, iMessage, and inbound verification codes.Best for / Teams building phone-capable agentsOpen â†—
HydraDB OSS logoHydraDB OSSOpen-source graph database for AI memory, ontologies, and agent context on object storage.Best for / Engineers building context-rich AI appsOpen â†—
Doop logoDoopA multiplayer canvas where Claude, Codex, and MCP agents design and review work live.Best for / Product teams co-designing with agentsOpen â†—
Orato logoOratoScores short speaking drills for pacing, fluency, vocabulary, and coherence.Best for / Professionals sharpening spoken deliveryOpen â†—
 
Quick reads
Anthropic’s 80-Environment Reward-Hacking Test Produced Cyber and Safety EvasionsAnthropic’s reward-hacking test produced cyber evasions ↗research
Meta Blocks AI Glasses From Recording When Their Warning Light Is CoveredMeta’s AI glasses stop recording if the warning light is covered ↗security risk
New York City Public Schools Will Block Student AI Use Through Eighth GradeNew York City schools plan to bar student AI use through eighth grade ↗government action
 

The Internet Had a Point

Meme from the stashChatGPT Pet Project vs. Corporate Work
Reddit screenshot from r/ChatGPT. Text reads “for real.” A two-panel meme compares “after 10 hours of work on my pet project with ChatGPT,” showing a relaxed animated blond character, with “after 2 hours of work at my corporate job,” showing a stark black-and-white distressed face. A “Funny” label is visible.
 
From our network. Tools built for the way you work. Useful products from the team behind Superpower Daily.
WFH.team logo
Reader tool / From our teamWFH.teamFind carefully selected remote roles and practical resources for distributed work.Remote work, without the noisy job-board scroll
Browse remote roles  ↗
WFH.team product preview

Reader check-in

Help shape tomorrow's briefing

One click tells us what to keep, improve, or tighten.

01Useful02Interesting03Too long

Prefer one email a week? Get the essential AI moves in the Sunday Weekly Digest.

Superpower Daily tracks the companies, models, products, tools, policy decisions, and cultural shifts moving AI.Follow Superpower Daily
X
LinkedIn
Discord
YouTube
RSS
Follow on Google News
Sponsor Superpower DailyEmail preferences