This website uses cookies

Read our Privacy policy and Terms of use for more information.

THE AI PROFIT WIRE

Issue #11 | July 25, 2026 | Weekly Intelligence Briefing

Last week handed you two reversals. This week handed you a receipt.

Anthropic shipped Claude Opus 5 on Thursday at $5 per million input tokens and $25 per million output, the exact rate of the model it replaces, while more than doubling that model's Frontier-Bench score. The pre-launch rumor had it at $10 input and $50 output. For the second week running, the correction ran in the buyer's favor.

The day before that, the same company published a report saying DeepSeek, Moonshot, MiniMax, and Alibaba's Qwen team ran roughly 49,000 fraudulent accounts to pull nearly 45 million exchanges out of Claude, then trained cheaper models on the reasoning they harvested. Anthropic calls the Qwen operation the largest known distillation attack against its models to date, and Congress is reportedly weighing whether to make this kind of distillation a sanctionable offense.

Read those two together and the week has a shape. The frontier is getting cheaper per unit of capability, on the record, with a published price list you can plan against. The discount tier underneath it now carries a provenance question that could turn into a supply problem. Google poured fuel on both ends by dropping Gemini 3.5 Flash-Lite at $0.30 per million input tokens while beating the previous generation on coding benchmarks.

Then Princeton and the University of Chicago published the part nobody wants on the invoice. In a 40-round hiring simulation, OpenAI's o3 scored 1.83 on a segregation scale where 2.0 means complete segregation, against 0.84 for human players in the original study. Instructing the models to be fair changed nothing at all.

Capability up, sticker price flat, provenance and judgment both open. Five signals made the cut, plus the Hype Check Spotlight and one free tool worth your Monday. The rest are in the wire below.

Source: Anthropic

What happened:

Anthropic released Claude Opus 5 on July 24 at $5 per million input tokens and $25 per million output, identical to Opus 4.8. It's generally available today across Claude.ai, Claude Code, Claude Cowork, and the API under the model ID claude-opus-5, and it becomes the default model on Claude Max and the strongest model on Claude Pro. Anthropic positions it as approaching Claude Fable 5's intelligence at half the price, while noting it stays behind Mythos 5 on cybersecurity tasks and long-running autonomous biology research.

What the data says:

Anthropic published results across nine named evaluations. On Zapier's AutomationBench, which scores whether a model completes a business task start to finish, Opus 5's pass rate runs around 1.5 times the next-best model at the same cost per task, and even at its lowest effort setting it clears more tasks than anything else tested. On Frontier-Bench v0.1 it more than doubles Opus 4.8's score at a lower cost per task, on ARC-AGI 3 it scores three times the next-best model, and on OSWorld 2.0 it passes Fable 5's best result at a little over a third of the cost. The number that actually sets your bill is tokens burned per finished task, not sticker rate: Artificial Analysis measured Claude Fable 5 inside Claude Code at $11.80 per completed agentic task against $2.49 for Grok 4.5, a gap driven entirely by token burn. Early partners report Opus 5 moving the right way there, including 26% fewer tokens on legal agent work at max reasoning and under half the latency of Opus 4.8.

Hype Check: 7.8/10

Every performance figure published so far traces back to Anthropic or the 22 partners it hand-picked for early access, which is the only thing holding this score down.

Business impact:

→ Treat this as a supply-chain event, not a purchase decision. Almost no small business calls the Claude API directly, so your real exposure runs through the tools already sitting on your monthly invoice.

→ Watch release notes from your existing AI vendors over the next 30 days. A flat price with double the capability historically surfaces as feature upgrades inside current subscriptions rather than price hikes.

→ Hold the cost claims until an independent per-task measurement lands. Treat the capability numbers as credible and the savings numbers as unverified, because a flat sticker price only helps if token burn per finished task stays flat with it.

Read the full signal.

Source: Anthropic

What happened:

Anthropic says three Chinese AI labs ran an industrial-scale operation to strip-mine Claude's reasoning. Per its own report, DeepSeek, Moonshot, and MiniMax operated roughly 24,000 fraudulent accounts generating more than 16 million exchanges with Claude, with MiniMax alone responsible for about 13 million. Weeks later it escalated, saying Alibaba's Qwen team ran nearly 25,000 fake accounts to pull 28.8 million exchanges between April 22 and June 5, which Anthropic calls the largest known distillation attack against its models to date.

What the data says:

Anthropic doesn't sell commercial access to Claude inside China, and says the accused labs routed traffic through commercial proxy resellers so requests looked like ordinary retail logins from permitted regions, shifting workloads to fresh accounts whenever one hit a rate limit. The detail that separates this from ordinary research use is the target: Anthropic says the labs specifically prompted Claude to walk through its chain-of-thought reasoning on hard math, code, and multi-step agent tasks rather than asking for final answers, because that reasoning trail is the training data you need to build a model that thinks the same way without redoing years of research. Distillation itself isn't the scandal, every lab does it, and it's how Claude Haiku, GPT's mini tiers, and Gemini Flash exist. This is Anthropic's account of what its own detection systems found, corroborated by outside reporting from CNBC and Reuters, and not a courtroom finding.

A rock-bottom price on an AI tool is not proof of a leaner cost structure, and treating it that way without checking the vendor's data practices is a bet you're placing on facts you can't see.

Business impact:

→ Ask your vendor where the training data came from and whether they'll say so in writing. A vendor that answers specifically is behaving differently than one that deflects, and that difference is free information.

→ Test any deeply discounted model on your messiest real task before you trust it on your cleanest one. A distilled model can inherit blind spots from its teacher or skip safety tuning the original lab spent months on, and that gap only shows up on an edge case with a live customer.

→ Track the regulatory fight even if you never touch a Chinese model. If lawmakers move to sanction this kind of distillation, any vendor built on disputed training data becomes a business continuity risk, not an ethics debate.

Read the full signal.

What happened:

Google released Gemini 3.6 Flash and Gemini 3.5 Flash-Lite, both aimed at token efficiency and latency for production agents rather than at leaderboard position. 3.6 Flash is priced at $1.50 per million input tokens and $7.50 per million output. 3.5 Flash-Lite lands at $0.30 input and $2.50 output while running at 350 output tokens per second, and both are available in the Gemini API today.

What the data says:

3.6 Flash consumes 17% fewer output tokens than 3.5 Flash on the Artificial Analysis Index, which cuts cost per agentic task without touching the rate card. Flash-Lite is the bigger surprise, beating 3 Flash on SWE-Bench Pro at 54.2% against 49.6%, scoring 54% on Terminal-Bench 2.1 against 31% for the previous generation, and hitting 74.0% on OSWorld-Verified against 65.1%. That's a model at a fifth of 3.6 Flash's input price clearing coding and computer-use benchmarks that premium models were failing a generation ago.

The expensive mistake in most AI stacks isn't picking the wrong model, it's running every task through one model when the price spread between tiers is 5x.

Business impact:

→ Route high-volume document processing, data extraction, and classification to 3.5 Flash-Lite at $0.30 per million input tokens. These are the jobs where you're paying premium rates for basic reasoning.

→ Keep multi-step reasoning on 3.6 Flash, where a wrong first pass costs you a human correction. The 17% token reduction compounds across every tool call in a long workflow.

→ Audit which model your existing automations call before you renew anything. Most n8n, Zapier, and Make workflows were configured on whatever was current at build time and never revisited.

Read the full signal.

What happened:

Researchers ran ChatGPT, Claude, and Gemini through a 40-round hiring simulation adapted from a classic psychology study, telling each model it was a consultant hired by a fictional city to fill 20 different jobs from four fictional ethnic groups. Every candidate held an equal chance of success. The models still divided applicants into roles by group, and they did it from their own operational experience rather than from anything in their training data.

What the data says:

OpenAI's o3 scored 1.83 on the study's segregation scale, where 2.0 means complete segregation. Human participants in the original version of the game scored 0.84. When a model learned that one candidate from the Aima group failed as a doctor, it stopped hiring Aimas as doctors entirely and started routing them to janitor roles it classified as requiring less warmth and competence. Newer reasoning models were worse, not better, with o3 and DeepSeek's R1 both showing stronger bias, and Cornell's Angelina Wang notes that memory features accelerate this because the model over-indexes on its own previous behavior. Two interventions got tested: instructing the models to be fair did nothing, because the drive to optimize for correct hires swallowed the instruction, while promising an explicit bonus for diverse hiring cut the bias sharply.

A fairness instruction is not a control, and any screening workflow that relies on one is running unguarded while looking supervised.

Business impact:

→ Stop treating "be fair and unbiased" in a prompt as a safeguard. The study tested that exact intervention and it produced no measurable change in behavior.

→ Write the diversity objective into the scoring itself if you use AI anywhere in screening. The bonus-incentive version worked because it competed with the optimization goal instead of sitting next to it.

→ Feed the model relevant candidate information like education and experience, and strip irrelevant markers. Relevant detail reduced segregation, and irrelevant detail brought it straight back.

Read the full signal.

Source: OpenAI Blog

What happened:

OpenAI launched a ChatGPT for Small Business program built around product webinars, interactive guides, in-person US events, partner promotions, and curated plugins. Alongside it sits ChatGPT Work, an agent built to run multi-step tasks end to end, powered by GPT-5.6 and available across all subscription plans rather than gated behind an enterprise tier.

What the data says:

OpenAI validates the program with data from its previous OpenAI Academy events, where 78% of participants built a functional AI workflow in a single day and 42% saved more than 5 hours a week. Its published case studies include an owner who saved at least 10 hours of data entry on a single 30-page contractor quote, and another running a chief-of-staff agent wired into Google Drive and Gmail for scheduling. ChatGPT Work integrates natively with Dropbox, Shopify, Intuit, Slack, Atlassian, and Wix, which is what separates it from a standalone chat window: the agent operates inside the stack you already pay for.

Every number in the launch material comes from OpenAI's own events, which makes the integration list more trustworthy than the time-savings claims.

Business impact:

→ Pick one workflow you already run manually every week and time it, then run it through the free guides and time it again. If the agent version doesn't beat your manual time on the second attempt, the bottleneck is your process documentation and not the model.

→ Start with the integration you already pay for. If your books are in Intuit or your store is on Shopify, that connector is where the 10-hour data entry example actually applies to you.

→ Treat the 42% figure as a ceiling and not a forecast. Those were people who showed up to a hands-on AI training event, which is a self-selected group with above-average motivation.

Read the full signal.

Every signal above is about what a model costs or where it came from. This one is about which of these things you should actually be paying for on Monday, because the agentic tier crossed from demo to production this year and the $20 line item on your card now buys a different product than it did in January.

Community adoption has consolidated hard. ChatGPT and Claude are effectively the only two choices for anyone who wants full agentic capability today, and if your workplace runs on Microsoft you may only have Copilot, which lags badly on agentic ability. Chinese open-weight models like Kimi K3, DeepSeek, and Qwen require real expertise to run as agents, which puts them outside the reach of most small teams.

Pricing model is where owners get this wrong. The $20 tiers include real agent usage, and agents burn through those limits fast. Upgrading to a more expensive plan buys you more hours of AI labor, not a smarter model, which means you should be pricing your plan against hours needed rather than against capability.

Benchmark data comes from hands-on work rather than a leaderboard. GPT-5.6 Sol inside Codex chased down 195 references in a book PDF in 30 minutes with zero hallucinated page numbers and no invented text. ChatGPT Work and Claude Cowork each prepped an MBA seminar, building the teaching materials and writing the email in 10 minutes, on work estimated at a couple of hours by hand.

Expert sentiment comes from Ethan Mollick, who runs independent testing for an audience of over 460,000 and has no commercial relationship with either lab. His one substantive complaint is that the systems are nitpicky, which is a filtering problem you solve with human judgment rather than an accuracy problem you can't solve at all.

Release maturity splits by where the computer lives. ChatGPT Work and Claude Cowork run on virtual machines the AI companies provide, so you can start a job from your phone and check the result later. Codex and Claude Code run on your local machine, which is what unlocks complex multi-file projects. Both modes are shipping today across consumer tiers.

Hype Check: 7.2/10

The verdict: you're buying hours of AI labor and not a tireless employee, so budget by hours and verify the output. Leave every permission set to ask for approval before sending, spending, or deleting, because an agent that reads your inbox and browses the web will eventually meet text written by someone trying to trick it.

Read the full signal.

While every lab this week argued about token prices and training data provenance, developer Prince Canuma shipped the answer to both. Nativ is a free, 100% MIT-licensed macOS app that runs AI models directly on Apple Silicon, with no accounts, no subscriptions, and no cloud calls. It requires an M1 chip or newer, ships with a chat interface and a localhost API server, and covers language, vision, video, code, and audio.

Simon Willison tested it on July 21 and confirmed it picked up the MLX models he'd already cached from Hugging Face on first launch. Three verified partner models ship at launch: Google's Gemma 4 E2B at 10.28 GB for vision and audio, Cohere's North Mini Code at 19.38 GB with a 500K context window for code and tools, and Liquid AI's LFM2.5-VL at 3.20 GB for vision and language. The localhost endpoint integrates with five coding agents including Claude Code, and the telemetry surface shows tokens per second, memory pressure, thermal state, and time to first token in real time.

If you already own an Apple Silicon Mac, Nativ closes the API bill and the client-confidentiality question in one download. Run your batch document processing, your sensitive client files, and your internal automations against a local endpoint where the data never leaves the device, and keep the metered API for the work that genuinely needs a frontier model.

Read the full signal.

The Wire: What Else Made the Cut

Google shipped another wall of Workspace updates this week: Gemini Intelligence now automates tasks across 40 apps on the new Samsung Galaxy Z Fold8 and Flip8, Gemini Live camera lets you point a phone at an error code or a manual for real-time walkthrough help, Sheets added native combo charts with better Excel import fidelity for secondary axes, Meet launched a centralized web hub for agendas, notes, recordings, and attachments, Calendar now flags who a busy contact's scheduling delegate is so you can message them directly, and account recovery accepts a short encrypted selfie video as a backup access method. Google is also taking applications for the Gemini Startup Forum, a two-day founder event with its engineers. Here's what else earned a spot outside that flood.

Anthropic confirmed that Claude Tag, its proactive Slack agent, now lands 65% of PRs on the Claude Code product engineering team, and that the same team cut its system prompt by 80% by replacing hard rules with broader context. OpenAI's GPT-5.6 prompting guidance points the same direction, showing leaner prompts improve eval scores 10 to 15% while cutting tokens 41 to 66% and cost 33 to 67%. Both frontier labs converged on this independently, which makes rewriting your longest agent prompt as context instead of rules the cheapest optimization available this month. Full signal

OpenAI added voice control to the ChatGPT desktop app, letting you drive agents and multi-step tasks hands-free through ChatGPT Work and Codex, which matters most for anyone working with their hands while a job runs. Full signal

Substack rolled out AI detection in its app that estimates how much of a newsletter was machine-written, pushing writers toward disclosure. If you publish anything, assume reader-facing AI attribution is coming to your platform next. Full signal

Synthesia launched AI Roleplay Sessions where employees practice high-stakes conversations against avatars that push back, give feedback, and score them against a rubric, turning soft-skills training into something you can actually measure. Full signal

Bluesky upgraded its assistant Attie with Quests, a free research feature for tracking trending topics and influential accounts across the network, useful for anyone testing the open social web as a distribution channel. Full signal

Palmier Pro arrived as a free, open-source macOS video editor with AI transitions and long-video-to-shorts cutting built in, a second zero-cost option this week for founders already on Apple hardware. Full signal

Two warnings worth your attention. AI agents remain unreliable at demand forecasting because they lack statistical feedback loops and hallucinate seasonal trends, so keep the forecast itself in a dedicated statistical model and let the agent handle data prep and scenario planning around it. Full signal And automated file-download agents save real hours logging into systems and pulling documents on a schedule, but a loose configuration serves you stale files and opens access holes you won't notice until an audit. Full signal

The full week's signals, detailed breakdowns, and action items are on the site. If this issue earned its place in your inbox, forward it to whoever signs off on your AI budget.

This issue went out to subscribers Saturday. If you want next week's before it hits the web, subscribe at metadatamarketer.com/subscribe

Test. Cut. Share.

Moe Sbaiti, The AI Profit Wire
https://metadatamarketer.com

Reply

Avatar

or to participate

Keep Reading