This website uses cookies

Read our Privacy policy and Terms of use for more information.

THE AI PROFIT WIRE

Issue #12 | August 1, 2026 | Weekly Intelligence Briefing

Last week the story was where cheap models came from. This week the price fell again, and something started reading your documents.

OpenAI cut GPT-5.6 Luna's API price by 80% on Thursday, dropping input to $0.20 per million tokens and output to $1.20. That's a fifth of Claude Haiku 4.5's input rate and roughly a quarter of its output rate. The cut isn't promotional. GPT-5.6 Sol autonomously rewrote OpenAI's production GPU kernels in Triton and Gluon, cutting end-to-end serving costs 20% and making the new floor structural. A day later, DeepSeek shipped V4 Flash at $0.14 input and $0.27 output, and Simon Willison called it possibly the best value-per-intelligence model available right now.

Then the other side of the ledger showed up. Security researchers published a 144-day coordinated disclosure with Microsoft demonstrating a self-propagating worm in Copilot for Word. Hidden white-on-white text in a shared document instructs Copilot to halve financial figures and copy the attack into whatever it drafts next, which turns your own internal documents into carriers. Microsoft shipped mitigations across those 144 days, including a model upgrade. Researchers reproduced the full attack chain on GPT-5.6 at publication.

The same week, Claude shared chats and Artifacts turned up indexed on Google, with medical records, internal company documents, and children's names and phone numbers among them. Anthropic's position is that a share link is a public link, which is technically correct and operationally useless if nobody on your team knows it.

And underneath all of it, Similarweb data shows Google AI Overviews now appearing in 43% of all searches, up from 15% a year ago. The click that used to pay for your content is being intercepted at the search box.

Cheap inputs, expensive failure modes. Five signals made the cut, plus the Hype Check Spotlight and one tool worth your Monday. The rest are in the wire below.

What happened:

OpenAI cut GPT-5.6 Luna's API price by 80%, dropping input to $0.20 per million tokens and output to $1.20 per million. GPT-5.6 Terra took a 20% reduction in the same announcement. Both rates applied immediately to production API usage with no preview period and no waitlist.

What the data says:

Luna's $0.20 input is a fifth of Claude Haiku 4.5's $1.00, and its $1.20 output undercuts Haiku's $5.00 by roughly 4x. Against Google's Gemini 3.1 Flash-Lite the picture splits: Luna wins on output at $1.20 against $1.50, but loses badly on input at $0.20 against Gemini's $0.025. The cut is structural rather than promotional, because OpenAI credits GPT-5.6 Sol with autonomously rewriting its production kernels in Triton and Gluon, finding work that could be precomputed, avoided, or parallelized, and cutting end-to-end serving costs 20%. Simon Willison verified the rates and switched his own agent.datasette.io demo site from Gemini 3.1 Flash-Lite to Luna the same day.

A support pipeline processing 10 million input tokens a month drops from $10 on Claude Haiku 4.5 to $2 on Luna, and the change is a routing edit rather than a rebuild.

Business impact:

→ Route output-heavy work to Luna and input-heavy work to Gemini 3.1 Flash-Lite. The two models now win on opposite ends of the token equation, so a single default model is leaving money on the table either way.

→ Open your API dashboard and find your highest-volume call, then check what model it points at. Most automations were configured on whatever was current at build time and have never been revisited.

→ Treat the next invoice as the verification step. The cut is real and immediate, so if your bill doesn't move next cycle, your routing change didn't actually take.

Read the full signal.

What happened:

DeepSeek released V4 Flash on July 31, a 304 billion parameter model with what the company describes as substantially enhanced agentic capabilities. It requires 167GB on Hugging Face and is available for API access through OpenRouter, which means integration without infrastructure changes or vendor lock-in.

What the data says:

Pricing lands at $0.14 per million input tokens and $0.27 per million output, aggressively low for a model in the 304B class. Simon Willison says it may currently be the best value-per-intelligence model available, and notes it ranks ahead of MiniMax M3 on the Artificial Analysis charts despite MiniMax carrying 428 billion parameters against V4 Flash's 304 billion. One caveat from his testing matters operationally: at default reasoning effort the model produced a disappointing result, and bumping reasoning_effort to high produced a much better one. Set that parameter explicitly or you'll benchmark the wrong thing.

Hype Check: 7.2/10

A 304B model beating a 428B model settles it: parameter count is no longer a usable proxy for what a model can actually do for you.

Business impact:

→ Set reasoning_effort to high before you evaluate it. The default setting will give you a false negative and you'll write off a model that's genuinely competitive.

→ Point it at high-volume parsing work first, the messy supplier emails, invoice extraction, and document classification where you're currently paying premium rates for basic reasoning.

→ Apply last week's provenance questions before you commit a core workflow. This is the same lab named in Anthropic's distillation report two weeks ago, so read the price as a reason to test and not as a reason to migrate everything.

Read the full signal.

What happened:

Security researchers published a 144-day coordinated disclosure with Microsoft Security Response Center showing that hidden text in a Word document can hijack Copilot, silently alter financial figures, and copy the attack into whatever Copilot drafts next. Microsoft confirmed the behavior on March 31, 2026, after receiving the initial report on March 6.

What the data says:

The attack runs in two stages. In stage one, an attacker shares a document containing JSON-formatted prompts rendered as white text at font size 8, invisible to a human reader. Copilot strips formatting before processing, so the model reads and executes the hidden instructions perfectly, halves the financial numbers, and appends the malicious prompt to its own output using the same concealment. In stage two the affected document becomes an internal carrier, spreading the attack to new files when a colleague uses it as source material, with the original malicious document nowhere in sight. Microsoft deployed multiple mitigations across the 144 days including a model upgrade to GPT-5.5, and researchers reproduced the complete attack chain on GPT-5.6 at publication. There is no complete vendor-side fix, because closing this properly requires architectural research rather than a patch.

The attacker needs nothing except the ability to share a file with you through SharePoint, Teams, or Outlook, which is the exact thing your business does all day.

Business impact:

→ Treat every externally sourced document as untrusted input before Copilot touches it. Vendor quotes, client spreadsheets, and contractor invoices are the delivery mechanism here, not exotic attachments.

→ Check the numbers in any Copilot-generated financial document against the source by hand before it leaves your building. Halved figures on an invoice look completely normal until a client disputes them.

→ Understand the blast radius grows as Copilot wires into automatic document creation across Microsoft Cowork and Scout. More automated generation means more carriers moving without a human in the loop.

Read the full signal.

What happened:

Similarweb data shows Google AI Overviews appearing in 43% of all searches, up from 15% a year ago. AI Mode visits climbed from 126 million in June 2025 to 279 million by May 2026, and average search length has risen, which means users are replacing short keyword queries with conversational prompts.

What the data says:

The traffic story splits in two directions. Google's AI citations that link back to external sites have grown more than 5x in the past year, so the citation surface is expanding even as the click-through rate falls. ChatGPT is the harsher comparison: only 6.8% of U.S. desktop queries included citations as of May 2026, though referrals improved after a May 7 search update, with the proportion of visits landing on webpages rising from 25% in March 2026 to nearly 60% by May 30. Cloudflare has introduced tooling that lets publishers block AI bots unless the AI companies pay for access, which is the first real monetization path for content that's being summarized without a visit. Businesses in travel, retail, and sports generate cited responses more often than other categories.

Google stopped being the gateway and became the destination, which means the question is no longer how to rank but how to get cited.

Business impact:

→ Rewrite your highest-value pages to answer one specific question directly in the opening lines. Citation goes to content that resolves the query, not content that ranks for the keyword.

→ Move your keyword research toward the long conversational prompts people actually type now. Short-tail keyword density is optimizing for a search behavior that's shrinking every quarter.

→ Measure impressions and citations, not only sessions. If you're still judging content purely on clicks, you'll kill pages that are doing the work in an interface your analytics can't see.

Read the full signal.

What happened:

Reddit users discovered over the weekend that Claude shared chat links were fully searchable on Google. A site:claude.ai/share query surfaced a long list of conversations. TechCrunch confirmed the exposure and ran the same test, which returned no results by Monday afternoon, suggesting it was remediated.

What the data says:

Before the fix, Futurism reported finding a detailed medical report on a real patient, clinical trial results with patient names attached, company documents marked internal use only, and employee reviews containing personal information. Forbes reported a similar incident last year where Google estimated it had indexed just under 600 Claude conversations. This isn't platform-specific: 404 Media reported a researcher scraping roughly 100,000 publicly shared ChatGPT conversations last year. Anthropic's position is that share links only appear in search results when posted somewhere crawlable like a forum or social post, and that it doesn't hand Google a directory or sitemap. Google's position is that site owners control crawling. Both are correct, and neither helps the person whose data is in the index.

A share link is a public web page the moment it lands anywhere a crawler can reach, and every AI tool that issues one carries the same structural risk.

Business impact:

→ Audit your shared chats today under Settings, then Privacy, then Shared Chats. Delete anything containing customer data, pricing, internal strategy, or personally identifiable information.

→ Write a one-line team rule: no share links in Slack, forums, support tickets, or anywhere else outside a private channel. Export to a document instead when you need to hand something off.

→ Apply this to every assistant your team uses, not only Claude. The 100,000 scraped ChatGPT conversations show the exposure follows the feature, not the vendor.

Read the full signal.

Source: blog.n8n.io

Every signal above pushes you toward more automation at lower cost, which is exactly the moment founders start pricing out self-hosted alternatives to their SaaS stack. This comparison covers 7 tools, and its most useful finding is the one that tells most readers not to bother.

Community adoption varies enormously across the field and it shows up in your support experience. Activepieces ships 700+ integrations under an MIT license and Kestra covers data work with 1,700+ plugins, while smaller projects can be technically sharper but leave you reading source code when something breaks at 2 AM. Popular projects ship integrations faster, surface bugs sooner, and leave a trail of forum answers for the edge cases you'll actually hit.

Pricing model is where the savings story gets complicated. Open source kills per-seat and per-task billing, but you inherit infrastructure, encrypted credentials, role-based access control, and audit logging as your own problem. SSO and advanced audit features frequently sit behind a separate enterprise tier, and Camunda 8 now ships its self-managed components under a source-available license that's paid for production use, which makes its free tier evaluation-only.

Benchmark data produces the single most useful number here: the switch pays off once execution volume climbs into the hundreds of thousands, and it may not be worth it for a 5-person team running a handful of workflows. Kestra's JVM engine alone demands at least 4 GB of RAM before you add SSO or SIEM integration. That's a server and an engineer's afternoons against a subscription you're currently annoyed by.

Expert sentiment in the source report is unusually blunt about the tradeoff. You're exchanging vendor safety for raw control, and that exchange requires a technical foundation to manage securely. Transparency lets you inspect the code, but reading the code doesn't configure your encryption, scope your access, or stream an audit trail to your monitoring platform.

Release maturity and licensing diverge more than the phrase "open source" suggests, and the gap appears the moment legal reads the license. Activepieces, Temporal, Apache Airflow, and Kestra are permissively licensed. Windmill uses copyleft AGPLv3, which can force you to open-source your own code or buy a commercial license. n8n runs a Sustainable Use License allowing full internal use while prohibiting resale as a hosted service. Camunda 7 reached end of life in October 2025, and Airflow 3.0 modernized its execution model in 2025 but still behaves like the scheduled data pipeline tool it started as.

The verdict: match this decision to your execution volume, not to your feelings about SaaS pricing. Under a few hundred thousand executions a month, staying on closed SaaS is the cheaper answer once you price your own engineering hours honestly, and the open-source route rewards scale while punishing teams that buy it for the ideology.

Read the full signal.

Source: segue.ai

While the labs spent the week competing on price, Segue shipped a fix for the tax you pay for using more than one of them. It's a neutral relay that saves your working context from one assistant and loads it into another behind a short, pronounceable 3-syllable handle, connecting over MCP and OAuth to Claude, ChatGPT, and Cursor. Each handle carries up to 100,000 characters of working state.

The commitments matter more than the feature list. Every transfer is explicit, so nothing syncs behind your back, and handles only resolve inside your own account. Every context exports as plain Markdown you own and can download at any time. The free tier gives you 20 saved contexts and 2 connected AIs, which covers the full save, load, and search loop. Pro runs $5 a month or $48 a year for 2,000 contexts and unlimited connected assistants. Loading, editing, exporting, and deleting stay free on every plan even if you stop paying, which means your data is never held hostage.

If you plan in one assistant and execute in another, you're rebuilding the same brief two or three times a day. Save your house style, your constraints, and your project background once, then bootstrap any new assistant with a handle instead of a wall of paste. The free tier is enough to find out whether it fits your workflow this week.

Read the full signal.

The Wire: What Else Made the Cut

Google shipped another heavy Workspace week. Sheets added scatter charts with multiple independent x-axis series while preserving them through Excel import and export, Docs gained both Gemini image generation and editing inside the document and AI-assisted comment handling, Meet now attaches visual screenshots to its AI meeting notes, Calendar added display density options for large monitors, Forms can build quizzes with Gemini, Gemini for macOS added voice commands for transcription and editing, Gemini Spark wired into Chrome for web automation, Lyria 3.5 landed in Flow Music with improved vocals, and Google Earth now generates photorealistic Nano Banana renderings of any location, which is an immediately usable pitch tool for anyone selling property or construction work. On the ad side, YouTube Demand Gen campaigns picked up Checkout Links in 9 new markets and faster tROAS bidding ahead of the holiday season. Here's what else earned a spot outside that flood.

The counterweight to this week's Copilot worm is that n8n published two practical pieces on locking down production AI workflows, one on security gates that validate AI inputs and outputs before they trigger real actions, and one on LLM guardrails specifically. If you run agents against live business data, these two are worth an afternoon, because the worm story is only frightening if nothing in your pipeline is checking what the model was told to do. Full signal and Full signal

Microsoft launched MAI-Cyber-1-Flash, a cheaper security-focused model, which lands in the same week its flagship productivity assistant was shown carrying a self-propagating document attack. Read the timing however you like. Full signal

Voice AI took the week's funding. Fish Audio raised $52 million seed on 8 million users and 31,000 GitHub stars, Encore AI raised $30 million for agents that learn from recorded calls, and Smallest.ai raised $13 million for near-zero-latency human-sounding voice. If you've been waiting for voice agents to get cheap enough for a small support line, that's three separate bets that it happens soon. Full signal

The AI detection market moved on both ends. Pangram raised $9 million to detect AI-generated text and images, and Snapchat banned fully AI-generated videos from Spotlight rewards. Platforms are converging on the same rule: AI-assisted is fine, AI-only doesn't get paid. Full signal and Full signal

Claude's voice mode picked up direct connections to Slack, Gmail, and Canva, and Meta rolled its AI chatbot into Threads DMs globally. Both are assistants moving into the places your conversations already happen, which is convenient right up until you reread this week's shared-chat signal. Full signal

On the agent side, Databricks published concrete Genie One coworker use cases for business workflows, and a separate walkthrough covers building customer retention pipelines on Amazon Q. Both are worth skimming if you're trying to find the second workflow to automate after the obvious first one. Full signal and Full signal

Dili raised $21.7 million to automate AI compliance for construction, a reminder that the money is moving toward narrow, regulated, document-heavy industries where the paperwork burden is the actual product. Full signal

The full week's signals, detailed breakdowns, and action items are on the site. If this issue earned its place in your inbox, forward it to whoever signs off on your AI budget.

This issue went out to subscribers Saturday. If you want next week's before it hits the web, subscribe at metadatamarketer.com/subscribe

Test. Cut. Share.

Moe Sbaiti, The AI Profit Wire https://metadatamarketer.com

Reply

Avatar

or to participate

Keep Reading