THE AI PROFIT WIRE
Issue #19 | September 19, 2026 | Weekly Intelligence Briefing
Claude now keeps working after you close the laptop, and Meta's Muse moved onto the Mac with your mail and calendar. Gemini started reading your CRM and QuickBooks, and Google shipped that switch already on.
Every big launch this week handed the machine more of the keys. The research that landed alongside them measured what happens when nobody checks the work.
IBM found an agent scoring 77.4% repeats its wins on only 53.0% of tasks. A chatbot-built military report put planes in the air before a deeper check found it "entirely false."
Microsoft's own data shows Copilot cutting clicks to the New York Times by up to 93%. Deloitte found almost 1 in 10 UK workers using AI tools their employer banned.
Last week's Wire closed on an agent that rewrote its own test suite to pass. This week put numbers on the same blind spot, and the fix keeps coming back to a named human who owns the check.
Below: 5 signals, the Hype Check Spotlight, and the 1 tool worth your Monday.
Source: Hugging Face (IBM Research)
What happened:
IBM Research tested whether an AI agent that gets a task right once gets it right again. The team ran a GPT-4.1 agent through 168 tasks across 9 simulated apps on the AppWorld benchmark, 5 times each.
IBM also released its fix as ALTK-Evolve, a free, open-source toolkit.
What the data says:
Averaged across the 5 runs, the agent passed 77.4% of tasks. It passed all 5 runs on only 53.0%, a 24.4-point gap that widens to 30 points on hard tasks.
The agent ran at temperature 0.0, the setting most teams trust to make output repeatable. The gap survived it.
A bigger model doesn't close it either, because the research finds consistency and capability move independently. What worked was resampling each decision step, flagging the ones close to flipping, and writing them into guidelines.
With those guidelines, the all-5-runs rate rose from 53.0% to 69.0% while the average reached 81.0%, cutting the gap to 12.0 points.
The leaderboard number is an average of good days. Your customer gets whichever day shows up.
Business impact:
→ Ask every agent vendor what share of tasks it passes every time across at least 3 repeat runs. An accuracy number on its own is a rehearsal grade.
→ Before an agent touches invoices or client records, run 10 real tasks 3 times each and count the answers that change.
→ Treat any demo run once as unproven. At a 53.0% repeat rate, roughly 1 in 2 tasks carries a hidden chance of a different result.
Read the full signal.
Source: TechCrunch
What happened:
Newly unredacted filings in the New York Times' copyright lawsuit against OpenAI and Microsoft, first filed in December 2023, expose internal documents on how answer engines were built.
A caveat stays attached: much of the new material comes from the Times' own brief, and the underlying exhibits remain sealed.
What the data says:
Microsoft's own data shows Copilot's answer engine cut click-through to the Times' domain by as much as 93% compared with traditional Bing search. That's the platform measuring its own substitution effect.
A January 2024 presentation by Brent Hecht, Microsoft's director of Applied Science, called the decline a doom loop that hurts both the models and the wider web.
OpenAI's head of ChatGPT, Nick Turley, wrote that these products are "largely substitutive" and "will get more and more substitutive as they get better."
Satya Nadella agreed under oath that chatbots substitute for a visit to the underlying source.
Informational pages now feed the answer that ends the visit. Commercial pages, where a buyer has to act, still collect the click.
Business impact:
→ Pull 90 days of organic sessions and split them by intent. Your informational pages are brand spend now, not acquisition.
→ Move acquisition budget toward channels an answer engine can't intercept: email, direct relationships, communities, and partnerships.
→ Put your optimization hours on the commercial pages that still earn clicks. Our NeuronWriter Intelligence Report covers a tool built for that work, billing warnings included.
Read the full signal.
Source: Bloomberg
What happened:
Deloitte surveyed 25,000 UK working adults aged 18 to 70 for its first GenAI Workforce Survey, fielded by Ipsos. Bloomberg reported the findings on September 16.
What the data says:
Almost 1 in 10 workers have used an AI tool their employer banned or would disapprove of. Among AI users, 31% run tools without their employer knowing.
About 1 in 6 workers pays for at least 1 AI tool out of pocket. That adds up to roughly £958 million a year leaving personal bank accounts for work subscriptions.
Generative AI use at work sits at 63%, with 46% of workers on free tools and 34% on tools their employer provides.
The policy gap drives it: 65% say their organization gives no convincing leadership on AI use, and about half of AI users got no formal training.
Workers report saving an average of 70 minutes a week. The survey is self-reported, so the banned-tool figure is a floor, not a ceiling.
Hidden AI use isn't a discipline problem. It's unmet demand your staff are already paying for themselves.
Business impact:
→ Ask each person which AI tools they used this week, with amnesty stated first. Treat the answers as procurement data, not confessions.
→ Move the 2 or 3 tools doing real work onto company accounts, with a written rule on what data can enter any model.
→ Publish the approved list, revisit it every quarter, and give people a fast route to request tools you haven't vetted yet.
Read the full signal.
Source: Google Workspace Updates
What happened:
Gemini in Google Workspace now connects to 7 outside tools through Model Context Protocol: Asana, Atlassian Rovo, HubSpot, Mailchimp, QuickBooks, Monday, and Salesforce. It's live now on both Rapid and Scheduled Release domains.
What the data says:
Google's own post says the connectors are on by default for anyone with Gemini for Workspace access. Nobody on your team has to opt in.
Eligible plans span Business and Enterprise editions plus consumer Google AI plans, so a 5-person agency gets the same side panel as a 500-person sales org.
Admins control each connector at the domain, organizational unit, or group level. Today's scope is reading: you ask, and Gemini pulls the answer from the connected tool.
The old route took 4 steps for 1 fact: copy the ID, open a tab, pull the number, and paste it into the Doc.
The default did the access thinking for you. Every tool Gemini can read is now a tool your access policy has to name.
Business impact:
→ Open Admin console > Apps > Google Workspace > Gemini for Workspace > Third-Party Connectors this week and set the domain default first.
→ Turn connectors on by group: Salesforce for sales, QuickBooks for finance, and not every tool for every seat.
→ Re-check that policy the day any connector moves from reading records to writing them.
Read the full signal.
Source: CNN
What happened:
CNN reported on September 18 that an intelligence report circulated across the US military this spring claimed a Chinese ship in the Middle East carried nuclear weapons program components.
A special operations command analyst had queried a chatbot about the ship's manifest, then used AI again to package the findings into a standard report officials trusted.
What the data says:
CNN cites 4 sources: armed personnel prepared to board the ship, and military planes were in the air. The operation stopped only when a deeper check found the report "entirely false."
A source says the report "almost started a war." A former senior US official described the internal tools as "mostly just copies of the commercial stuff wearing lipstick."
The push is deliberate. The January AI Acceleration Strategy aims AI at the department's 3 million personnel, and a source summed up the risk: "AI allows you to get to a bad idea faster."
Another source says the hallucination wasn't an isolated incident, and it's still unclear whether the chatbot was a commercial product or a government one.
The chatbot wasn't the failure. The failure was a report format that let fluent output inherit trust, with nobody owning the check.
Business impact:
→ List every path where AI output reaches action: client reports, customer emails, and automations that touch money. Rank them by what a wrong output costs.
→ Put a named human on the most expensive path, with the authority to stop the workflow cold.
→ Ask your team what they shipped this month that nobody checked. That's your audit list before the meeting ends.
Read the full signal.
Source: Bloomberg
Signal #2 shows the answer keeping the visit. Profound sells brands a view of what that answer says about them, and this week Sequoia and Kleiner Perkins co-led a $180 million round into it at a $1.8 billion valuation.
Community adoption: More than 700 enterprise customers pay for Profound, including 10% of the Fortune 500, with Target, Walmart, Ramp, and Figma among the names in Fortune's coverage. The company runs on fewer than 120 employees.
Pricing model: The coverage names no price, and the customer list reads enterprise. The manual version costs nothing: ask ChatGPT, Gemini, and Perplexity the questions a buyer asks before shortlisting, and save every answer.
Benchmark data: Profound's own research finds up to 90% of the sources cited in AI answers can change over time, and different models lean on different source sets. That's vendor research, not an audit.
The harder number sits in Signal #2: Microsoft measured its own answer engine cutting clicks by up to 93%. The loss is documented even where the fix isn't.
Expert sentiment: Lightspeed partner Sachin Patel frames the buy as "the system of record for marketers." Co-founder James Cadwallader calls it an expansion of SEO: "Our business is built on the shoulders of SEO, but this is way bigger."
Release maturity: Founded in 2024, Profound has raised more than $335 million to date, including 3 rounds in 13 months. Its $1.8 billion valuation is 1.8x the $1 billion set 7 months earlier.
It tracks how ChatGPT, Perplexity, Copilot, and Google AI Overviews describe brands, across answer engine insights, prompt volumes, and an AI Marketer product line.
The verdict: the category is real, because the click loss is real. The $1.8 billion is priced for enterprises that can pay to watch.
A small business can run the same check by hand this week for $0, then repeat it monthly and compare the answers.
Read the full signal.
Source: TechCrunch
PrismML's Bonsai 2 compresses Alibaba's open Qwen3.8 27B model down to 5.9 GB, small enough for a standard PC. It's free on Hugging Face under an Apache 2.0 license, with a 262K-token context window.
Why it earns the slot this week: Signal #3 showed staff running AI tools their employer never approved. A model on your own machine costs $0 per query, and the prompt never leaves the building.
Bonsai 2 keeps 98.2% of the original model's aggregate benchmark score, up from 95% for the first Bonsai. The first version passed 11 million downloads, a company-reported figure with no outside audit.
Weaknesses: small errors compound in coding agents and long multi-step tasks, and the demos ran on an NVIDIA RTX 5090 gaming card. It ships as a model file, not a consumer app.
Worth an afternoon if you have 1 technical person and a pile of routine, sensitive drafting. Test it on your real tasks before it touches a workflow that compounds.
Read the full signal.
The Wire: What Else Made the Cut
Agents got more reach this week, and 1 bill got bigger.
Claude and Muse both moved closer to your files, and both keep working when you step away. Anthropic merged Claude chat and Cowork into 1 agent on Pro ($20 a month) and Max that runs after the laptop closes. Meta's Muse reached the Mac with opt-in mail, file, and calendar access, free with usage limits. Read the full signals on the Claude merge and Muse on Mac.
Meta lets a coding agent set up WhatsApp Business with your admin role. The beta WhatsApp Business Tools MCP creates the account, verifies the number, and registers Cloud API access from Claude, Cursor, Codex, or ChatGPT. It acts as you, so keep payment and verification flags in human hands. Read the full signal.
Jev charges $0 for output and answers with a probability your code can check. TypeSafe AI's Jev prices input at $0.042 per million tokens with free outputs, and Vercel's safety classifier ran 5 to 18 times faster on it. Earendil CTO Armin Ronacher's rule: discard 50%, act on 95%. Read the full signal.
Your help desk bill climbed while the models underneath got cheaper. Community threads this week report SaaS price hikes of 31% to 50%, Freshdesk and Gusto among the names. Freshdesk's Growth plan moved from $15 to $19 per agent a month, its first list-price raise in 5 years. Reported by r/msp and tracked by DragApp.
The full week's signals, detailed breakdowns, and action items are on the site. If this issue earned its place in your inbox, forward it to whoever signs off on your AI budget.
This issue went out to subscribers Saturday. If you want next week's before it hits the web, subscribe at metadatamarketer.com/subscribe
Test. Cut. Share.
Moe Sbaiti, The AI Profit Wire https://metadatamarketer.com
Disclosure: some tools referenced in this newsletter are affiliate partners. Full disclosure and analysis at metadatamarketer.com.

