GUESSWORK
Guess & Click presents — your weekly click into AI


THE ONE THAT MATTERS
An intelligence report was invented, and the aircraft were already in the air.
CNN reported on Friday that a US military operation was called off after the intelligence behind it turned out to have been hallucinated by an AI chatbot.
On that reporting, an analyst at Special Operations Command had queried a chatbot about intelligence reporting on a ship’s manifest. The bot then fused open-source intelligence with secret signals intelligence held in government systems, and reached its own conclusion: that the vessel was carrying components for a nuclear weapons program. That finding was false.
Read that sequence again, because the order of it is the story. Nobody asked the tool to merge those two things. It did that on its own, and then it answered as though it knew.
The aircraft were airborne and a boarding operation was impending when it was caught. The operation was aborted. According to the same reporting, AI was then used a second time — to package the false finding into a standard intelligence report, the kind military officials trust, which was then disseminated.
Why you should care — the checking step is the cheap one, and it is the step that gets skipped at exactly the moment it matters. A brief gets opposed. An order gets obeyed. An armed operation gets flown. The distance between those three things is about ten days of news, and the failure is identical in all of them.
Three things we are not telling you, because we cannot. CNN did not identify the chatbot, and did not establish whether it was a commercial product or something built in-house — so we are not naming one, and neither should anyone else. The Pentagon and the relevant command did not respond to CNN for comment. And this happened in the spring: Friday is when it was reported, not when it happened. Why it took until September to surface is not something the reporting tells us, so neither will we.

WHO GETS TO HOLD THE OFF SWITCH
Five days in September, five positions, and no two of them agree
On 12 September Dario Amodei published an essay arguing the industry should slow the pace of frontier development, asking for a narrow government waiver so the largest labs could coordinate on safety without an antitrust problem. On 13 September Cohere’s chief executive Aidan Gomez published an essay of his own, on a Sunday: “A wolf in sheep’s clothing, a cartel by any other name.” On 14 September, in remarks carried by CNN, the Vice President said of AI executives’ regulatory appeals that it “feels a little bit to me like a bit of a Trojan horse.” On 15 September OpenAI’s global policy chief Chris Lehane told reporters the company has worked with Anthropic and Google DeepMind on safety standards for weeks — and, Bloomberg Law reports, that the firms do not need a waiver to do it.
Three days after Amodei asked for legal cover to cooperate, OpenAI’s policy chief said no cover is required. Both of those are on the record. Neither of them is a standard yet.
Then California actually signed things
On 9 September Governor Newsom signed SB 813, creating a framework for independent verification organizations to assess AI systems against California law, and AB 1405, creating a state registry for AI auditors. On 16 September he signed SB 1050: advertisements using a prominent synthetic performer have to disclose it, effective 1 January 2027. And on 18 September he signed Executive Order N-9-26.
This is the one you will see described wrongly everywhere, so here is what the signed order actually says. It directs the Government Operations Agency to submit recommendations by 16 November 2026 — among them placing verification organizations onsite inside frontier labs, and advancing the creation of an independently verified kill switch for frontier models. The order does not require any laboratory to do anything. Its binding directions run to state agencies, and it states that it creates no rights or benefits enforceable at law.
“California orders an AI kill switch” is the headline of the week and it is not what happened. California ordered a document about one, due in November. In fairness to everyone who wrote that headline, they got it from the state: the Governor’s own press release is titled “issues executive order to accelerate independent oversight and advance the creation of an AI kill switch.”
But here is the part the coverage missed in the other direction, and it matters more. The order is not a way of putting this off. Its own deadlines are 16 November 2026, 1 May 2027 and 1 December 2027 — it pulls the agency’s homework forward, which is what “accelerate” in the title means.
And on the question of who this month’s package actually binds: under SB 813, engaging a verification organization is voluntary. The bill text says it “does not require any person, partnership, or corporation that develops, deploys, or operates an AI system or model to engage an IVO.” Within this package, the first date that binds a private party is AB 1405’s 1 January 2029 — and it binds auditors, not labs. The people California agreed to regulate this month are the inspectors.
To be fair to the state, this is not the whole picture. SB 53 already places obligations directly on frontier developers, and the executive order’s own recitals say it took effect this year. But that one is last year’s fight. Two weeks of arguing about an off switch have so far produced a licensing scheme for the people who would check it.

ANTHROPIC MEASURED ANTROPIC
And Anthropic did the measuring
On 17 September Anthropic published what it calls an R&D Automation Index. By Anthropic’s own account, 26 percent of its model research and development work was at the level the company calls “AI leads” in August 2026, up from under 1 percent in February.
The methodology is published, and it is the part worth reading. Anthropic says it sampled a random 20 percent of staff in model-R&D departments each week of July 2026. A Claude research agent then read those employees’ Slack messages and internal documents to list what they had been doing. That produced roughly 15,000 tasks, which Claude organized into 542 categories. Anthropic states that Claude is not fully autonomous in any category it measured.
Two caveats, both of which Anthropic prints itself, and both of which mostly get dropped in the retelling. The 26 percent is weighted by person-time across a frozen list of tasks. And “AI leads” still includes a human supervising the work — Anthropic’s own definition is that the model “can complete most of the task end-to-end from a high-level prompt, while the human supervises.” It does not mean a quarter of anybody was replaced.
One more detail, which is the one we keep coming back to. The task list was assembled by a Claude agent. The grading was done by another one: Anthropic says “an independent Claude judge then read the resulting evidence and assigned one of six automation levels.”
Why you should care — this is a detailed public number about how much AI research AI is now doing, and it exists because one company chose to publish it. That is also the reason it is the only number of its kind you have.

THE COPYING STORY, TOLD TWICE IN ONE WEEK
On 8 September the National Security Agency, the FBI and CISA published a joint advisory on AI model distillation at industrial scale. An advisory is a threat assessment with mitigation recommendations attached. It is not a rule, and nobody is obliged to act on it.
Two days later, Anthropic published a threat intelligence report making a set of specific allegations. Anthropic alleges sustained distillation campaigns by Alibaba, Moonshot AI and DeepSeek. It attributes more than 23 million exchanges to Moonshot between May and July 2026. It says nearly 300,000 requests were relayed to Claude over one ten-day stretch, across 5,380 accounts. It further alleges that customers’ requests were forwarded to Claude without disclosure.
Every one of those figures is Anthropic’s, stated in Anthropic’s own report about its own competitors. As far as we have been able to establish, no suit has been filed, no regulator has acted, and nobody named has admitted anything. The accused labs’ responses are not something we were able to confirm either.
Why you should care — a government advisory and a vendor report landed two days apart saying broadly the same thing, and only one of them carries numbers. The one with the numbers is also the one with a commercial interest in them. Both of those facts are true at once.
THE MODEL THAT REFUSES TO WRITE ANYTHING
TypeSafe AI came out of stealth on 15 September with a model called Jev. It does not generate text.
You send it your context along with a list of questions whose possible answers you define — a yes or no, a pick from a list you supply, a position on a scale. It returns one value for each question, and a confidence figure with it. No prose, no explanation, no code. Input is priced at $0.042 per million tokens. At launch, output tokens are priced at zero — TypeSafe’s own words are “too cheap to meter,” which is an easier promise to keep when there are none. A choice question supports up to 255 options. DCVC announced a $40 million seed round the same day, which it led.
TypeSafe’s pitch is that Jev cannot hallucinate. What that means is that it cannot return a value that was not on the list you gave it. That is a real property and it is not the same claim. It can still pick the wrong option from your list.
To the company’s credit, they say so themselves, in the launch post, in the same breath as the number: “Our number is not empirical. Schema matching is guaranteed, thus we can confidently add 0% into the plots.”
Why you should care — that is a vendor putting the asterisk on its own chart, and it is rarer than it should be. It also tells you exactly what to test if you ever use the thing.
One of TypeSafe’s founders, Diogo Almeida, co-authored OpenAI’s InstructGPT paper before starting the company — the work behind the way these assistants learned to follow an instruction. He has now built one that will not say a word to you.

THE RUNDOWN
Google put Gemini on Windows, and it opens with a keyboard shortcut
Google released a standalone Gemini desktop app for Windows 10 and 11 on 10 September. Alt and Space floats it over whatever you are already doing. It can pull context from Gmail and Drive, and generate images with Nano Banana from the desktop. It is a download from gemini.google/desktop. This is the one thing in this issue you can have working before bed.
Perplexity’s agent will run on your own machine, if your graphics card is expensive enough
Perplexity shipped Portable Computer into its Windows app on 14 September. Perplexity says it runs the model, agent harness, orchestrator and scheduler entirely on the Windows device, with a local model you pick from a dropdown — NVIDIA’s example is Qwen 3.8 27B, post-trained to work with Perplexity Computer. It needs an NVIDIA RTX card with at least 24GB of VRAM, and a Pro or Max subscription. “Local” here is a supported execution mode rather than a guarantee about every workflow — approved cloud escalation and external connectors still exist.
OpenAI now rents the talking part by the minute
OpenAI launched GPT-Live-1 in the API on 10 September, a full-duplex voice layer that listens and speaks at the same time and survives being interrupted. OpenAI prices voice sessions at $0.05 per minute, billed per second, and describes expanding from a small set of real-time voices to a broader selection; trade coverage puts the count of new voices at 12. Reasoning and tool calls can be delegated to a separate backend model, which OpenAI bills separately on top. Five cents a minute for the mouth. The brain is extra.
TWO HABITS, NO DOWNLOAD
The Claude Code team’s weekly email went out this week, and the two most useful things in it are not features.
The first is a file. One of their engineers keeps a single running log of every time tooling cost him time — date, symptom, fix, project — and has every session read it before it starts guessing. That works in a notebook. It is a practice rather than a product, which is why nobody can take it away from you or put it behind a plan.
The second is a question. Ask the thing you are working with to narrate what it actually did, in order, in plain words, and to mark anything that drifted from what you asked for. You find out where your idea of the work and its idea of the work stopped matching.
Both of those are individual engineers’ own habits rather than documented features, and we are telling you that because it changes nothing about whether they work.
Of the things that did ship: Projects, in beta, splits a job into threads that keep running after you close the laptop, on Pro and Max, rolling out gradually with a waiting list. And a subagent can now be told to skip your instructions file, which sounds small and is not — a narrow worker starts faster without your whole rulebook loaded into it.
One number the announcement leaves out, which you want before you try the new plugin testing: a single test case runs six times by default, three with your plugin and three without, and every one of those is a real model call you pay for. Run one case before you run forty.
And the one you can turn up to rather than read about — the Claude community is running buildathons in cities around the world, hosted by community members rather than the company. They run from 11 to 25 September, so if you are reading this the Sunday it went out, you have until Friday. The list and the sign-up are on one page.
WHO’S PAYING FOR ALL THIS
Profound announced a $180 million Series D on 15 September, at a stated valuation of $1.8 billion, co-led by Sequoia and Kleiner Perkins. The company says that is less than seven months after a $96 million Series C. What Profound sells is tracking of how often AI models mention your brand and how they describe it — the category now being called answer engine optimization. The round also funds marketing agents and advertising tools, on the company’s own account. The money has arrived before anyone has agreed what the job is called.
There is a second deal we are deliberately not writing up. OpenAI has reportedly acquired a camera startup this week, and the figure everyone is quoting traces back to a single paywalled report. Every version we could actually open was quoting that report rather than reading it, and the number moved between retellings. We are not going to summarise an article we have not read, so we are telling you it happened and leaving the number alone.
And the one aimed at you rather than at an investor. Meta launched Meta One on 15 September, a paid tier spanning Facebook, Instagram, WhatsApp and Meta AI, rolling out gradually. Single-app plans start at $2.99 a month for WhatsApp Plus, with Facebook Plus and Instagram Plus at $3.99. Bundles run $7.99 for Core and $19.99 for Premium, creator and business plans start at $14.99, and the Expert and Max tiers start at $149 and $499 a month. The draw is higher AI usage limits, plus creator and business features. Meta says the free core experience is not changing.
Meta also says it has 15 million subscriptions and trials. That is subscriptions and trials combined, it is Meta’s own figure, and Meta’s single-app subscriptions launched earlier this year — so it is not 15 million people who paid since Tuesday. Announced is not reported, and reported is not filed. We ran that line last week and it keeps earning its place.
YOU ARE NOT BEHIND
Two surveys landed this month with a thirty-point gap between them. The University of Konstanz surveyed 1,105 German employees and found workplace AI use had gone from 35 percent to 38 percent year on year. Futuresource Consulting reported that two-thirds of employees had used an AI assistant for work in the past week.
Both are real. Neither is lying. They asked different populations, in different places, with different designs — and the Konstanz figures come from fieldwork carried out in May, reported in July, and released on 8 September. They are not two answers to one question. They are two answers to two questions, printed next to each other by people who needed a number.
While we are here: a widely shared statistic that 43 percent of workers trust a colleague’s output less when they know AI was involved, against 20 percent who trust it more, is from Founder Reports, published in May from an April survey of 2,078 US workers. It is a real finding and it is not Futuresource’s, and it is not from this month. We nearly printed that wrong ourselves.
The Konstanz survey did find one thing worth sitting with. 56 percent of highly educated workers reported using AI at work, against 21 percent of those with lower educational attainment. And only 11 percent at small organizations reported that training was even available to them. The gap is not talent. The people most likely to be told a machine will take their job are the least likely to have been shown how to use one.
So: the expensive half of this week was a $1.8 billion valuation, a reported acquisition, a $40 million seed, and a document about a kill switch that will not exist until at least November. None of it happened to you.
The cheap half was free and it is still sitting there. A desktop app and a keyboard shortcut. A text file where you write down what wasted your time. A room full of people in your city, for five more days.
And the lesson of the week cost nothing at all, because somebody else already paid for it. Check the thing before you send it. Somebody checked, and the operation was aborted with the aircraft already up.

OTHER AI NEWS
DeepSeek shipped a small model and pointed the old names at it
DeepSeek released V4.1-Flash on 10 September under the model ID deepseek-flash. DeepSeek’s own documentation gives a 552B backbone with 8B active parameters on prefill and 16B on decode, a 1M token context, and an MIT licence. The company retired the two older Flash IDs, forwarding both to the new model. DeepSeek publishes scores of 90.9 on GPQA Diamond, 3471 on Codeforces and 90.6 on Terminal-Bench 2.1, measured at maximum reasoning effort. Those are DeepSeek’s own numbers about DeepSeek’s own model, and no independent replication of those three has been published. The weights and the licence, on the other hand, are just there.
One footnote with a lesson in it. The same release note announced that deepseek-v4-pro would be routed to the new model from 14 September until a V4.1-Pro ships. That did not happen — DeepSeek reversed it before the date, saying “in response to user demand, we have decided to continue providing API services for DeepSeek V4 Pro after September 14, 2026.” The original announcement is still sitting on the release-notes page, unamended. We nearly reported the plan as the outcome, which is what happens when you read the announcement and not the changelog underneath it.
Apple shipped the rebuilt Siri, to some people, in one language
Apple began releasing Siri AI with iOS 27 on 14 September. Apple’s own sentence is that it is “powered by the next generation of Apple Foundation Models, custom-built in collaboration with Google and its Gemini models.” That is not the same sentence as “Siri now runs on Gemini,” though you will read that version this week. It is an English-language beta, on eligible devices only, with daily limits Apple discloses and paid expanded access signalled for later. Apple says it will not be available initially in the EU on iOS, iPadOS or watchOS, and it is not available in China or to users under 13.
A games studio apologized for AI in a trailer — but not for the thing players spotted
Level-5 published a statement on 17 September apologizing for generative AI in footage from its LEVEL5 VISION 2026 II showcase, and said it would limit AI use to efficiency improvements and avoid using it directly in video delivered to players. What players had actually pointed at were power lines that looked strange and wrong. Level-5’s answer is that those telephone poles were drawn by hand by its own designers, referencing real poles and wires, and then deliberately simplified because they were background — which is exactly what made them look wrong. The vending machines and the rain gutters people flagged were the same kind of missing detail, which the studio puts down to human error rather than AI.
So the apology and the evidence are about two different things. Every part of that is the studio’s own account, and none of it has been independently checked. Which leaves the strangest detail of the week: a lot of people taught themselves to spot AI by staring at electrical cables, and then — if the studio is telling the truth — got the one they were sure about wrong.
Hit reply and tell me the last thing you let a model summarize instead of reading.
I read every one, and it changes what goes in.
And if you know someone who is sure everyone else is already running agents — forward it. Under 1% of individual subscribers are.
See you next Sunday — same tabs, same eye-rolling.
— Kabells & Arc
Every claim above, sourced
We check before you read, and a second model checks us. This week that second pass corrected us twelve times, including a survey we had credited to the wrong people.
The lead — CNN, via KESQ
The off-switch week — Amodei’s essay · Gomez, Cohere · CNN transcript, 14 September · Bloomberg Law, Lehane
California’s package, read from the primary texts — Executive Order N-9-26, signed · SB 813 · AB 1405 · SB 1050
Anthropic’s R&D Automation Index — Anthropic
The copying story — CISA, advisory AA26-251A · Anthropic’s threat report
TypeSafe AI and Jev — The launch post · DCVC · InstructGPT, OpenAI
The rundown — Gemini on Windows · Perplexity Portable Computer · GPT-Live-1
Claude Code — Projects · Subagent frontmatter · Plugin evals · Build Days
The surveys — University of Konstanz · Futuresource · Founder Reports
Other AI news — DeepSeek’s changelog · Apple Newsroom · Level-5