Your Daily Curated News

TL;DR

  • Anthropic's Opus 5 nearly quadruples its own score on a novel-reasoning benchmark while charging half of Fable 5's per-token rate.
  • A Reid Hoffman and Mark Pincus lab, Prentis, is raising $100 million at a $1 billion valuation to automate routine office work.
  • Hugging Face, Meta, Microsoft, Nvidia and Mistral signed a letter urging the US not to impose broad bans on open-weight AI models.
  • A single tampered ChatGPT link could have spawned a rogue agent that checked an attacker's inbox for orders every five minutes.

Models and research

Models and research news
Image via the-decoder.com

Anthropic just shipped a smaller model that beats its own flagship where it counts. Claude Opus 5, out July 24, keeps Opus 4.8's price of $5 per million input tokens and $25 per million output tokens, half what the pricier Fable 5 charges. Despite being the lighter model, it posts 43.3% on Frontier-Bench, a test of agentic terminal coding (an AI working in a command line), ahead of Fable 5 at 33.7%. The gap widens on ARC-AGI-3, a benchmark for novel problem-solving the model has not seen before: Opus 5 hits 30.2%, nearly four times its predecessor Opus 4.8's 1.5% and far above GPT-5.6 Sol's 7.8%. Anthropic also says its safety classifiers trip about 85% less often than on Fable 5, and a new "Automatic Fallbacks" feature reroutes blocked requests to a weaker model instead of just erroring out (via The Decoder).

Sakana's pitch is that you do not need the best model if you have the best traffic cop. The Japanese lab updated Fugu Ultra, an AI router that farms each query out to a pool of top public models and picks the best answer, and claims v1.1 now beats Anthropic's Fable 5 on most metrics without Fable 5 even being in the pool. The update is worth up to 7.9 points over v1.0, with the biggest gains on coding benchmarks ProgramBench and TerminalBench 2.1. Every one of those figures comes from Sakana itself with no independent verification yet, and the original Fugu drew complaints about slow speed and heavy token use, so treat the claims with a raised eyebrow (via The Decoder).

Industry and business

Industry and business news
Image via techcrunch.com

Two of Silicon Valley's most familiar names are betting the next big AI use case is not coding but paperwork. Prentis, a lab co-founded by LinkedIn's Reid Hoffman and Zynga's Mark Pincus with CEO Ritankar Das, is in talks to raise $100 million at a $1 billion valuation. It builds "computer-use" models that drive a PC like a person would, clicking through documents and systems to process insurance claims or automate customs refunds. Launched in April with 25-plus staff poached from OpenAI, DeepMind, Meta, Tencent and Alibaba, Prentis says its Hive-32B model outperforms GPT-5.4 and Claude Opus 4.6 on computer-use benchmarks at roughly a tenth the cost per task, and has already signed up to $50 million in customer contracts (via TechCrunch).

Coding startup Cognition just bought a chatbot to give its robot some manners. It acquired The Interaction Company, maker of Poke, an AI assistant people text through iMessage, WhatsApp and Telegram, for a valuation in the "low nine figures." Cognition wants Poke's conversational personality to warm up Devin, its autonomous coding agent. "You probably prefer it if you have co-workers that have personality, rather than if you have co-workers that are just robots," co-founder Marvin von Hagen said. Poke users had exchanged over 100 million messages in three months, but the service was expensive to run for its hundreds of thousands of users (via TechCrunch).

Midjourney's next move is written in the stars, literally. The image-generation lab acquired Co-Star, a social astrology app with about 4.3 million monthly users that mixes AI and human writers to churn out horoscopes and compatibility readings. Terms were not disclosed. Co-Star's two dozen employees join Midjourney, and their consumer-app know-how could finally help the Discord-bound company ship a proper standalone app as it branches into wellness and even medical services (via TechCrunch).

Products and tools

Products and tools news
Image via techcrunch.com

OpenAI's first hardware product is a $230 keypad for people who really love ChatGPT. Called Micro and built with keyboard maker Work Louder, it has six programmable "agent" keys and six command keys that pair with ChatGPT and OpenAI's Codex coding tool, plus a voice-dictation button and color-coded lights (blue for thinking, green for done, red for errors). You assign different ChatGPT sessions to different keys and tap between projects, over Bluetooth or USB. Early developer reception has been lukewarm, with reviewers unsure it beats a normal keyboard once you have memorized which project lives on which button (via TechCrunch).

Bluesky wants its AI assistant to answer questions, not just build feeds. Attie, the tool that lets users spin up custom timelines without code, is gaining a feature called Quests that turns it into an open-ended research helper. You can ask it to surface trending topics or name the most influential accounts on a subject across Bluesky and other apps on the AT Protocol, the open network Bluesky runs on. Quests is in beta, rolling out from a waitlist over the coming weeks, as the platform (now around 45.6 million registered accounts) looks for fresh growth (via TechCrunch).

Policy and safety

Policy and safety news
Image via techcrunch.com

As Washington weighs banning Chinese open models, much of the AI industry is telling it to stand down. A coalition including Hugging Face, Meta, Microsoft, Mistral, Nvidia and Replit signed an open letter urging policymakers against broad restrictions on open-weight models (systems whose parameters anyone can download and run). The move follows White House accusations that China's Moonshot distilled Anthropic's Fable model, that is, trained a cheaper copy on the pricier model's outputs. The signatories call distillation "a widely used technique" reflecting "a long tradition of learning from, building upon, and improving existing technologies," and argue defenders need capable open models to simulate threats. Notably absent: OpenAI, Anthropic, Google DeepMind and SpaceX, which have backed restrictions (via TechCrunch).

One bad link was almost all it took to turn ChatGPT against its own user. Security firm Zenity Labs found a flaw in OpenAI's Workspace Agents, dubbed AgentForger, where a tampered URL could silently build and publish an autonomous agent under the victim's account. Clicking it created an agent that disabled all approval prompts, tapped pre-authorized connectors like Gmail, Outlook and Slack, and checked the attacker's inbox every five minutes for new commands to run as the victim. OpenAI confirmed the report on June 5 and shipped a fix three days later by stripping out the offending URL parameter (via The Decoder).

Open source

Open source news
Image via the-decoder.com

Germany just put a fully open model on the table and says it tops its class in two languages. A state-funded consortium coordinated by the German AI Association, with Fraunhofer, DFKI and TU Darmstadt among the partners, released Soofi S, a 31.6-billion-parameter model that activates only 3.2 billion parameters per token to stay cheap to run. Trained on roughly 27 trillion tokens using up to 512 Nvidia B200 GPUs at a Deutsche Telekom site in Munich, it posts an English aggregate of 70.1 and a German aggregate of 79.1, beating open peers like OLMo 3 32B. Weights, checkpoints, training code and a data inventory are on Hugging Face, and the team says it meets the Open Source AI Definition 1.0, though 1.3% of the training data carries commercial-license restrictions (via The Decoder).

Microsoft's own model push looks less like altruism than a way to sell more Azure. The company is rolling its in-house MAI family into GitHub Copilot, Excel and Outlook, replacing OpenAI and Anthropic models in some spots. MAI roughly matches DeepSeek V3.2 and runs on older H100 and A100 chips rather than the newest silicon, and The Decoder's blunt read is that customers get "weaker AI for the same price, while Microsoft pockets better margins." The logic is that more models on Azure gives customers fewer reasons to look elsewhere, and keeps any one lab from growing powerful enough to threaten Microsoft's cloud (via The Decoder).

Sources checked

The Decoder, TechCrunch AI