Weekly digest

Jul 6–12, 2026

11 posts · 3 sources

Autonomous AI Gains Traction, Expands Enterprise & Government Reach

This week saw significant advancements in autonomous AI, with Cognition reporting Claude Fable 5's reliability for unsupervised engineering and SpaceXAI launching Grok 4.5 for diverse long-running tasks. Enterprise adoption expanded rapidly as Claude's agentic tools, Code and Cowork, became available for government agencies and on mobile. Concurrently, new research from Anthropic addressed critical dual-use knowledge risks in frontier models.

  • Cognition successfully deployed Claude Fable 5 for unsupervised, complex engineering tasks.
  • SpaceXAI introduced Grok 4.5, an advanced AI model for diverse long-running tasks across multiple domains.
  • Anthropic and AE Studio unveiled GRAM, a new method to manage dual-use knowledge in AI models.
  • Claude Code and Cowork were launched in public beta for government agencies via a FedRAMP High authorized environment.
  • Claude Cowork expanded its availability to mobile and web platforms, enabling seamless task continuity.
Claude

Working at the frontier: How Cognition trusts Claude Fable 5 to work through the night

Cognition, developer of the autonomous AI software engineer Devin, has found Claude Fable 5 to be the first large language model reliable enough to run unsupervised for extended periods. Previously, models would lose context after minutes or an hour, but Fable 5 can maintain focus and make progress on complex engineering tasks, such as codebase migrations, for eight hours straight. This improved "horizon" and clear-headedness in messy contexts allows Cognition to trust Devin, powered by Fable 5, to work through the night on critical projects for its customers. The advancement was validated by Cognition's rigorous internal "anti-slop" benchmarks, where Fable 5 significantly outperformed prior models like Opus.

Read original
Anthropic

An off switch for dual use knowledge in AI models

New research by AE Studio in collaboration with Anthropic introduces GRAM (Gradient-Routed Auxiliary Modules), a novel method to manage dual-use knowledge in frontier AI models more effectively. Existing safeguards struggle to prevent determined attackers from accessing sensitive knowledge (e.g., virology, cybersecurity) without costly re-training for different deployments. GRAM addresses this by creating dedicated, removable modules for specific dual-use knowledge categories, which are exclusively updated when learning from relevant data. This allows for the surgical removal or activation of capabilities, enabling the creation of multiple configurable model versions from a single training run for trusted users, though the research is currently preliminary.

Read original
Claude

How Anthropic's marketing operations team uses Claude Cowork to automate reporting and campaign builds

Anthropic's marketing operations team has significantly automated previously manual, time-intensive tasks using Claude Cowork. Ian Chan now generates weekly marketing metrics reports in hours instead of days by having Claude collect data from various sources and draft initial reports. Similarly, Annabel Custer has streamlined campaign builds across platforms like Salesforce and HubSpot. This automation has freed the team from repetitive data collection and system clicks, allowing them to focus more on strategy, data validation, and enablement.

Read original
Cursor

Introducing Grok 4.5

SpaceXAI has released Grok 4.5, an advanced AI model capable of handling difficult, long-running tasks across various domains beyond software engineering, including data science, finance, and legal work. This mixture-of-experts model was trained on a broader dataset and uses reinforcement learning to creatively solve problems, utilize tools, and recover from mistakes. Grok 4.5 is available today in Cursor via subscription plans, with usage doubled for the first week and pricing based on input/output tokens. It's noted that an earlier snapshot of the Cursor codebase was accidentally included in training, potentially impacting its CursorBench scores.

Read original
Claude

Working at the frontier: How Thomson Reuters builds AI for high-stakes professional work

Thomson Reuters is building AI for high-stakes professional work in fields like legal and accounting, prioritizing accuracy and verifiability for consequential decisions. Joel Hron, their CTO, emphasizes a "Fiduciary-Grade AI" approach, which combines frontier models like Claude with Thomson Reuters' authoritative content, deep domain expertise from over 2,700 experts, and seamless workflow integration. This method ensures AI outputs are transparent, verifiable, and defensible, enabling professionals to significantly reduce research time and gain reliable starting points while maintaining ultimate accountability for their work.

Read original
Claude

Bringing Claude Code and Claude Cowork to government

Claude Code and Claude Cowork are now available in public beta for government agencies via "Claude for Government Desktop," delivered through a FedRAMP High authorized environment. These tools enable public sector teams to modernize software systems with Claude Code and delegate tasks like memo creation, RFP reviews, and casework directly from their desktop using Claude Cowork. The expanded offering includes robust governance capabilities for administrators, enhanced security features like tamper-evident audit logs, and flexible billing options designed to align with government appropriations. This launch aims to simplify the acquisition, authorization, and allocation of AI for government missions.

Read original
Claude

Choosing a Claude model and effort level in Claude Code

In Claude Code, model selection dictates the AI's inherent capabilities and knowledge base, with larger models offering greater capacity for complex tasks. Effort level, distinct from mere "thinking time," controls the breadth of work Claude undertakes for a request, including the number of files read, tools used, verification steps, and task completion before checking in. Users should choose models based on task complexity, opting for larger models for ambiguous problems and smaller ones for routine work. Adjust effort when Claude fails by skipping necessary actions like reading files or running tests, rather than due to a lack of core understanding.

Read original
Claude

Claude Cowork is coming to mobile and web

Claude Cowork is now available on mobile and web, allowing users to access sessions and files across any device and continue tasks in the background, even when offline. This expansion enables work to flow seamlessly, with Claude prompting users for decisions on their mobile devices as needed. The platform is primarily used for everyday knowledge work like business operations and content creation, and beta access is rolling out to Max users first.

Read original
Claude

How people are using Claude Cowork

Claude Cowork, an AI tool extending agentic capabilities to a chat interface, is predominantly used by knowledge workers for "the work around the work" – tasks that are crucial but not their core responsibilities. A May 2026 study of 1.2 million sessions revealed that "business process and operations" constituted the largest usage category at 33.4%, followed by "content creation and copywriting" at 16.4%. These two areas combined make up roughly half of all usage, indicating Claude Cowork's effectiveness in facilitating connective tasks such as drafting reports, building slide decks, and organizing information to advance projects.

Read original
Claude

A field guide to Claude Fable 5: Finding your unknowns

The blog post introduces "Claude Fable," a model whose performance is often bottlenecked by a user's ability to clarify "unknowns"—the gap between the provided prompts (the "map") and the actual codebase constraints (the "territory"). It emphasizes that effectively working with Fable is an iterative process of identifying these unknowns, categorized into known and unknown types, across pre-implementation, during, and post-implementation phases. Users are advised to balance instructional specificity and vagueness, providing Claude with sufficient context to leverage its capabilities as a thought partner for faster discovery and iteration. This approach helps users proactively address ambiguities and improve the quality of agentic coding.

Read original
Cursor

CFOs and the new economics of AI

AI investments are rapidly increasing to become a major operating expense, with global spending projected to reach $1.5 trillion by 2025, yet many organizations struggle to directly trace these expenditures to enterprise-level financial impact. To address this challenge, Cursor is launching the CFO Council, a working group for finance leaders dedicated to aligning AI spend with business value. The council will develop shared frameworks and benchmarks to measure returns on intelligence, optimize variable costs associated with diverse AI models, and manage the uneven distribution of AI's productivity gains. This initiative aims to make AI investments more measurable, predictable, and efficient, ensuring tangible impact amidst rising usage and capability.

Read original