Need to know
- xAI's Grok Voice Think Fast 2.0 tops the Tau Voice agentic benchmark at 56.5%, beating OpenAI and Google at half the price.
- Microsoft's Satya Nadella confirmed a unified Copilot super app merging chat, GitHub Copilot, and AI agents will ship this year.
- Anthropic's Claude Mythos found a key-recovery attack on HAWK-256 post-quantum crypto and sped up a known AES-128 attack by up to 800x.
- GPT-5.6 Sol autonomously rewrote OpenAI's production GPU kernel, cutting external inference costs 20% before the model was public.
- Paxos disclosed its Slack-native Hoplites agent now generates 15% of all merged pull requests across company repositories.
New Releases
xAI released Grok Voice Think Fast 2.0, a speech-to-speech model that scores first on the Tau Voice agentic benchmark at $0.08 per audio minute.
- Benchmark results: 82.9% on Artificial Analysis Speech-to-Speech Index (second overall behind Alibaba's Qwen Audio 3.0 at 84.1%), 56.5% on Tau Voice agentic performance, beating GPT-Realtime-2.1 at 45.7% and Gemini 3.1 Flash at 37.7%.
- Pricing and rollout: At $0.08/min, it is roughly half the cost of comparable OpenAI voice API tiers; existing Grok Voice users auto-migrate on August 5 unless they opt out.
Satya Nadella confirmed on Microsoft's earnings call that a single Copilot app will merge Copilot chat, GitHub Copilot, and AI agent capabilities before year end.
- Scope: The unified surface covers both consumer and enterprise users, collapsing what are currently separate products and distribution motions into one application.
- Strategic read: This is Microsoft's answer to the fragmentation critique — and a direct competitive move against the standalone productivity AI apps gaining enterprise traction.
Google Cloud moved Agent Memory Bank, Agent Runtime, Agent Identity, Agent Gateway, Agent Registry, Agent Evaluation, and Agent Observability to general availability on its Gemini Enterprise Agent Platform.
- What's now GA: The seven capabilities span the full agent lifecycle from identity and runtime through evaluation and observability, giving enterprise builders a governed deployment stack rather than experimental APIs.
- Competitive context: The GA timing follows Okta's Agent Gateway launch and Microsoft's agentic security announcements, signaling that governed agent infrastructure is becoming table stakes across cloud platforms.
Anthropic launched two foundational certification exams — Claude Certified Developer Foundations and Claude Certified Architect Foundations — designed to test practical skills over knowledge recall.
- Market move: Certifications create a skills-verification layer that enterprises can use in hiring and vendor evaluation, mirroring the playbook AWS and Salesforce used to entrench their ecosystems.
- Timing: The launch coincides with Project Glasswing expanding Mythos Preview to roughly 50 enterprise security partners, accelerating demand for verified Claude expertise.
Funding
Centralize closed $19M led by NEA with Salesforce Ventures and Slack co-founder Stewart Butterfield participating, signaling that relationship mapping for complex enterprise deals is where GTM AI investment is concentrating.
Encore AI (formerly Insait IO) raised a $30M Series A led by Team8 for its enterprise agentic customer interaction platform, which studies top human agents to train AI that autonomously closes leads and recovers balances in regulated sectors.
Case Studies
Paxos disclosed that Hoplites, its Slack-native AI coding agent, now accounts for 15% of all merged pull requests company-wide, having generated over 2,000 PRs since launch with more than half merged.
- Operational model: Hoplites runs as an autonomous layer on top of LLMs inside Slack, handling tasks beyond code generation including test writing and documentation, without requiring engineers to context-switch to a separate tool.
- What 15% means: At a company the size of Paxos, that share implies the agent is handling work that would otherwise occupy multiple full-time engineers — and it is measurable enough to report externally.
Cloudflare reported a tenfold increase in bug-finding efficiency through Project Glasswing, Anthropic's program giving select vendors early access to Claude Mythos Preview for vulnerability discovery.
- Scale of findings: Across roughly 50 partners in the program's first month, hundreds of high- or critical-severity vulnerabilities were found, with Microsoft separately reporting Mythos surfaced 90 vulnerabilities in April alone — faster than patch teams could respond.
- Verification cost: Anthropic noted each cryptographic finding cost roughly six figures in API usage and hundreds of staff hours to verify, framing Mythos as a research-grade tool, not a fire-and-forget scanner.
Trending on X
- OpenAI cannibalizing enterprise customers The AI community is debating whether OpenAI and Anthropic's direct product expansions — into coding, sales, and vertical workflows — make them structural threats to the SaaS companies built on top of them, with the Figma precedent cited repeatedly.
- ARC-AGI-3 benchmark legitimacy François Chollet posted a clarification distinguishing legal general-purpose API settings from illegal benchmark-specific harnesses after OpenAI claimed a state-of-the-art score on GPT-5.6 Sol via a custom run not verified by ARC Prize.
- GPT-5.6 rewrote its own inference kernel Practitioners are reacting to OpenAI's disclosure that GPT-5.6 Sol autonomously optimized its own production GPU kernel and cut inference costs 20%, treating it as the clearest evidence yet that AI self-improvement is moving from research to operations.
- Frontier models withheld from public Discussion is building around the observation that the strongest AI systems are tested and deployed internally — or in programs like Project Glasswing — long before public access, effectively bifurcating the market into those with and without frontier access.
- Sam Altman hypes science acceleration models Altman posted that OpenAI is 'very close to models that will significantly accelerate scientific discovery,' amplifying debate about whether the academic researcher program announced today is a genuine science play or a distribution move.