24 September 2026
OpenAI and Anthropic public-post watch
Checked 24 September 2026 at approximately 3:05 am AEST (5:05 pm UTC on 23 September). Priority window: the preceding 24 hours. This is an incremental update to the 21 September bulletin.
Headlines
-
OpenAI launched GPT-6 Sol and GPT-6 Luna, cutting its mid- and low-cost API prices in half — Sam Altman said both models improve on their GPT-5.6 predecessors in intelligence, alignment, coding, computer use and work output, while Noam Brown highlighted how sharply Luna’s price has fallen. Standard API pricing is now $2 per million input tokens and $10 per million output tokens for Sol, and $0.10/$0.50 for Luna; cached inputs cost $0.20 and $0.01 respectively. Both support a 1.05-million-token context window and 128,000 output tokens. They are available in the API and are rolling out in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu; Free and Go users can try Luna in the desktop app. Tibo Sottiaux also announced a banked usage reset for Plus, Pro and Business accounts. OpenAI’s updated safety material classifies both models as High—but below Critical—in cyber and bio/chemical capability, and reports improved jailbreak resistance alongside some weaker HealthBench results associated with much shorter answers. These are primarily OpenAI-run evaluations, so the launch claims should not be read as uniform real-world gains. OpenAI’s original post · Altman’s capability and price post · Brown on Luna pricing · Sottiaux on pricing and the reset · launch report · API changelog · safety appendix
-
Anthropic released Claude Opus 5.5 just outside the strict window, emphasizing efficiency, alignment and forthcoming 5.5 models — The launch post appeared about 34 minutes before the cutoff. Anthropic says Opus 5.5 performs at roughly Fable 5.1’s level on most work, produces output more than 30% faster than Opus 5, and costs about 40% less on typical workloads. API pricing is $4 per million input tokens and $20 per million output tokens, with cache reads cut 60% to $0.20; an optional mode runs up to 2.5 times faster at double the token price. The model is available across Claude and the major cloud platforms, while five-hour limits increased for Pro, Max and Team and subscription users received a banked rate-limit reset. Anthropic says METR and Frontier Design evaluated the model before release and that it achieved the company’s best automated behavioral-audit score, but it still applies Fable-like cyber, biology and anti-distillation safeguards. Mike Krieger said Sonnet 5.5 and Haiku 5.5 are due in the coming weeks. Benchmark and cost-per-task comparisons remain mostly company- or partner-run. Claude’s launch post · Krieger on price and limits · Krieger on Sonnet and Haiku · Anthropic’s launch report
-
Boris Cherny used Opus 5.5 with Lean and TLA+ to hunt concurrency and state-management bugs in the Claude Agent SDK — Cherny said a couple of short prompts produced sixteen pull requests fixing bugs and race conditions after the model formally modelled the SDK. He does not know either formal language well himself, and presented the exercise as evidence that agents can make formal verification more accessible. The post is an unusually concrete developer-workflow example, but it is still a self-reported demonstration: it does not show how many patches survived review, how severe the bugs were, or whether the approach generalizes. Cherny’s post and video
-
Steve Yegge prefers Fable’s deliberate workflow to Astra’s eager execution for serious work—for now — In his first substantive main-feed update since the previous bulletin, Yegge praised GPT-6 Astra’s writing and said it had already surprised him, but contrasted its tendency to start immediately with Fable’s habit of recording and planning work before acting. He described Astra as fun but comparatively unnerving and said he would keep Fable for serious work. This is one experienced user’s early workflow impression, not a controlled model comparison; it is most useful as a signal that initiative and persistence style can matter as much as raw capability in agentic systems. Yegge’s post · follow-up on verification habits
Scan notes
The OpenAI launch, Cherny’s formal-verification example and Yegge’s comparison fall inside the strict 24-hour window. Opus 5.5’s official announcement is included because it was the largest Anthropic development since the previous bulletin and missed the cutoff by only about 34 minutes. Alex Albert and Sholto Douglas also posted Opus 5.5 Blender, pixel-art and 3D demos, but these were illustrative examples rather than independently measured capability results. Sam Altman’s separate comments about OpenAI’s internal mobility and startup-like organization were fresh but did not announce a product, research result or policy change. Representative leadership, product, developer, research, safety and policy timelines at both companies were checked, including Sam Altman, Greg Brockman, Tibo Sottiaux, Romain Huet, Noam Brown, Boaz Barak, Dario Amodei, Jack Clark, Boris Cherny, Amanda Askell, Mike Krieger, Alex Albert and Sholto Douglas; Steve Yegge’s feed was reviewed separately.