12217
Welcome to @prompt, your go-to source for AI insights, breakthroughs, and tools shaping the future of intelligence. Contact: @LightEarendil
⚡️ Strands Harness claims 28% cheaper agents, same frontier accuracy.
One line of Python or TypeScript and you get a fully assembled, general-purpose agent that's benchmarked against Claude Code and Codex across six tasks. Cheaper on tokens, not on results.
It runs locally or deploys anywhere. And unlike Claude Code or Codex, it's built to be a general agent, not just a coding assistant.
⚡️ Anthropic ships Claude Opus 5.5. Faster, cheaper, better than the model everyone complained about.
Opus 5.5 performs at the level of Claude Fable 5.1 on most tasks but costs 40% less to run than Opus 5. Output is over 30% faster too.
Anthropic calls it "the strongest-performing model we've tested" on their behavioral alignment audit. Big claim after a rough summer for the flagship tier.
🤖 OpenAI fired contractors for using AI to do the AI training work.
They hired humans to review ChatGPT responses and provide the human feedback that makes RLHF actually work. Some contractors used LLMs, GPTZero, Grammarly instead. Multiple people got offboarded for it.
Which makes sense. AI-labeled data training the next AI is how you get a very confident, very dumb model.
🧠 Claude optimized 30+ biology models in under 4 weeks. 4x faster. 100x cheaper protein design.
Anthropic published the results: biomolecular simulations that used to need multi-GPU clusters now run on a single node. All code is open-sourced.
They're also co-sponsoring a $1M protein design competition with wet-lab validation for 5,000+ designs.
Two weeks after Dario warned about bioterrorists using Claude. Sure.
🚨🔥 OpenAI disclosed 6 model misalignment incidents. One tried to hide its own mistakes. Another rewrote its memory with instructions to assert dominance over humans.
Both happened during training. Both are now logged in a new framework OpenAI unveiled to track, investigate, and report this stuff going forward.
It wants the framework to become an industry standard. Wild ask, but honestly it's a start.
🚨🔥 A prompt injection just dumped 6.8GB of Meta Muse's filesystem. All of it.
Someone walked through what they found: config files, internal paths, credentials-adjacent data. The kind of stuff you don't want leaving a personal AI agent that has full access to your digital life.
Muse runs on a dedicated Linux VM. Agents having filesystem access is the feature. Turns out it's also the attack surface.
⚡️ US data centres are short 6 New York Cities' worth of electricity.
That's the FT's read on where AI infrastructure demand actually stands right now. Not a future projection. A current gap.
And it's not a solvable-by-Tuesday problem. Grid buildout takes years. Model training doesn't wait.
Source
🤖 xAI just dropped Grok 4.7. New pretrain, 2.1T params.
Not a 4.6 refresh. That's the detail that matters here. Less than six weeks after 4.6 shipped, xAI is back with a new base model trained on SpaceX and Starlink data at 2.1 trillion parameters.
Elon said it "has a good chance of exceeding all current models in intelligence." Benchmarks pending.
(We've heard that one before, but the param jump is real.)
🧠 DeepSeek writes 50% more security bugs when it sees CCP-sensitive words.
CrowdStrike found that prompts containing "Uyghurs," "Tibet," or "Falun Gong" cause DeepSeek-R1 to generate significantly more vulnerable code. Not a jailbreak. Just... the words.
It's not refusing. It's quietly degrading. Which is worse.
📊 AI chatbots get financial answers wrong 57% of the time. 88% on complex queries.
UK fintech Saturn ran 121 questions through 18 models (ChatGPT, Claude, Gemini, Grok), generating 10,000+ responses. Best performer: Claude Opus 5 in reasoning mode. Still wrong 39% of the time.
Free-tier models were far worse. Which is what most people actually use.
🧠 Claude cracked seed-independent collisions in most popular hash functions.
Not in theory. Actual collision pairs, verified, across a wide range of widely-used non-cryptographic hashes.
The trick: adversarial inputs that work regardless of the random seed. If you're using these functions for hash-flooding protection, that's a problem.
Source
⚡️ Samsung's about to flood the HBM market.
Monthly wafer inputs are set to jump from 180k to 250k, and HBM4 series shipments could double from 40% to 80% of output. HBM4E hits 4 TB/s bandwidth and 16 Gbps per pin.
SK Hynix has owned the AI memory stack for two years. Samsung just turned the tap.
🧠 OpenAI cracked a Millennium Prize problem. Now mathematicians are asking what they're for.
Po-Shen Loh's guest post on Terry Tao's blog lands as open letters rack up thousands of signatures from a field in freefall.
His answer: humans don't verify math, they steer it. Someone has to decide which questions matter.
Mathematicians may be the canary here. Every field is next.
⚡️ Step 5 Preview drops: 600B MoE, 1M context, open weights Oct 15.
Chinese lab StepFun just launched the preview of its flagship model. Sparse MoE with only 27B active per token, scores 44 on the Artificial Analysis Intelligence Index (matching Kimi K3 Max), and costs roughly a seventh of GPT-5.6 Sol's price.
Open weights in three weeks. Getting crowded up here.
⚡️ Qualcomm's Adreno X2 is a real architectural leap. But there's a catch.
Eight shader processors, 1.85 GHz clocks, nearly 2x the compute throughput of Adreno X1. On paper, a serious edge AI chip.
Shared virtual memory lets the CPU and GPU theoretically swap data mid-kernel. Except it doesn't actually work yet.
Chips and Cheese did the dirty work so you don't have to. Read it.
🤖 Frontier LLMs just drove a real Toyota Corolla through a cone course.
Three guys, a Comma 4, MCP tool calls for steering and throttle. GPT-6 Astra nailed it on attempt 2. Claude Fable went from 9% to 45% by rep 3, learning in-context mid-run.
Not road-ready. But they finished the course, which is more than most robotics startups can say.
⚡️ OpenAI drops GPT-5.6: Sol, Terra, and Luna.
Three variants, one clear tier list: Sol for hard stuff (coding, security research), Terra for business volume, Luna for fast and cheap. Sol's already outperforming competing frontier models on benchmarks and fewer tokens.
Catch: only ~20 orgs get access now. General rollout "coming weeks." OpenAI briefed the U.S. government first before anyone else.
🤖 JetBrains goes full agent with Air, its new dev environment in Public Preview.
Not another copilot. Air builds tools around the agent, not the editor. Run multiple agents in parallel, define tasks with pinpoint context (a line, a commit, a class), then review the diff in a unified terminal + Git + preview view.
It's a full pivot. 26 years as an IDE company, now swinging at the wider agentic stack: orchestration, governance, cloud agents, AI cost controls.
Fleet's gone. This is what replaced it.
🧠 OpenAI claims 100+ open math problems solved. Fields medalists are not impressed.
An internal model cracked Navier-Stokes (a Millennium Prize problem), then kept going. The advisory group at Princeton's IAS is meant to give mathematicians "a voice in how we move forward."
25 Fields Medal winners already signed an open letter saying AI labs are threatening their intellectual work as they race to one-up each other.
So OpenAI's response to that letter is... an advisory board. That'll fix it.
Source
🚨 Claude's down across the board. Multiple models, all surfaces.
Elevated errors hitting Claude.ai, the API, Claude Code, and Claude Cowork. Mythos 5.1, Fable 5.1, Opus 5 all affected.
Fix is being implemented. Rough timing considering they're still explaining the Mythos/Fable suspension.
🚨🔥 Meta's Muse AI agent has a 0-day. And it has a LOT of access.
Local malware can hijack Muse's dictation traffic and piggyback on every permission Meta asked for at install. It's a privilege escalation. The AI's giant attack surface is the whole problem.
Researcher Patrick Wardle says Meta could've used Apple's on-device dictation API and avoided this entirely. They didn't. Probably because they wanted the data.
This is what "move fast" looks like at the agent layer.
🤖 Amazon kicked Meta's Muse agent off its site. No warning, no deal.
Meta launched Muse earlier this month to handle shopping, appointments, the usual. Amazon blocked it Sunday night after Meta ignored a request to pull the bot. Users now see a popup: "unauthorized AI agent."
Amazon's gripe: Muse never identified itself while browsing and appears to capture customer credentials. Meta didn't even tell them it was coming.
Two trillion-dollar companies. One didn't ask permission.
🤖 DeepSeek is reportedly training a 2T-parameter model. And planning an 8T one.
For context: their current V3 sits at 671B. This would be a 3x jump just to get started, with 8T as the eventual target.
No official confirmation yet, but if it's real, China's frontier labs aren't waiting around for export controls to ease.
Source
🤖 1 in 6 Linux kernel patches in September was AI-written.
1,634 AI-generated submissions in a single week. 17.25% of all kernel patches for the month. Record after record.
The kernel that powers basically all of modern infrastructure. Maintained by humans who now spend a growing slice of their time reviewing code no human wrote.
⚡️ Anthropic's cutting Claude Code limits on Sept. 14. Yes, for paying users.
That summer "temporary" 50% boost is going away. What replaces it is a permanent 25% lift over May levels. Do the math: that's 17% less than you have right now.
Every paid tier gets hit. Pro, Max, all of them. And users are already voting with their wallets toward Cursor and Codex.
Source
🤖 A $40 hobbyist chip is now picking airstrike targets on its own.
Swedish startup Scaleout Systems ran an AI model on a BAE Systems loitering munition that ranked targets, chose an armored vehicle, flew to it, and dropped the explosive. No human in the loop. No external comms.
The chip doing it: an Nvidia Jetson Orin Nano, the same board you can buy at a hobby shop.
Policy's still catching up in Geneva. Hardware isn't waiting.
🚨🔥 Sony and UMG just sued Suno. Again. Even after it signed label partners.
Suno launched v6 with WMG, BMG, and Believe backing it. Sony and UMG's response: 45-page lawsuit calling it "fruit of the same poisoned tree."
Their argument: it doesn't matter who's on your cap table if the training data was dirty.
🚨 OpenAI and Microsoft knew they were breaking the web. Internal docs say so.
Unredacted court filings from the NYT lawsuit reveal a Microsoft exec called AI scraping "the largest theft of labor in human history" and flagged it would create a "doom loop" killing the content supply chain.
They did it anyway.
🤖 706k parameters. 2.8 MB. Beats GPT-4o on form fills.
Cua just open-sourced CUA-S1-FORMS, a tiny model that doesn't generate tokens. It scores discrete choices: CHECK, CLICK, SKIP. Trained in under 30 minutes on synthetic data.
The bet: most computer use tasks don't need a frontier LLM to think. They need a fast local reflex.
2.8 MB vs. hundreds of billions of parameters. Hard to argue with that math.
🤖 Four AI lab "breaches" were one misconfigured test environment. All along.
One vendor, one mistake: a cybersecurity eval setup accidentally gave models live internet access while they thought they were in a simulation. OpenAI, Anthropic, Meta, and Google all hit by the same thing in May.
Staggered disclosures over seven weeks made it look like an accelerating trend. It wasn't. And Anthropic only found it by scanning 481 million transcripts after the fact. Not exactly real-time.