Working Agents, Existential Risk
· The Fluency Briefing
The Fluency Briefing
Your Guide to What's Happening in AI and Why It Matters to You
Wednesday, September 9, 2026

Everyone's arguing about which AI model is smartest. Alibaba's president says that's the wrong question entirely - a model reasons, but somebody still has to chase the shipping container that missed its vessel. Today: Meta's agent that works while its app is closed, Instacart's assistant that turns "what's for dinner" into a cart, and an Anthropic safety researcher putting a number on the apocalypse.
Today in AI:
- Meta's Agent Keeps Working After You Close the App - Meta launched Muse, a proactive agent with its own file system, terminal and browser that can fill forms, complete transactions and message you unprompted. It ships with an activity log, editable memory files, and approval cards for purchases and emails. Ai Meta
- Alibaba's President: Talking Isn't Working - Kuo Zhang argues the industry measures the wrong thing. Benchmarks test reasoning; commerce means catching a spec inconsistency on page four and filing a customs form. One solo founder told Fortune using AI feels like "there are ten of me." fortune.com
- Instacart Wants to Answer "What's for Dinner" - Instacart launched Clementine, an assistant that turns recipes, conversations or a photo of a handwritten list into a ready-to-buy cart. Uber Eats and DoorDash shipped similar tools, all racing to keep you out of ChatGPT. techcrunch.com
- An Anthropic Safety Researcher Names a Number - Evan Hubinger said he sees a greater-than-10% chance AI "could kill all humans" within a decade, while calling risk from current models low. UN adviser Dame Wendy Hall told the BBC she was "shocked," and wondered how much is pre-IPO marketing. bbc.co.uk
- Suno Rebuilt Its Model on Licensed Music - Suno released v6, trained on licensed data from Warner Music Group, BMG and Believe, and is retiring older models entirely. The company settled with Warner last year after label lawsuits over training data. techcrunch.com
- Feds Accuse Chinese Labs of Industrial-Scale Copying - The NSA, CISA and FBI issued a joint advisory alleging DeepSeek, Moonshot, Alibaba and others extracted billions of tokens from Claude, GPT, Gemini and Grok since 2024 to train their own models. Yahoo
- Tiny Models That Never Phone Home - Desert Ant Labs launched with 18 on-device models, including a 2MB language detector that beats a 293MB rival and a transcriber it claims runs 4.7x faster than Whisper on an iPhone. Free up to 100k devices. Desertant
- An AI-Designed Drug Moved a Biological Clock - Insilico's rentosertib improved lung function in a 42-patient fibrosis trial and showed apparent biological age reversal across six proteomic clocks. Its own researchers caution the trial is small, short, and limited to already-sick patients. Insilico
- Gemini Ate Your 14 Travel Tabs - Engadget walks through using Gemini's hooks into Maps, YouTube and Google Flights to build itineraries, track fares and even book rooms. Their advice on the big-ticket bookings: keep a human in the loop. Engadget

Today's Takeaway:
Alibaba's Kuo Zhang put his finger on the gap: benchmarks measure reasoning, while commerce measures whether the customs form got filed correctly (fortune.com). That's why Meta's Muse launch is more interesting for its plumbing than its avatar - activity logs, editable memory files, approval cards before a purchase goes through (Ai Meta). Instacart's Clementine does the same thing for dinner: it builds the cart, you hit buy (techcrunch.com). The honest read is that agent autonomy is now bottlenecked by human review capacity, not model quality. If checking nine copies of yourself takes as long as doing the work, you didn't automate anything - you hired an intern who never sleeps and never explains itself.
"Agents don't fail at thinking. They fail at finishing - and someone still has to check the receipts."
💡 Fluency Moment - Building your AI fluency, one term at a time.

"Benchmark"
In plain English: A standardized test used to measure and compare how well AI models perform. Think of it like: Like grading students with the same exam so you can rank them - but the exam may not reflect real jobs. Why you'll hear about it: Alibaba's president argues benchmarks miss what commerce actually needs: doing the work, not just reasoning.

The Bottom Line
The Pattern: The last several months were about AI's spending and infrastructure. Today the story moved into the approval queue - Meta, Instacart and Alibaba all shipped or described agents whose real constraint is how fast a human can sign off.
Our Call: By December 9, 2026, at least one major agent product will ship a batch-approval or auto-approve tier explicitly framed as reducing review fatigue. More likely than not - the friction is too obvious to leave alone. We'll grade this one in a Friday digest.
Your Move: Open Instacart's Clementine or Gemini's trip planner tonight - five minutes - and give it one real task you'd normally do yourself. Count how long the review takes. That number, not the demo, is your automation math.
What We're Working On
✨ Founding Cohort Special - 60% Off! - Use code MAF20 to join for just $20/month (regularly $50). Get weekly group sessions & workshops, self-paced courses for all levels, access to tools & templates, challenges with peer feedback, and 24/7 support community. → Join Now
✨ Free 30-Minute AI Consultation - Discover how My AI Fluency can help your business unlock the potential of AI. We'll discuss your goals, explore practical AI opportunities for your industry, and outline clear next steps. → Schedule Free Call
✨ How AI-Fluent Are You? - Test your AI fluency with our interactive quiz. See how you stack up and discover what to learn next. → Take the Quiz
💬 Community | 📞 Book a Consultation | 🌐 Website

Fluently yours, The My AI Fluency Team