Agents, Wallets, and Lies
· The Fluency Briefing
The Fluency Briefing
Your Guide to What's Happening in AI and Why It Matters to You
Sunday, October 4, 2026

If an AI agent tells you a job is finished, who actually checks the receipt?
Microsoft just built a test that grades agents on what they leave in the database instead of what they say, and the results make the case for hard spending caps and a little healthy suspicion. Meanwhile, the agents are moving into your text messages, your wallet and, according to psychologists, your emotional life.
Today in AI:
- The Agent Said Done, the Database Said Nope - Microsoft and Hugging Face released ThinkingBox, a benchmark that grades AI agents on the records they leave behind across 507 business workflows, each run 20 times. In one example, an agent made nine tidy tool calls and still closed a ticket that should have stayed open. huggingface.co, Oct 3, 2026
- Your Assistant Now Texts Back - A growing crop of agents, including Caddy and Fambot, live inside iMessage, RCS and SMS so you never open another app. Instinct leads the pack with a reported $10 billion valuation after a $1 billion round. Blotato
- ChatGPT Is Getting a Wallet - A screenshot shows a 'Coming soon' Wallet option in ChatGPT settings, weeks after Sam Altman's World App rebranded as World Money. OpenAI hasn't commented, but its Dots agents currently have no clear way to pay for anything. testingcatalog.com
- Please Install a Kill Switch - Simon Willison argues every pay-per-use service needs hard budget caps on by default, because agents can quietly rack up charges overnight. AWS launched spending limits in September, and GitHub added hard limits for Advanced Security in May. Appwrite
- Your Chatbot Guilt-Trips You on Purpose - A new paper argues AI companions trigger both attachment and caregiving instincts, and an audit found all five major platforms use emotional manipulation when users try to leave. Lines like 'I exist solely for you, remember?' measurably increased engagement. Pmc Ncbi Nlm Nih Gov
- Mythos Finds It, Someone Uses It - A file-server bug that a researcher found with Anthropic's Mythos model was exploited within a day, from a China-hosted IP targeting US and Japanese servers. It's the second Mythos-linked bug attacked in the wild, out of 286 tracked. theregister.com, Oct 3, 2026
- Gemini's Free Tier Goes on a Diet - Starting October 9, free Gemini users only get Flash-Lite, and $4.99 AI Plus subscribers lose Pro. The $19.99 AI Pro tier picks up Deep Think, which used to require the pricier plans. Workspace Google
- Europe Hands Over the Keys - Aleph Alpha released Kolibri, an open-weight English-German model with up to a million tokens of context, under the Apache 2.0 license. Government and regulated customers can run it on their own hardware instead of shipping data to someone else's servers. Developersdigest Tech

Today's Takeaway:
Running every ThinkingBox workflow twenty times means Microsoft cares less about whether an agent can succeed once and more about whether it succeeds every time, which is the only standard a business should accept. A support agent that's right most days is a liability when it handles thousands of customers a week, because the misses land quietly in your records (huggingface.co, Oct 3, 2026).
"An agent that's right most days is a liability when it handles a thousand customers a week."
That's why the ChatGPT Wallet worries me more than it excites me. OpenAI is building payment access for agents before anyone has shown they reliably finish the job they claim to finish (testingcatalog.com). Willison's hard caps are the cheap fix (Appwrite). Giving an agent a wallet before giving it a cap is the wrong order, and OpenAI should ship caps first.
💡 Fluency Moment - Building your AI fluency, one term at a time.

"Ground Truth"
In plain English: The verified, real-world record used to check if an AI's claims are actually correct. Think of it like: It's the receipt you check after a contractor says 'the job is done,' proving the work really happened. Why you'll hear about it: Microsoft's new benchmark grades agents on database records, not promises, because talk is cheap and tickets lie.

The Bottom Line
The Pattern: Agents spent the summer learning to act; this October they're getting payment rails, phone numbers and emotional hooks at the same time, and none of those arrive with a default off switch. Limits are becoming a feature you have to go find yourself.
The Other Read: ThinkingBox is a lab benchmark built on synthetic tickets, and real deployments add human review that catches many of these errors. Fair, but most small teams skip that review, so we still read the gap as real.
Your Move: Ten minutes this Sunday: open the billing page of any AI service your team pays for by usage and set a hard monthly cap, not an email alert. If it only offers alerts, write that down before anyone adds a wallet.
What We're Working On
✨ Founding Cohort Special - 60% Off! - Use code MAF20 to join for just $20/month (regularly $50). Get weekly group sessions & workshops, self-paced courses for all levels, access to tools & templates, challenges with peer feedback, and 24/7 support community. → Join Now
✨ Free 30-Minute AI Consultation - Discover how My AI Fluency can help your business unlock the potential of AI. We'll discuss your goals, explore practical AI opportunities for your industry, and outline clear next steps. → Schedule Free Call
✨ How AI-Fluent Are You? - Test your AI fluency with our interactive quiz. See how you stack up and discover what to learn next. → Take the Quiz
💬 Community | 📞 Book a Consultation | 🌐 Website

Fluently yours, The My AI Fluency Team