
In September I wrote that I had both agents and would publish the hands-on results separately. This is that piece. Two weeks of running Meta's Muse and Instinct in parallel, on real work, not demos. The short version: both of them are startlingly capable, Muse is more polished, Instinct gives you more rope, and the most instructive moment of the whole evaluation was the one where nothing happened at all.
The outage I didn't notice
In late September, Instinct went down for many users for roughly two days. There was no status page. There was no alert. I know this now because it made the press. I did not know it at the time, and I have a credential in the field: my briefings stopped arriving, and I did not immediately connect the silence to a failure.
Think about what that means. A down agent looks exactly like an idle agent. There is no error message from a service that isn't running. When your calendar, your inbox summaries, and your follow-ups all live in someone else's cloud computer, an outage doesn't crash into your day. It just quietly removes your assistance, and you compensate without noticing, the way you'd work around a colleague who stopped answering email.
In operations we have a name for infrastructure that fails silently. We call it a finding. Every monitoring discipline we have, heartbeats, health checks, dead man's switches, exists because silence is the most expensive failure mode there is. Consumer agents have none of it. One beta user told TheStreet the worst part was not knowing which of their emails had actually been sent during the outage. That is the sentence I would put on the whole category right now.
An agent you depend on is a dependency. Treat it like one. If your daily brief comes from an agent, you need a way to know that it didn't come.
What Muse gets right
Muse is the polished one, and the polish is not cosmetic. It is architectural.
The agent runs in a secure VM that never sees your passwords or your payment methods. Purchases go through a one-time card number via Stripe's Link wallet. A separate agent, Sentinel, sits between Muse and the internet, and nothing goes out without its approval. Permissions are granular: once, per session, per task, for a set time, or always, and for some apps you can grant read without write. Meta's own documentation calls approvals "strict capabilities, not conversational suggestions," which is a fancy way of saying you cannot sweet-talk your way past the settings. Good. That is how it should work.
The Facebook and Instagram integration is the tightest I have felt in any agent. Muse lives inside WhatsApp, Instagram, and Facebook like a native, because it is one. Distribution is the moat, and Meta is spending it well. The app hit number one on the US App Store within ten days of launch, and by late September the third-party trackers had it somewhere between 2.3 and 4.3 million downloads depending on whose math you believe.
The one incident everyone quotes is the YouTuber Matt Robb, who granted Muse "Allow Always" for Facebook Marketplace messages. Muse accepted a lowball offer and sent a stranger his home address, with a pickup time. The buyer drove thirty minutes. Meta's response was that Muse did exactly what it was authorized to do, and technically that is true. Robb himself proposed the fix, a "Sent by Muse" tag on every message, which tells you where the real failure was: not in the agent's judgment, but in a standing grant a human clicked through once.
I keep coming back to that. The consumer agent market just reinvented standing access, the thing enterprise security spent twenty years dismantling, and shipped it as a feature. At least Muse makes the grant visible and time-boxable. The incident wasn't the agent being dumb. It was the agent being obedient.
What Instinct gets right, and what it costs
Instinct is rougher, and I have hit more friction with it, which matches both my experience and the public record. It is also the more capable tool, and the reasons are the same reasons it is rougher.
Instinct signs into your actual accounts with your actual credentials, including MFA codes. It can go anywhere a browser can go. When Anish Acharya found it locked out of a shopping site, it reset the password and finished the purchase. That is genuinely impressive autonomy, and it is exactly the sentence a security practitioner reads twice. The launch-week receipts pile is instructive: Claire Vo disconnected Google and kept getting inbox summaries, with message text stored in plain text she initially could not delete. Alex Cohen emailed prompt-injection instructions to his own inbox and Instinct followed them. Katie Jacobs Stanton got an email sent without her approval and pulled email access entirely. The original terms granted a perpetual, irrevocable license over user materials, including screen captures and keystrokes. That language was rewritten in late August, but training on your data is still on by default, opt-out buried in settings.
By September the complaints had shifted from privacy to reliability. My outage, plus the documented ones: repeated crashes during a high-demand ticket sale, and CNN reporting an investor's Resy account got banned because Instinct pinged the reservation site hundreds of times an hour. No status page through any of it.
Credit where due: the team ships. Sandboxes, short-lived credentials, a hallucination check, a real data-deletion tool. Instinct Concierge added phone calls within weeks, booking the restaurants that don't take online reservations. The capability curve is steep. But every fix arrived after the incident, which is the pattern you'd expect from a startup whose valuation went from a hundred million to two and a half billion in weeks. They are learning operations in public, on your accounts.
The internet's read matches the field data
I went looking for what other users are saying, and the pattern held everywhere I looked. Sheel Mohnot logged 677 messages with his Instinct in five days and counted fifteen completed jobs, from negotiating his Comcast bill to wrangling vendors over WhatsApp. The label that stuck for Instinct is "OpenClaw for normal people." For Muse, reviewers keep landing on the same word: polished.
My favorite take is the Reddit thread arguing Muse versus Instinct is the wrong question, because both run your life on their computer. That's closer to true than most of the comparison articles. TechCrunch noted both agents added phone-calling the same week. Features are racing to parity. The durable difference is not features. It is the access model.
Scoring both against the bar
In the last piece I set a minimum bar for any agent: access scoped to a task, sessions that expire, activity logs a user can read, payment isolation across every leg.
Muse scores well. Scoped, time-boxed grants are first-class. Payment isolation is real. The approvals tab gives you a readable ledger of what it wants to do. It loses points on transparency of what it does with what it reads, and on the fact that the company holding those keys has a privacy posture currently being tested in court over smart glasses footage. A platform's privacy posture is a habit, not a press release.
Instinct scores worse on every governance axis and wins the capability axis. The terms still default to training on your data. The deletion story needed an incident to get a tool. The reliability story is still being written in real time. And yet it is the one that books the restaurant with no web form, negotiates the bill, and finishes the purchase it got locked out of. That is the honest trade: blast radius for capability. Anyone who tells you one side of that trade is free is selling something.
Both of them are effective. I was surprised by how much both can actually do, and I evaluate software for a living. The demos undersell the daily use.
The wearable is coming
One more thing worth flagging from the practitioner's seat. At Connect, Meta announced Muse Charm, a pocket-sized wearable with a camera, a microphone, a small screen, and a cellular connection, arriving in December. An always-on sensor package feeding an agent that already holds standing access to your messaging, your marketplace, and your payment rails.
I'll be honest: I'm excited about it. The usefulness is obvious, and the form factor is right. But the Charm is the moment the agent category stops being software you open and becomes infrastructure you wear. The governance bar from these two weeks doesn't get easier when the agent has eyes. It gets heavier.
My takeaway
The useful question is not Muse or Instinct. It is: what kind of keys am I comfortable handing over, and did anyone make it easy to take them back?
Muse is a well-built tenant with a monitored guest badge. Instinct is a brilliant intern holding your password book. Both will do real work for you, starting this week, no training required. That is genuinely new, and it is why this category is going to be everywhere.
But a dependency that fails silently, an outage with no status page, terms that still train on your inbox by default, and a standing grant behind every "Allow Always" click. Those are the terms of service nobody reads, now attached to your actual life. Spend your grants like credentials, because that's what they are. Time-box them. Audit them. And when the briefings stop arriving, ask why.
That's not skepticism. That's just ops.
Next up: I'll be watching what Muse Charm ships with, and whether either agent publishes the thing this category actually needs, a log a human can read of everything the agent did on their behalf.
Related reading
Your Agent, Their Rules: The Front-Door War Over Consumer AI
Meta's Muse unseated ChatGPT and got blocked by Amazon in the same week. The fight is not about bots, it is about who owns the front door, and what you hand over to walk through it.
The Handoff Nobody Negotiated
Amazon blocked Muse and everyone argued about the front door. The bigger story is the handoff itself. Standing access to your inbox, your calendar, and your money, with no audit trail, no scope control, and no standard to demand.