This feels like the busiest week of the year, we’re in that frantic ‘post labor day’ heading into conference season with Holiday 2026 bearing down us - Thanksgiving is 72 days way - basically we’re in that critical '2 months left to get all our ducks in a row’ period.
Going to RetailClub or GroceryTalk or ShopTalk?
At ReFiBuy we have folks at all three shows - Jason and I are at RetailClub where we’ll be doing live podcasting of both Retailgentic and the Jason and Scot show. If you’re not going, don’t worry, we’ll be streaming many of these pods live from the show floor (well sand, it’s on the beach!) so be sure to subscribe to our YouTube channel and turn on notifications.
Heading into Holiday 26…
This is a great time to update some of our big picture Agentic Commerce Optimization thoughts so everyone can focus on getting the best bang for the buck in this short time. The Good news? → The easiest thing to change and also the highest impact is how you describe your products to agents.
In this part 1 of the 2 parter we focus on: "Personal Shopping Agents enter the Arena”
In part 2 coming after RetailClub on 9/29 we’ll cover the rest of the update and have a corresponding podcast/video presentation that puts both parts together and gives you all the data and slides/graphics so you can use them in your Q4/27 planning.
New: Latest datapoints on Agentic Commerce Scale and Growth
Updated: Anatomy of a Product Card
Updated: 5 Levels of Autonomy
Updated: ACO Everywhere strategy map
Updated: ACO 11 Steps Sequential Strategy
Personal Shopping Agents Enter the Arena
In the last two weeks we have seen an entirely new category of Agentic Commerce come seemingly out of nowhere that we call Personal Agents - these are horizontal agents meaning they do many more activities, such as email, calendaring, event planning, travel booking, but they do all anchor on and toug their Agentic Commerce capabilities. These are different than the vertical shopping agents:
Top Five Brand and Retailer Questions about Personal Agents used for Agentic Commerce
These agents are gaining traction rapidly and generating a lot of question for retailers and brands. I’ve rolled them into these 5 questions:
Q1: Should I block these?
Q2: How do I track personal agent transactions?
Q3: How do I optimize for personal agents?
Q4: How do these personal agents work and where are they going?
Q5: How is this different than Answer Engine-driven Agentic Commerce?
In Part 1, of this series we’re going to dig into and answer these questions, then in part 2, we’ll zoom back out and answer question 5 by updating what’s going on in the rest of the world of Agentic Commerce which is not holding still at all either.
To answer these questions, we need to dig into background, talk a bit about Instinct and GrokBot as those are new here on Retailgentic (We have detailed Muse walk through with purchase flow here).
Personal Agent Backgrounder
To answer the big five questions, we need to look at:
The underlying exponential model improvements - it’s the rising tide lifting all the capabilities we cover here.
History of Personal Agents - The history here will help us see where this is going and how fast.
What’s a ‘Harness’? - You’ve probably seen this floating around, Harnesses and Loops are all the rage in the AI builder community. We stay focused on 'commerce’’ here at Retailgentic, but to answer the five questions, it will help to have a working knowledge of what a Harness is, what it does and what it doesn’t do.
Two Modes of Agentic Browser Use - The most important part of these harnesses that impacts us and leads to the answers for our five questions is Agentic Browser Use.
Underlying it all: Exponential LLM Improvements
At this point, you probably gloss over the charts showing exponential growth of LLM intelligence. Some look like this:
Or this→
For example, Claude Haiku came out 10/25 and scores an indexed 638 on this meta-benchmark aggregator and now Fable 5.1 is at 7,798. The models are improving at a compounding 10x rate and everytime we think there’s a slow down, a new innovation, busts through that. This time next year, it will be at 80,000. 🤯
But, just like humans, the model’s knowledge is based on what it learns, or is trained on and this creates what AI researchers call ‘jagged edges’. LLMs like Fable 4.1, Sol 5.6 can write code for hours if not days now that’s better than humans, but some simpler tasks, they completely fail at. If you want the best examples of these follow ‘Husk’ on YouTube or X.
Here’s a fun example: (remember this model can also code for days on end)→
How can something so smart also be so dumb? Jagged edges.
One of the edges the model companies have been working hard on is browser/computer use - training the models to understand human interfaces so they can better do work on our behalf. In our world we’re most interested in the Research→Find→Buy cycle, but computer use includes things like: “Create a ticket in my crm system”, “open up SAP and create a new widget”, “Load up my PIM and translate this SKU into 93 languages.”.
In fact, fact the Information recently reported:
The root cause of the emergence of these super smart peronal agents is due to the new (last 6-12 months) focus from the frontier labs on models trained for computer use.
There are ~50 different benchmarks for browser/computer use and this one aggregates them:
To put this in perspective, Gemini Pro 3.1 Preview (the one on the far right that scored 457) came out in 2/26. Today the models are scoring 1720 on an indexed scale, that’s a 4-5x improvement in 10 months. Many of the benchmarks have scores that have a % of tasks completed. These models are coring 93%, 96%+ at tasks that are much much more comlex than buying something online.
Part 1 of the story is in the last 12 months, the underlying agents haven 10x smarter at general intelligence and 4X smarter at computer use. Researching→Finding→Buying a product on a consumer website is child’s play**
**with the right context, memory, tools, etc.
What’s a “Harness”?
That ** above is important. Think of the models as raw intelligence - computer brains, but they kind of all over the place and wild, they need direction. A harness, borrowed from a horse harness metaphor, gives the LLM direction for a specific use-case.
The picture above illustrates how the harness ‘surrounds’ the agent and wraps it with the pieces it needs to guide the raw intelligence, keep it focused, keep it grounded, help it remember where it went right and wrong (writing things downs in a scratchpad turned out to be a huge breakthrough - who knew!).
Harness are increasingly adding advanced capabilities:
Agentic Loop - On a timed or event-based basis, the agent wakes up and takes pre-determined actions. For example, when I get a new calendar invite, do x, y, z or closer to home: if I get an email about a product I’m looking for that’s less than the 30-day average, go buy it.
Multi-player capabilities - On the consumer
Compounding self-improvement -
Browser use -
If you’ve used chatgpt, claude - those are light harnesses. If you’ve used Claude cowork or ChatGPT Codex→Work, those are heavier more meaty b2b harnesses.
To understand the consumer agent harness progression, we need to go way way back to the earliest days of consumer - way back to January of 2025 (or Oct 25 depending on how we’re counting).
Retailgentic’s History of Personal Agents
The pre-cursors to these agents is OpenAI Operator, in fact ChatGPT’s co-founder, Greg Brockman has been saying since early 2025 that ‘computer use’ is a top priority for ChatGPT→
Jan and July you can see OpenAI working on computer use with the introduction of Operator which became Agent. Then they folded that into Apollo - their browser Atlas (here’s our first Atlas LIVE Retailgentic demo if you want to jump in the time machine for a spin.)
IMHO, what really broke this open, was the introduction of OpenClaw in 11/25 which by March 26 had a thriving community of people extending the harness to Agentic Commerce.
In March 26, we posted this and worked with friend Ryan Eade to demonstrate 3-4 use cases that are now table stakes (minus the ray-ban piece) in these consumer personal agents.
After OpenClaw got progress, some ‘clones’ came out that simplified the install, added some multi-player improvements and beefed up the security. Examples in the is category are Hermes, Perplexity Computer and Microsoft Scout is transparently based on OpenClaw. At Google I/O they announced and launched Gemini Spark.
I’ve tried all of these it’s very hard to get the right mix of proactivity/safety/pedantically asking permissions.
On one side you have OpenClaw that’s scary aggressive about access to your computer and all your files. On the other end of the spectrum you have Google Spark which asks for permissions so much and for every task and sub-task that it’s basically unusable (you’ll see in the demo below).
The Q3_26 ‘class’ of personal agents broke through this. They have a great balance of getting up fast, not bothering you for permissions every five seconds for permissions, but asking permissions at the right break points (e.g. hey you’re about to spend money, review this transaction for me).
Two Modes of Browser Use: Server and Local
Back in the Agentic Browser wars days, we spent some time walking through how the first ‘computer-use’ agents used servers to do their browsing. The challenge with this model is it’s easy to block - the destination (e.g. Amazon) or the infrastructure layer (e.g. cloudflare) can detect the bank of IPs via fingerprinting and say: “Oh this is a Meta muse bot, I will shut it down”. What seems to always happens is these systems start with server-based (or a virtual machine, or VM) and then add the ability to use the local machine. When you’re on the consumer’s local machine, you are effectively unblockable - the agent’s traffic looks exactly like a consumer. As a countermeasure, companies are developing technology that looks at how fast you’re using the browser, mouse movements and what-not to try and decide: “human or agent”. The problem with this is ‘false positives’ - I frequently, especially when using some specific sites and retailers get put through Captcha sequence (pain) or the slidey-puzzle thing, etc.
Of note - all of these ‘prove you are a human’ tests have fallen with the growth in computer use, so they are placebos at this point.
This all ties to the ‘track-ability’ of personal agents which also ties to ‘optimization’ so It’s worth going through this diagram in more detail:
The user is on their local computer - they have what we call a consumer IP
They use a personal agent which, even if it’s a desktop or mobile app, is running in a server/vm.
That agent needs to do something with a browser, so spins one up in it’s VM (usually chromium-based so it “looks like” chrome).
For Agentic Commerce the browser tool decides if it wants to crawl the merchant’s website, use UCP (or some other protocol like ACP, WebMCP, etc.) for discovery and transaction or some mix of those.
Alternatively, the agent may decide to connect back to the consumer’s computer and use a local browser, a browser extension, or IP routing.
In fact, literally as I was working on this post Elon conveniently posted this→
They very clearly say the intent is to get around server-side IP blocking - nobody is hiding the ball here.
Demos: Gemini Spark, GrokBot, Instinct and Muse
To demonstrate the agentic commerce capabilities of all these systems, I’ll briefly walk through Spak, GrokBot and Instinct to show you the experience.
The Veggiedent test
Long time readers have probably recognized that one of my favorite use cases is the dog treat →
On the surface this seems easy, but the lowly veggiedent has broken many an agent because, just like all of e-commerce, there is a lot of complexity hidden in this image:
These are not widely available
They have these variations:
FRESH and 2 other ‘styles’ of chew
Dog size - this is small 11-22lb
Pack size - this is the 30 chew
Agents struggle to get this right - spoiler alert, at least one of our agent demos will not survive the Veggiedent test!
Also, I happen to know that Chewy and Amazon battle it out for this SKU and they are two of the best blockers of agents.
Gemini Spark Walk Through With Transaction
After finding 10mins how to get to it (mobile only, personal account not business, pro or something, hidden in a menu, accept/yes/yes/yes/yes/yes/yes/yes/ok/upload pick/accept/ok/ok/ok/phew). I was able to get to shopping:
GrokBot Walk Through With Transaction
I’m in very deep in Grok Bot for work usage - i have a fleet of 15 bots doing things for me every day and collabing on many projects - it’s amazing for that use. case. But let’s try it for ecommerce. I go in and setup a bot, I called mine cleverly: “Scot’s Shopping Bot”
I’m including the setup here because it’s pretty cool and gets a lot of the nagging questions out of the way in one big chunk. The most interesting one→
You can see I freed my shopping bot to buy anything under $50 without bugging me. Let’s hope that doesn’t backfire.
Only Chewy was able to block GrokBot - kudos to the team for figuring out how to get Amazon to work - not easy. (Note: I’m not using the local IP, not available to me at the time of this writing, that would get around the Chewy block though for sure).
From here let’s look at the Amazon buying flow:
To finish, I have to login to my amazon account - I can give this info to Grok to hand over to Amazon or take over. Both Muse and GrokBot have secure VMs that the LLM can ‘write to’, but not read from so their claims are that your credentials are secure. Personally I’m increasingly becoming a fan as a consumer of the Stripe Link System - it’s one click that has a complete review and you’re done - pretty amazing. Speaking of which, let’s look at Instinct now.
Instinct Walk Through With Transaction
First of all, Instinct is in private beta and you need an invite code, they are pretty hard to come by, but if you’ve made it this far and need one - DM me on LinkedIn - we have a pool of Instinct Invites here at ReFiBuy HQ for our BFFs.
Instinct funs right in your iMessages and it’s best characterized as “insanely proactive”. You have to be very clear with it will go and do things you may not have intended. But it always has the best intentions and that proactively has really been a game changer for me - it’s saved my 🥓 a couple of times and I've been using it 2wks.
Instinct famously is dealing with a crushing load of new users and can’t get servers fast enough, so there was a glitch with the image upload that was easy to fix→
Notice it got through to chewy, but not Amazon in this use case.
Then I clicked on the link to see the product before buying and…
Womp womp - Instinct has a bit of a ready/fire/aim thing going on and it got caught in the Veggiedent trap - (these are the extra small vs. small) - third variant got it.
Lesson - always check the transaction before approving. Speaking of which, I ordered a cooler using Instinct for tailgate season and, like Muse, it uses the Stripe Link connection for payment which I really am enjoying→
When I click on that link (I won’t show that as it has tons of PII), it lays out the whole transaction and once you authorize it and TFA, it magically works.
After you order something, Instinct checks your email and pings the order site proactively so you know where it is in it’s journey to you.
Some people have even wired Instinct up to a phone calling service and have it call merchants and do things IRL for it - wild stuff, on my weekend fun experiment list.
Muse Product Walk Through
Ok, I realized that I didn’t show the Veggiedent example in my Muse walkthrough so let’s do that real quick to be fair→
From there, Muse found a $20 off and processed my order with a stripe link!
Recap: Muse vs. Grok Bot vs. Spark vs. Instinct
Spark came in last for me: while it didn’t trip up on Veggiedent, it is so far behind the others, capability-wise it has a lot of room to make up. GrokBot is more a business agent for me, but it did well.
Muse is the real standout. I’m not sure why Chewy/Amazon aren’t blocking it yet, perhaps it’s early, or, if you’re META, you have one infrastructure for FB, WhatsApp, Instagram, Threads, etc. all that traffic, including Muse looks like ‘meta family traffic’ -perhaps it’s hard for the blockers to separate that out?
The real winner: Distribution
The hardest thing to conquer with a consumer business is getting consumer attention. In this field of 4 Gemini and Muse have a huge advantage - existing billions of consumers. All you have to do one on of the existing platforms is put a little ‘try muse now button’ and BOOM, off to the races. In fact, there’s early evidence they are doing exactly that→
Back to the Five Big Merchant Questions Personal Agents and Agentic Commerce
Q1: Should I block these?
A1: For me there are 3 angles on this one: retail/dtc brand/non-direct brand
Retail Answer: This is the toughest question you face right now and comes down to theology and where you are as a company. You see Amazon, Chewy and many of the top retailers in a category increasingly actively blocking. The argument here, is “I don’t want to be a logistics company and give away the front door.”
DTC Brand Answer: For DTC Brands, they are always looking for more distribution, so they tend to open the door wide. Because many of them both DTC and wholesale. they are already used to ‘giving away the front door’, so theologically they are ok - the argument here goes: “meet the customer where they are” and where they are is increasingly using Agentic Commerce across all these new surfaces.
Non-DTC Brand Answer: This is the easiest one - you should make sure your brand site is open so all agents can get all the product info they can consumer and work with your retailers to optimize for their on-site agents and if they are embracing other outside agents (Walmart is prime example), collab with them on that.
At the end of the day, the theology question seems to come down to: How badly do you want new customers? Can you afford to lose the sale to a site that doesn’t block?
Q2: How do I track personal agent transactions?
A2: In their current configuration you can see with GrokBot, where i’m using my credentials, all you see that’s different is a new login from a new device. Eventually GrokBot’s IP range will become known, but again, they are already implementing local which gets rid of any hope of tracking that model.
I think the good news is once you start using these super efficient, fast and easy agents, you open up your mind to getting out of the middle of the checkout - the Stripe Link thing got me over that very quickly.
We’re looking into specifics, but there’s definitely a way to track that so at least you can bucketize “agents that paid with Stripe Link” and put a number on this category and see how big it gets, bounce it against your customer file and see the existings vs. new and what-not.
Q3: How do I optimize for personal agents?
These bots all see the meta data on your site. There’s speculation Muse is already or soon will start looking at UCP, they use it in other parts of the FB family. The WebMCP people see a world where everyone switches to that.
To be honest, the computer/browser use is so good, I’ve been using this very heavily for 30 days and it’s not getting stumped with anything. The discovery process does break down sometimes and you can see some of the agents are using intermediate product databases (Muse and Spark do this as best I can CSI) and in there they maybe using the UCP part of discovery.
We’ve said this here maybe 1000 times, but it works here too - the single thing you can due regardless of retailer/brand/etc. is work to make your catalog agent friendly - keep chipping away at those 5 levels of enrichment (details here).
The harnesses need context and content to chew on - feed them well!
Q4: How do these personal agents work and where are they going?
You can see in the timeline here that in just 12 months these agents have gotten very sophisticated. Using the Retail 5 levels of Agentic Commerce Autonomy as a guide→
These consumer agents are squarely stradling 3/4/5 with different elements. Where they need to go next:
A ‘wallet’ for your favorite preferred types, loyalty programs, etc.
Multi-cart/wish-list and deal watching (like Gemini multi-cart)
Extracting your purchase history from email (they can do it, but requires prompting and they are forgetful compared to Answer Engines)
Prediction of needs - Insight is getting there, very cool and helpful.
Q5: How is this different than Answer Engine-driven Agentic Commerce?
That is the main topic of Post 2, we’ll be back after RetailClub!





































