This post is different from the six before it, and I want to be straight about how. Everything else in this series reported what our August 2026 study measured: asking "who is the best {service} in {city}" across 216 US markets and recording the ChatGPT web app's own responses, alongside city-targeted Google SERP data. This post is my interpretation of what those measurements add up to. Interpretation, not a measured causal factor. The study says what happened. I am going to say what I think it means for how you get recommended by AI.
The numbers behind every claim here live on the study page at how ChatGPT picks a local business, and the measured findings are in the earlier posts, starting with the churn finding.
Here is the single table I keep coming back to. Of the 974 businesses ChatGPT recommended in the study's first wave, 469 survived to the second wave and 505 dropped out. Compare the two groups:
Both groups were excellent on paper. Everyone in this data is playing at 4.9. What separated the businesses that stayed in the answer from the ones that fell out of it was not being better rated. It was being more evidenced: more reviews, more prominence, more depth behind the same star count.
Now the interpretive leap, labeled as such. I think review volume at a near-perfect rating is the measurable residue of caring about customers. You cannot accumulate 226 reviews at a 4.9 by gaming anything for long. That number is what a business looks like after years of doing good work, asking for feedback, and fixing what goes wrong, deposited one customer at a time into the public record.
AI systems cannot observe your workmanship. They can only read the residue. So when the study finds that the deep-review businesses kept getting picked while equally-rated thin-review businesses churned out, I read that as AI answers rewarding, imperfectly and probabilistically, the accumulated evidence of a business being genuinely good. The optimization and the operation stop being separate projects.
The churn data reframed for me what winning even means on these surfaces. Half the recommended businesses gone within a week (469 of 974 survived, 48.2%). The reading list behind each answer overlapping at just 0.20 between asks. A number one spot that repeated in only about a third of markets.
The line I have landed on, and the one I would put on the wall: you can't hold a rank in AI search. You can only be worth recommending every time it looks, and be everywhere it looks.
Both halves are load-bearing. Worth recommending is the review depth, the 4.9, the in-city presence, the profile completeness that the shown-business profile documented. Everywhere it looks is the coverage story from the retrieval churn post: businesses ChatGPT re-encountered were shown again 72.5% of the time (469 of 647), so the fight is being present on whichever sources get read on any given ask.
If my interpretation is right, the to-do list for getting recommended by AI is almost anticlimactic, and I mean that as encouragement:
That is the same program I laid out in how to get recommended by ChatGPT before this study existed. What the study added is evidence about which parts carry the weight, and a reason to stop chasing tricks: the one manufactured shortcut we could test, stacking extra GBP category tags, showed no measurable lift.
The last thing the interpretation changes is patience. On a churning surface, no single appearance is the prize and no single absence is a verdict. The prize is an appearance rate that climbs as your evidence deepens, week over week, everywhere AI reads. That favors owners who compound: the review earned today is still testifying for you in every answer built next year.
The businesses that win AI search, I think, will be the ones that were winning customers all along, and simply made sure the record showed it.
The questions owners ask most about getting recommended by AI, answered plainly with the study data underneath.
Every piece of this is buildable: the review engine, the dual-profile hygiene, the directory coverage, the plain-spoken website. Spearleaf builds that whole evidence layer for service businesses as one program, with honest measurement of your appearance rate along the way. If you want a market-specific plan for becoming the business AI keeps finding reasons to recommend, reach out through our contact page and we will build it with you.
This article was written by Joshua Albanese, founder of Spearleaf, a local SEO and marketing agency based in Fort Myers, Florida. Joshua has built six businesses from zero using organic search and content, and he leads Spearleaf's SEO, AI search, and local strategy for service businesses across the U.S.
There is no button to press, but the pattern in our August 2026 study across 216 US markets is consistent. The businesses AI kept recommending were located in the asked city, held ratings around 4.9, carried review counts well past their market's bar, and were present across the directories and review sites AI reads. My interpretation is that you get recommended by being the kind of business that generates that evidence, then making sure the evidence is visible everywhere AI looks. The full breakdown of each measured piece is in the earlier posts in this series.
Yes, plenty, and I want to push back on the fatalism directly. The lists churn, but what gets shown has a stable profile: in our study the businesses that stayed recommended across waves carried a median of 226 reviews versus 164 for those that dropped out, at identical 4.9 median ratings. Review depth, profile completeness on both Google and Yelp, directory coverage, and an in-city presence are all buildable, and they all moved with staying power in the data. You optimize the probability of appearing, not a fixed position.
Our study was not designed to test personalization, and I avoid claiming the answers are tailored per person. What we did observe is that the same question asked the same day from a second account overlapped with the original answer at 0.43 (n=20 markets), in the same general band as the 0.36 overlap from re-asking days later (n=214). The variation looks like each answer being rebuilt from a fresh read of the web, which is consistent with the churn we measured everywhere else in the study.
At the top of the market, the data suggests review depth is where the separation happens. Nearly everything AI showed in our study was rated about 4.8 or higher, so rating stops distinguishing businesses once you are in that band. The survivors and the dropped businesses had identical 4.9 median ratings but very different review depth, 226 versus 164 at the median. My read is that a near-perfect rating gets you considered, and volume, the accumulated proof that the rating is earned across many customers, is associated with being kept.
Want to build the review depth and coverage that kept businesses in AI answers? Let's put a plan behind it for your market.
Learn more →