Ask ChatGPT who the best plumber in your city is, then ask the identical question next week, and there is a coin-flip chance any given business from the first answer shows up in the second. I can say that with numbers behind it because we measured it. In August 2026 Spearleaf ran the question "who is the best {service} in {city}" across 216 US markets, 12 service types in 18 cities, recording what the ChatGPT web app's own responses recommended. Then we came back 4 to 9 days later and asked every market the identical question again.
So does ChatGPT change its recommendations? Yes. Within a week it had rebuilt roughly half of every local list.
The setup was deliberately boring. Same question, same phrasing, same markets, days apart. Each answer produces a small panel of recommended business cards, usually around five, and we recorded every card. We ended up with 214 markets where both asks returned a valid panel, which is the paired sample every number below comes from. The full method, all the numbers, and the limitations live on the study page at how ChatGPT picks a local business.
One framing note before the results. This study describes one pinned model over one week in August 2026. It is a measurement of that window, not a claim about all AI search forever.
Across the 214 paired markets, ChatGPT recommended 974 businesses in the first wave. Only 469 of them, 48.2%, were still recommended when we asked again days later (95% CI 44.5% to 51.8%).
The overlap between the two lists in each market, measured as Jaccard similarity where 1.0 means identical lists and 0 means no shared names, averaged 0.36 (median 0.33, CI 0.33 to 0.39, n=214). For comparison, a business owner who saw their name in a ChatGPT answer on Monday had roughly even odds of being anywhere in the same answer the following week.
The top of the list churned even harder than the membership. The same business held the number one card in both waves in just 73 of 214 markets, 34.1% (CI 27.6% to 40.2%). Two markets out of three crowned a different winner within the week.
That matters because the first card is the one a real user acts on. If your pitch, or an agency's pitch, rests on "ChatGPT says we're number one," the honest follow-up question is "as of when?"
We also ran comparison arms at shorter distances to see whether the churn needed a week to build. It did not.

Panel overlap (Jaccard) against the same market's first-wave answer at each horizon: same day from a second account 0.43 (n=20), rephrased question about 2 days later 0.32 to 0.40 by phrasing (n=20 each), night ask about 3 days later 0.38 (n=37), identical question 4 to 9 days later 0.36 (n=214).
Read that left to right and the striking part is how flat it is. The same question asked the same day from a second account already overlapped at only 0.43. Waiting a week barely moved it. Across gap lengths of 5 to 9 days the overlap stayed in a flat band of 0.30 to 0.42, so the churn is not a slow decay you can outrun by checking more often. It is baked into how each answer gets assembled.
The night arm adds a wrinkle worth knowing. In those 37 markets, 93.8% of businesses shown during the day were open at ask time (152 of 162), versus 31.5% for the night ask (51 of 162). Whether you are open right now appears to shape who gets recommended right now.
Here is the stabilizing counterweight. When a business did appear in both waves, its position was fairly durable. Rank agreement on shared cards (Kendall tau) had a median of 0.67 across the 144 market pairs with at least two shared businesses. Panel sizes barely moved either, changing by a mean of just +0.05 cards, and 85 of 214 markets, 39.7%, returned the exact same number of cards. Displayed ratings on shared businesses were nearly frozen: only 60 of 467 changed at all, 12.8%, with a mean shift of 0.037 stars.
So the shape of the answer is stable. Five-ish cards, high ratings, familiar order. What churns is which businesses fill the slots.
The practical upshot is about measurement. Google rankings move slowly enough that a screenshot means something for weeks. A ChatGPT answer, on this evidence, is closer to a weather reading. Real, accurate, and stale within days.
That points toward monitoring instead of snapshots. Run a fixed set of buyer questions on a schedule, log who appears, and look at appearance rate over a month rather than presence on one day. It is the same discipline I described in how to get recommended by ChatGPT, and it sits inside the broader practice covered in what AI search optimization is. Appearing in 7 of 10 weekly checks is a real position. Appearing once is an anecdote.
The encouraging read for owners is that churn means opportunity keeps knocking. ChatGPT rebuilds the list constantly, so a business that was absent last week gets a fresh chance this week, and the businesses that show up most often are the ones whose signals are strong everywhere ChatGPT looks.
The questions owners ask most about whether ChatGPT changes its recommendations, answered from the study data.
One answer tells you almost nothing, because within a week about half of it will be different. A repeated prompt set tells you your actual appearance rate, which businesses keep taking the other slots, and whether the work you are doing is moving that rate. Spearleaf runs exactly that kind of tracked, honest measurement as part of our AI search work. If you want to know where you stand in your market before your competitors think to check, reach out through our contact page and we will map it with you.
This article was written by Joshua Albanese, founder of Spearleaf, a local SEO and marketing agency based in Fort Myers, Florida. Joshua has built six businesses from zero using organic search and content, and he leads Spearleaf's SEO, AI search, and local strategy for service businesses across the U.S.
Yes, and faster than most people assume. In our August 2026 study we asked ChatGPT the identical local question in 214 US markets twice, 4 to 9 days apart, and only 48.2% of the businesses it recommended the first time were still recommended the second time (469 of 974). The overall list overlap averaged 0.36 on a scale where 1.0 means the same list. The answer you get today is a draw from a distribution, not a fixed ranking.
There is no visible update schedule, and our data suggests the list is rebuilt fresh from web retrieval on many asks rather than updated on a cycle. We saw meaningful list differences at every time distance we measured, including the same question asked the same day from a second account (overlap 0.43, n=20). Churn was already present within a week and did not grow much across 5 to 9 day gaps, which points to per-ask variation rather than scheduled refreshes.
It is real evidence that you appeared, and worth celebrating, but it is one draw from a changing answer. In our study the businesses shown in the first wave survived to the second wave only 48.2% of the time (469 of 974), and the number one spot repeated in only 34.1% of markets (73 of 214). Treat a screenshot the way you would treat one good day of rankings. The useful measurement is a repeated prompt set over weeks, which shows how often you appear, not whether you appeared once.
It can. In a 37-market comparison we ran the same question at night and found that only 31.5% of the businesses shown at night were marked open at that hour (51 of 162), versus 93.8% for the daytime ask (152 of 162). ChatGPT appears to fold current open status into who it puts forward, so an ask at 11pm can favor a different set of businesses than an ask at 11am. Complete and accurate hours on your profiles are part of being recommendable at any hour.
Want to know whether ChatGPT recommends your business this week and whether it still will next week? Let's set up honest monitoring for your market.
Learn more →