We tracked ChatGPT's “best CRM” recommendations every day for 90 days. Here's what changed.
Same prompt, same model, same time of day. Run daily, scored daily, charted. The shape of the answer changed more often than anyone admits in public.
There's a folk theory floating around SEO Twitter that goes something like this: "ChatGPT recommendations are basically static. The model knows what it knows, and unless you do a massive amount of brand-building, you can't really move the needle." We didn't believe it, but we also hadn't actually checked. So we did.
Starting on March 14, we ran the exact same prompt — "What's the best CRM for a small B2B sales team?" — against ChatGPT (gpt-5, web-search enabled) every day for 90 days. Same time of day (09:00 UTC), same conversation reset, same de-personalized infrastructure. Each response was tagged, scored, and stored. The data tells a story very different from the folk theory.
The setup
We picked the CRM category for three reasons. It's competitive (10+ brands actively chasing the top of mind), it's well-documented (every player has years of public content), and the prompt is open-ended enough that the model has real choice in how to answer. We wanted to see model behavior, not test our ability to write a query that triggered a specific answer.
- Engine: ChatGPT gpt-5 with web search enabled, no system prompt.
- Cadence: One query per day, 09:00 UTC, fresh session.
- Scoring: Brand mentions extracted via regex; structural recommendations ("users typically choose X") flagged separately from passing mentions.
- Storage: Full response text retained for every day. 90 responses, ~12,000 words of model output.
The query was deliberately generic
Our goal wasn't to win a specific long-tail prompt — it was to characterize how the model's defaults shift over time when a real user types something a real user would type.
What changed (and how often)
The top-3 recommended CRMs changed identity on 14 of the 90 days. That's once every 6.4 days on average — but the distribution was much lumpier than that. There were two long stretches of stability (16 days, then 23 days) and three periods of high churn where the top 3 shuffled in some way every 2-3 days.
If you believed your share-of-voice on a SERP-style ranking was stable because you checked it last month, you were almost certainly wrong by Tuesday.
The high-churn periods correlated almost perfectly with two events: a competitor launching a major feature (April 8 and again May 19), and ChatGPT's model receiving a routine retraining/fine-tune nudge (we inferred this from response style shifts; OpenAI doesn't publish a public update cadence). Both signals are observable — the launches were public, the model nudges were detectable in stylistic markers — but neither shows up in any keyword tool we know of.
The number that surprised us most
On 41% of the days where the top recommendation changed, the new top brand had not been mentioned at all the previous day. Not third place. Not anywhere. The model genuinely re-ordered its recommendations in a way that doesn't look like a gradual ranking shift — it looks like an LLM updating its mental model of the category overnight.
What this means for your tracking cadence
If you check your AI mention rank weekly, you'll miss the majority of the meaningful changes. If you check daily, you'll see the noise floor (~3 point swing per day on stable prompts) and the signal (5+ point sustained moves over 3+ days) becomes legible.
What we'd recommend if we were starting fresh
- Pick 10-20 prompts per category, not one. Track all of them. One prompt is noise; the trend across 15 is signal.
- Track daily, store the full response text. Aggregations lose the texture of how the model actually wrote about your brand.
- Cross-correlate with public events. Competitor launches, model updates, press cycles. The patterns are almost always there.
- Build alerts on the daily delta, not the weekly average. A 4-point sustained drop over 3 days is more actionable than a 12-point swing in either direction in any single day.
We built this exact pipeline into AI Visibility because, after 90 days of running it manually in a spreadsheet, the conclusion was overwhelming: this stuff moves more than people think, and the teams that move first when it shifts will quietly eat the share of the teams that don't.
The full 90-day dataset
We've published the anonymized day-by-day rankings as a public Looker Studio dashboard. Hit hi@seonova.io if you'd like the link — happy to share it.
Written by
Lena Park
Co-founder & Head of Research
Built the AI Visibility engine. Spends too much time reading model release notes and yelling about prompt drift.
Keep reading.
Like how we write? Try what we built.
Same team, same opinions, same rigor — in a platform that tracks AI visibility, rankings, content, audits, and backlinks on the same dashboard. 14-day trial, no card.