Getting your B2B software cited when the answer replaces the list
How AI assistants build vendor shortlists, why entity consistency matters more than page optimisation, and a measurement approach that works without click data.
Andrei Saioc B2B & SaaS SEO consultant
Ask an assistant for the best tools in your category. If you are not in the three names it returns, you were never in the evaluation, and there is no impression in any dashboard to tell you it happened.
That is the shape of the problem. It is not that AI search is taking clicks, though it is taking some. It is that a shortlist is now being assembled before anyone visits a website, and the criteria for making that shortlist are not the criteria for ranking first.
How the shortlist gets built
Simplifying considerably, three things are happening at once when an assistant answers a category question.
There is what the model absorbed during training, which is a diffuse aggregate of everything written about your category up to some cutoff. There is live retrieval, which pulls current pages, usually a mix of search results and specific trusted sources. And there is the synthesis step, which has to reconcile the two and produce something confident.
The practical consequence is that what gets said about you across the open web matters more than what you say about yourself. Your own site is one source among many, and often not the most trusted one for a comparative question. Review platforms, listicles, forum threads, and community discussions carry disproportionate weight because they read as third-party assessment.
This is uncomfortable if your instinct is to optimise pages. Most of the work here is not on your domain.
Entity consistency, which is the boring foundation
Models are consensus machines. If five sources say your product is a marketing automation platform and three say it is a CRM, the answer will be hedged or wrong, and hedged descriptions do not make shortlists.
So the first pass is unglamorous data cleaning. What your company does, who it is for, where it is based, when it was founded, what it costs, and what category it sits in, stated consistently on your site, in your structured data, on Crunchbase, on LinkedIn, on G2 and Capterra, on Wikidata if you qualify, and in your press materials.
We audit this by asking four assistants to describe the client’s product and comparing the answers to reality. On the last six accounts, every single one had at least one significant inaccuracy propagating — a discontinued pricing tier, a wrong founding year, a category description from a positioning the company abandoned in 2022.
Those errors come from somewhere specific and usually from one or two sources with disproportionate reach. Finding and fixing the source is more effective than adding a correction to your own site.
What makes content citable
From testing across client accounts, the content properties that correlate with citation are fairly consistent.
Specific numbers with attribution get lifted more than qualitative claims. “Median implementation is 19 days across 340 deployments” is quotable. “Fast implementation” is not.
Clear, self-contained definitions near the top of a page. Retrieval systems chunk documents, and a chunk that answers a question completely without needing surrounding context is far more useful to them.
Comparative and balanced framing. Pages that acknowledge trade-offs get cited more often than one-sided ones for comparison queries, which is the same effect that makes honest comparison pages convert better with humans.
Recency signals that are real. Dated content with a genuine last-reviewed date.
Structured data, which is the one genuinely technical lever. Organization, Product, SoftwareApplication, FAQ, and Dataset markup give machine readers unambiguous facts rather than requiring them to infer from prose.
The third-party half
If a listicle titled “12 best [category] tools” is the source three assistants cite, being on that list matters more than anything on your own site.
The practical work is: identify which sources actually get cited for your category, which takes a few hours of running prompts and reading the citations. Then get accurate, current information into those sources. That means claiming and maintaining your review platform profiles properly, reaching out to listicle authors with corrections when your entry is wrong or missing, participating honestly in the community forums where your category gets discussed, and encouraging customer reviews through legitimate means.
None of this is new marketing work. What is new is knowing which sources to prioritise, and the citation data tells you.
Measuring it
There is no rank tracker for this, so build a panel.
Take 100 to 200 prompts your buyers plausibly use — category shortlists, comparisons, feature questions, “is X good for Y” questions. Run them weekly across the assistants your market uses. Record whether you appear, in what position within the answer, what is said about you, and which sources are cited.
That gives you four trackable metrics: citation rate, share of voice against named competitors, sentiment, and source concentration. All four move, and they move faster than organic rankings do, in both directions.
Then watch referral traffic from assistant domains separately in analytics, and add the self-reported attribution field on your demo form, which catches a share of this that nothing else will.
Should you block the crawlers
For almost every B2B software company, no.
The argument for blocking is that models are reproducing your content without sending traffic. The argument against is that being absent from the retrieval corpus of the systems your buyers use to build shortlists is a visibility decision with a direct commercial cost.
If your business is selling content itself, that calculus changes. If your business is selling software and your content is marketing, blocking is choosing to be invisible in a growing share of the research process.
What I would push back on is making that decision by accident, which is what a default robots.txt from a template amounts to. Look at what you are currently allowing and disallowing, and decide deliberately.
How much of this is actually new
Perhaps 30%. Authority, clarity, third-party corroboration, and accurate structured data all mattered before and matter more now.
The genuinely new parts are the measurement approach, the weight on sources you do not control, and sentiment as a variable — an assistant can mention you and describe you unfavourably, which has no analogue in a blue-link result. Those are worth building a specific workstream around. The rest is SEO done properly.
Andrei Saioc
B2B & SaaS SEO consultant
Four years working exclusively on B2B and SaaS search. I run every engagement myself, which means the person who writes the strategy is the person who implements it and the person who explains it when a month goes badly.