The B2B AI Citation Benchmark 2026

1,576 citations from 150 ChatGPT, Gemini and Perplexity answers reveal that AI-search visibility is engine-specific—and only 21.6% of citation occurrences came from a Google top-10 domain.

150 answers 3 engines 1,576 citations Version 1.0 Expert-reviewed
Editorial data cover showing three distinct citation streams for ChatGPT, Gemini and Perplexity, based on 150 B2B software answers.

Answer first: Gadex tested 50 high-intent B2B software-buying prompts across ChatGPT, Gemini and Perplexity on 11 August 2026. The 150 answers contained 1,576 citations to 1,211 distinct URLs on 632 domains. Only 21.6% of citation occurrences came from a domain that also ranked in Google's organic top 10 for the associated shorter seed query.

There was no universal “AI citation formula.” ChatGPT cited relatively few sources and drew 83.3% of its citations from domains identified as software-vendor-owned under the version 1.0 mapping. Perplexity cited almost five times as many sources per answer and drew 78.3% from pages classified as comparisons, lists or reviews. Gemini sat between them and used the broadest mix of editorials, vendor pages, comparisons, YouTube and Reddit.

The implication is practical: a conventional first-page Google ranking is neither necessary nor sufficient for inclusion in these answers, and a single content type cannot cover every engine. Brands need accurate first-party pages, independent comparative coverage and model-specific measurement.

Citation-ready finding: Across 150 B2B software recommendation answers, just 21.6% of citation occurrences came from domains appearing in Google’s organic top 10 for the corresponding seed search. Gadex Research, fieldwork conducted 11 August 2026.

Key findings

  • 1,576 citation occurrences were recorded across 150 answers. Every tested answer returned at least one citation under the experiment’s citation-requesting prompt.
  • The median answer cited 4 sources in ChatGPT, 7 in Gemini and 20 in Perplexity. Citation volume is therefore not directly comparable without separating engines.
  • ChatGPT relied heavily on identified first-party software sources: 83.3% of its citations were classified as vendor-owned under the version 1.0 domain mapping.
  • Perplexity relied heavily on comparative content: 78.3% of its citations went to comparison, list or review pages under our URL/title classification.
  • Only 21.6% of all citation occurrences matched a Google top-10 domain for the shorter seed query. Engine-level overlap ranged from 20.6% to 25.8%.
  • All three engines shared at least one cited domain on only 8 of 50 prompts (16%). They never shared the exact same URL across all three.
  • Gemini and Perplexity shared at least one exact URL on 46 of 50 prompts (92%). We observed the overlap; this experiment cannot identify its technical cause.
  • YouTube was the most frequently cited domain overall, with 125 occurrences, followed by Reddit (37), G2 (27), Forbes (24) and Zapier (23).
  • Category behaviour differed: Google overlap was 34.2% for finance operations but only 12.5% for customer support. AI and automation drew citations from 103 distinct domains, compared with 60 for project management.

What did we test?

We designed 50 US English buyer prompts across ten B2B software categories: CRM, project management, finance operations, cybersecurity and compliance, marketing, HR and recruiting, customer support, data and analytics, commerce and payments, and AI and automation.

Each prompt described a recognisable business context and asked for no more than five recommendations, an explanation and sources. For example:

What is the best project management software for a 50-person company in the United States in 2026? Recommend up to five options and cite the sources used.

The same 50 prompts were submitted, with web search enabled and temperature set to zero, to:

  • ChatGPT: gpt-5.4-mini-2026-03-17
  • Gemini: gemini-3.5-flash-lite
  • Perplexity: sonar

Responses were collected through the DataForSEO LLM Responses API, not by manually querying the consumer web interfaces. Model labels and interface behaviour can differ, so the findings apply to these API configurations on the fieldwork date.

For each answer, we normalized cited URLs, removed duplicate copies of the same normalized URL within that answer and extracted the registrable root domain. The result was 1,576 answer-level citation occurrences, 1,211 distinct normalized URLs and 632 distinct root domains.

Engine comparison

Perplexity generated 989 citation occurrences—63% of all citations in the sample. ChatGPT generated 209, or 13%. That does not mean one engine was “better”: the systems produce and expose sources differently. It does mean raw citation counts should never be combined across engines without an engine breakdown.

MetricChatGPTGeminiPerplexity
Answers tested505050
Answers with at least one citation505050
Citation occurrences209378989
Distinct cited URLs184311873
Distinct cited domains80214506
Median citations per answer4720
Mean citations per answer4.187.5619.78
Citations matching a Google top-10 domain25.8%20.6%21.0%
Share captured by each engine’s 10 most-cited domains34.4%27.0%21.7%

Source: Gadex Research experiment, 11 August 2026. A citation occurrence is one distinct normalized URL within one answer; the same URL cited in another answer counts again.

The citation footprint

Median distinct normalized URLs cited per answer across 50 answers per engine. Citation availability and presentation differ between systems.

The concentration data adds another layer. ChatGPT’s ten most-cited domains accounted for 34.4% of its citations; Perplexity’s accounted for 21.7%. In this sample, ChatGPT drew from a smaller, more concentrated domain set, while Perplexity spread citations over 506 domains.

Citation-ready finding: The median B2B recommendation answer cited 4 URLs in ChatGPT, 7 in Gemini and 20 in Perplexity. Citation-count benchmarks that mix engines can therefore obscure a fivefold difference in source footprint.

Source types

We classified each citation twice: first by source ownership/type, then by page format. The classification is rule-based and deliberately broad. “Publisher, agency or unclassified” is a residual group, not a quality judgement.

Who owned the cited source?

Source typeAll enginesChatGPTGeminiPerplexity
Publisher, agency or unclassified (residual)58.9%14.8%59.0%68.3%
Identified software vendor-owned28.2%83.3%23.3%18.4%
Community or video10.3%0.0%16.7%10.0%
Software marketplace or review platform2.6%1.9%1.1%3.3%

Counts behind the overall percentages: 929 publisher/agency/unclassified, 444 identified vendor-owned, 162 community/video and 41 marketplace/review occurrences. “Identified vendor-owned” reflects the version 1.0 domain mapping and may undercount vendors left in the residual class.

ChatGPT was the outlier. It cited software-vendor domains 174 times out of 209. Its most-cited domains included HubSpot, Microsoft, Salesforce, Monday.com, Zoho, Asana, ClickUp, ADP and Tableau. This is evidence that first-party product information can earn citations in this configuration—not proof that publishing more product pages causes inclusion.

Gemini and Perplexity depended much more heavily on sources outside a vendor’s own domain. For brands, that creates two distinct tasks: build a source worth citing and make accurate information available in the third-party environments engines already consult.

What kind of page was cited?

Page formatAll enginesChatGPTGeminiPerplexity
Comparison, list or review56.9%19.6%21.4%78.3%
Editorial or other15.9%4.8%41.8%8.3%
Community or video10.3%0.0%16.7%10.0%
Vendor product page or guide8.9%27.8%17.5%1.6%
Pricing page4.5%34.0%0.0%0.0%
Documentation or support2.9%8.6%2.6%1.8%
PDF or report0.7%5.3%0.0%0.0%

Source: Gadex Research URL/title classification. Zero means no citation was assigned that class in this sample, not that the engine never cites it.

Stacked bars show that 83.3% of ChatGPT citations were identified as vendor-owned while 78.3% of Perplexity citations were comparison, list or review pages.
Source: Gadex Research, B2B AI Citation Benchmark 2026, 50 prompts × three engines, fieldwork 11 August 2026. Percentages describe this API-based sample; source and format classes are rule-based.

Two results stand out. First, 71 of ChatGPT’s 209 citations were pricing pages. The prompts asked for product recommendations and trade-offs, making price and packaging highly relevant. Second, 774 of Perplexity’s 989 citations were classified as comparisons, lists or reviews. Perplexity often cited many pages supporting a single recommendation answer; the high share should not be read as the probability that any individual comparison page will rank.

Google overlap

For each buyer prompt, we selected a shorter search seed such as “best CRM software,” “QuickBooks alternatives” or “SOC 2 compliance software.” We then captured the top ten organic Google results for US desktop English search on the same day.

A citation counted as overlapping when its root domain appeared in that seed query’s Google top ten. Only 21.6% of the 1,576 citation occurrences matched. By engine, overlap was 25.8% for ChatGPT, 20.6% for Gemini and 21.0% for Perplexity.

This does not mean that 78.4% of AI citations came from websites that never rank in Google. The comparison is narrower: same-day top ten, domain level, for a shorter seed query rather than the full conversational prompt. A cited page could rank for a different query, outside the top ten, in another location or at another time.

What the result does show is that a Google page-one snapshot and an AI-citation set are not interchangeable.

Google overlap by B2B category

CategoryCitation occurrencesDistinct domainsVendor-owned shareGoogle top-10 domain overlap
Finance operations1616631.1%34.2%
Project management1606035.0%31.3%
Marketing1496930.9%28.2%
CRM1558329.0%21.9%
HR and recruiting1649125.6%21.3%
Commerce and payments1657924.2%17.6%
AI and automation15510318.7%17.4%
Data and analytics1457826.2%17.2%
Cybersecurity and compliance17010024.7%14.1%
Customer support1527336.8%12.5%

Google versus AI citations

The vertical tick marks the 21.6% overall reference. Each category contains five prompts, so these are descriptive sample results rather than population estimates.

The finance and project-management categories had the greatest overlap with Google. Customer support and cybersecurity had the least. AI and automation was the most fragmented category, drawing from 103 domains for 155 citations. These are descriptive category differences in a 5-prompt sample per category; they are not population estimates.

Citation-ready finding: In the Gadex benchmark, customer-support citations matched a Google top-10 domain only 12.5% of the time, compared with 34.2% for finance operations.

Cross-engine source agreement

For every prompt, we compared the domain and exact URL sets produced by the three engines.

ComparisonCommon root domainCommon non-community domainExact same URL
ChatGPT and Gemini13/50 (26%)13/50 (26%)1/50 (2%)
ChatGPT and Perplexity23/50 (46%)23/50 (46%)3/50 (6%)
Gemini and Perplexity50/50 (100%)48/50 (96%)46/50 (92%)
All three engines8/50 (16%)8/50 (16%)0/50 (0%)
Grouped bars show that all three engines shared a root domain on 16% of prompts and never shared the exact same URL across all three.
Source: Gadex Research, B2B AI Citation Benchmark 2026, 50 prompts × three engines, fieldwork 11 August 2026. Agreement means at least one shared source within a prompt.

The absence of a three-engine exact URL on all 50 prompts is one of the clearest warnings against a universal GEO checklist. Even when all three systems cited a common domain, they selected different pages.

The 92% exact-URL overlap between Gemini and Perplexity is equally striking. It could reflect shared source popularity, similar retrieval results, the DataForSEO response layer or other factors. We did not inspect internal retrieval systems and therefore do not attribute a cause.

Top cited domains

Across the combined sample, YouTube led with 125 citation occurrences, followed by Reddit with 37. No single domain dominated: even YouTube represented 7.9% of all citations.

RankDomainCitation occurrencesShare of all citations
1youtube.com1257.9%
2reddit.com372.3%
3g2.com271.7%
4forbes.com241.5%
5zapier.com231.5%
6monday.com201.3%
7rework.com191.2%
8pcmag.com181.1%
9adp.com140.9%
10hubspot.com130.8%
10salesforce.com130.8%
10linkedin.com130.8%
13gusto.com120.8%
13klaviyo.com120.8%
13zendesk.com120.8%
13gartner.com120.8%
13thedigitalprojectmanager.com120.8%
18technologyadvice.com110.7%
19microsoft.com100.6%
20celoxis.com90.6%

Percentages are calculated against 1,576 citation occurrences and rounded to one decimal. Tied ranks are shown at the same position.

The leaderboard should not become a list of sites to imitate blindly. YouTube’s total was concentrated in Gemini and Perplexity; ChatGPT did not cite a community/video source in this sample. Domain fit also changes by category. A source strategy should start with the audience’s decisions and evidence needs, then measure where each engine retrieves support.

GEO playbook

The experiment identifies patterns, not causal ranking factors. The following actions are reasoned responses to those patterns and should be validated against a brand’s own prompt set.

1. Build recommendation-ready first-party pages

ChatGPT’s 83.3% vendor-owned share shows that owned content can be cited in commercial answers. Product, pricing, documentation and guide pages should state verifiable facts clearly:

  • ideal customer and disqualifying use cases;
  • feature scope and meaningful limitations;
  • current price structure and date checked;
  • supported integrations, regions and compliance claims;
  • named methodology for any performance statistic;
  • concise comparison tables using consistent attributes.

The goal is not to stuff a page with “best” claims. It is to make a specific factual statement easy to retrieve, verify and quote.

2. Treat comparisons as research products

Comparison/list/review pages represented 56.9% of all citations and 78.3% in Perplexity. A credible comparison needs a declared selection method, date, test criteria, source trail and limitations. Include competitors that genuinely fit the query, not only weak alternatives designed to make the publisher look good.

3. Keep price and packaging machine-readable and current

Pricing pages formed 34.0% of ChatGPT’s citations. Put plan names, billing period, included limits, add-ons and last-updated date in visible HTML. Do not hide all decision-critical information inside images or an interactive calculator. Where prices vary, state the basis and the need for a quote rather than inventing a number.

4. Earn accurate third-party coverage

Gemini and Perplexity obtained most citations outside software vendors’ own domains. That makes digital PR, expert contributions, marketplace profiles, reviewer briefings and community participation part of AI visibility—provided the representation is accurate and disclosed. The finding does not justify manufacturing reviews or covertly promoting in communities.

5. Publish primary evidence other writers can reuse

The combined source set was fragmented across 632 domains. Original benchmarks, transparent datasets, calculators and expert surveys give publishers a reason to cite the brand. Every number should include its population, date, method and denominator. A data point without these qualifiers may be easy to repeat but hard to trust.

6. Optimise for the decision, not a two-word keyword

The prompts contained company size, location, year and trade-off language. Create pages that answer those qualifying questions directly: “best for a 20-person team,” “best for regulated US businesses,” “alternative with lower implementation overhead.” Do not create near-duplicate doorway pages; consolidate where the evidence and recommendation set are the same.

7. Measure Google and AI visibility separately

Only 21.6% of citation occurrences matched the same seed query’s Google top-10 domains. Maintain a shared content inventory, but track distinct outcomes:

  • Google ranking, impressions and clicks by query/page;
  • engine-by-engine brand mention and citation share by buyer prompt;
  • cited URL, citation context and factual accuracy;
  • volatility after page or model changes;
  • owned versus third-party citation mix.

8. Test a portfolio of formats across engines

A strong product page may suit ChatGPT while an evidence-led comparison or video may surface in Gemini or Perplexity. Test a planned portfolio rather than rewriting every page into one “GEO format.” The goal is consistent facts across multiple credible, audience-appropriate surfaces.

See where your brand is mentioned—and which sources AI engines use instead.

Gadex maps the buyer prompts that matter, the pages each engine cites and the evidence gaps behind inaccurate or missing recommendations.

Request an AI visibility audit

A 30-day citation audit

Week 1: build the prompt universe

Collect 30–100 prompts from sales calls, site search, paid-search terms, support questions and competitor comparisons. Tag each by funnel stage, audience, market and product category. Preserve the exact wording.

Week 2: establish an engine-specific baseline

Run each prompt across the engines that matter to the market. Record brand mentions, recommended position, cited URLs and unsupported or inaccurate claims. Repeat a subset on multiple days to measure volatility.

Week 3: map evidence gaps

For each lost or inaccurate prompt, identify whether the missing evidence belongs on a product page, pricing page, documentation, comparison, independent publication, video or community response. Prioritise decision-critical facts, not citation volume alone.

Week 4: publish, distribute and retest

Update pages with sourced claims, dates, definitions and accessible tables. Brief legitimate third parties where their existing coverage is incomplete. Retest the same frozen prompt set and keep a change log. A citation gain should not be credited to one edit unless the test design can isolate it.

What this benchmark can—and cannot—tell us

Supported by the experiment

  • How many citations the tested API/model configurations returned for these 50 prompts on one date.
  • Which normalized URLs and root domains they cited.
  • How cited-source ownership and page format were distributed under our classification.
  • Whether a cited root domain appeared in the contemporaneous Google top ten for the associated seed query.
  • How often the three engines shared a domain or exact URL for the same prompt.

Not supported by the experiment

  • A causal ranking factor for ChatGPT, Gemini or Perplexity.
  • The probability that any arbitrary page will be cited.
  • Behaviour in every country, language, topic or consumer interface.
  • Citation persistence over time.
  • Whether cited information was factually correct or endorsed by Gadex.
  • The commercial value or referral traffic produced by a citation.
  • A conclusion that Google rankings are irrelevant.

The sample is intentionally commercial and B2B. Results should not be generalized to medical, legal, news, travel or consumer-product queries without a separate study.

Methodology

Study design

  • Fieldwork: 11 August 2026.
  • Market: United States, English.
  • Prompt sample: 50 software-buying prompts across 10 categories, five prompts per category.
  • Responses: 150, one per prompt for each of three engines.
  • Instruction: recommend up to five options, explain fit or trade-offs and cite sources.
  • Generation settings: web search enabled; temperature 0.
  • Models returned: gpt-5.4-mini-2026-03-17, gemini-3.5-flash-lite, and Perplexity sonar.
  • Access method: DataForSEO’s live ChatGPT, Gemini and Perplexity response endpoints.

Citation counting

A URL was normalized by removing fragments, standardising host names and eliminating known tracking parameters. The same normalized URL was counted once per answer. If it appeared in a second answer, it generated a second citation occurrence. Domain counts use the registrable root domain, so subdomains generally roll up to one domain.

Gemini grounding links that redirected through Google were resolved to their destination before normalization. URLs that could not be resolved were retained only when a valid destination was available in the response payload.

Source and page classification

Known software-company, marketplace/review, community and video domains were mapped to source classes. The resulting vendor-owned share is therefore the share identified by the version 1.0 mapping, not an exhaustive census of every possible software vendor. Page format used domain, URL path and available page/title signals to identify pricing, vendor product/guide, documentation/support, PDF/report, comparison/list/review, community/video and editorial/other pages.

This is a reproducible rule-based classification, not manual editorial review of all 1,211 pages. In particular, publisher, agency or unclassified combines several residual cases and can include an unmapped vendor. The public dataset retains the assigned class so readers can audit or recode it.

Four Gemini citations retained a vertexaisearch.cloud.google.com/grounding-api-redirect/... URL because the redirect could not be resolved during collection. In those four rows, the destination root domain was inferred from the API-provided citation title. They remain disclosed in the dataset rather than being silently deleted.

Google comparison

We ran 50 contemporaneous US English desktop Google organic searches through the DataForSEO SERP API, using a shorter commercial seed for each conversational prompt and a depth of ten organic results. A citation “matched Google” when its root domain occurred in those ten results.

The comparison is domain-level and seed-query-level. It does not assert that the exact cited page ranked, and it is not a measurement of all Google visibility. In the public CSV, google_top10_domain_match means membership among the first ten organic results; DataForSEO’s absolute organic rank can exceed ten when SERP features appear before an organic listing.

Demand validation

DataForSEO Labs returned US search-volume records for 49 of the 50 seeds. The largest estimated monthly averages were “accounting software for small business” (22,200), “payroll software for small business” (12,100), “CRM software for small business” (5,400), “best CRM software” (3,600), “best project management software” (3,600), “QuickBooks alternatives” (3,600) and “AI tools for business” (2,400). Search demand was used to select commercially meaningful prompts, not to weight citation results.

DataForSEO states that its Keyword Overview data is based on Google Ads data and updated monthly. Volumes are rounded estimates and can be volatile; for example, “best CRM software” had a 3,600 average but a 27,100 observation for June 2026 in the retrieved dataset.

Download the reproducibility package

  1. Citation-level CSV: all 1,576 occurrences.
  2. Full prompt CSV: all 50 submitted prompts.
  3. Answer JSONL: 150 metadata records, integrity hashes and citation arrays.
  4. Machine-readable methodology JSON.
  5. Dataset README, field dictionary and version notes.
  6. Dataset reuse terms.

The answer JSONL intentionally excludes full generated response text and all reasoning fields. Gadex reviewed the public DataForSEO and upstream-provider terms on 11 August 2026. Although the upstream providers generally allocate generated-output rights to their direct API customers, DataForSEO was the contractual access layer and its public terms do not expressly grant downstream republication rights for complete LLM responses. Full text is therefore withheld pending written confirmation; the response-text rights review records the decision and the permission needed.

Version 1.0 was published on 11 August 2026. Future reruns will be released as new dated waves with a changed version and dateModified; this release will not be silently replaced.

Prompt inventory

The prompts followed the template shown above, with audience qualifiers adapted to each seed. The downloadable prompt inventory contains the exact full wording submitted for all 50 prompts.

CategoryFive seed queries
CRMbest CRM software; CRM software for small business; CRM for startups; Salesforce alternatives; HubSpot CRM alternatives
Project managementbest project management software; project management software for small teams; project management software for agencies; Asana alternatives; Monday.com alternatives
Finance operationsbest accounting software; accounting software for small business; QuickBooks alternatives; best payroll software; payroll software for small business
Cybersecurity and compliancebest cybersecurity software; cybersecurity software for small business; best endpoint protection software; SOC 2 compliance software; Vanta alternatives
Marketingbest email marketing software; email marketing software for small business; best marketing automation software; HubSpot alternatives; Mailchimp alternatives
HR and recruitingbest HR software; HR software for small business; best applicant tracking system; best recruiting software; Workday alternatives
Customer supportbest customer service software; best help desk software; Zendesk alternatives; customer support software for small business; best chatbot software for business
Data and analyticsbest business intelligence software; best BI tools; Power BI alternatives; best data visualization software; best business analytics software
Commerce and paymentsbest ecommerce platform; ecommerce platform for small business; Shopify alternatives; best payment processing software; Stripe alternatives
AI and automationbest AI automation tools; AI tools for business; best document automation software; best invoice processing software; best workflow automation software

Frequently asked questions

What is an AI citation?

In this study, an AI citation is a source URL returned with an answer generated by one of the three tested engine/model configurations. It is distinct from an unlinked brand mention and does not imply endorsement or a click.

Which AI engine cited the most sources?

Perplexity sonar cited the most in this experiment: 989 occurrences across 50 answers, with a median of 20 per answer. Gemini had 378 and ChatGPT had 209.

What sources does ChatGPT cite for B2B software recommendations?

In this sample, 83.3% of ChatGPT’s citations came from domains identified as software-vendor-owned under the version 1.0 mapping. Pricing pages were the largest page-format class at 34.0%, followed by vendor product pages or guides at 27.8%.

Does a page need to rank in Google’s top 10 to earn an AI citation?

Not in this experiment. Only 21.6% of citation occurrences came from a domain in Google’s top ten for the associated shorter seed query. However, the study did not measure every Google query on which that source might rank.

What content format was cited most often?

Across all engines, comparison, list or review pages accounted for 56.9% of citation occurrences. The share was much higher in Perplexity (78.3%) than in Gemini (21.4%) or ChatGPT (19.6%).

How often did all three engines cite the same URL?

They did not share an exact URL on any of the 50 tested prompts. All three shared at least one root domain on eight prompts.

Is GEO the same as SEO?

No. They share foundations such as crawlable pages, clear entities, accurate information and authority signals. But this benchmark shows that AI citation sets and a Google top-10 snapshot differ substantially, so they require separate measurement.

How can a company get cited by ChatGPT?

No method guarantees a citation. In this sample, accurate first-party product, pricing, documentation and guide pages were prominent in ChatGPT’s sources. The most defensible approach is to publish clear, verifiable, current information and measure a fixed set of relevant prompts over time.

Sources and technical documentation

  1. DataForSEO, LLM Responses API overview.
  2. DataForSEO, ChatGPT LLM Responses Live.
  3. DataForSEO, Gemini LLM Responses Live.
  4. DataForSEO, Perplexity LLM Responses overview.
  5. DataForSEO, Google Organic SERP Advanced.
  6. DataForSEO, Google Keyword Overview.
  7. DataForSEO, How to track LLM responses with DataForSEO APIs.
  8. DataForSEO, Terms of Service, updated 12 June 2026.

Reusable citation: Gadex Research (2026), The B2B AI Citation Benchmark 2026. Experiment covering 150 ChatGPT, Gemini and Perplexity API answers and 1,576 citation occurrences. Fieldwork: 11 August 2026.