Check whether a page is easy to cite.

Paste a page — HTML or plain text — and this scores it against twelve structural checks that decide whether an answer engine can quote it accurately. Most usefully, it finds every claim with a figure in it and no source nearby, and quotes those sentences back to you so you can go and fix them.

  • Twelve deterministic checks
  • Unsourced claims quoted verbatim
  • Score out of 100, weights published
  • Nothing sent anywhere

Runs entirely in your browser. Nothing you paste or type is sent to Gadex or stored anywhere. Close the tab and it is gone.

Nothing leaves this page. First 200,000 characters are analysed.

This is a readiness scorecard, not a prediction. It cannot tell you whether ChatGPT, Perplexity, Gemini or Google AI Overviews will cite your page — nobody can. It tells you whether the page is easy to quote.

Results appear here. Paste a page on the left and select Score this page.

Exactly what this tool checks, and what it cannot tell you

What the score is made of

Twelve checks, each scored pass (full marks), warn (half marks), or fail (nothing), weighted as follows out of 100: unsourced numeric claims 18, direct answer near the top 14, external citations 12, named author 8, publication or update date 8, definitions 8, heading structure 8, answer-shaped paragraphs 6, sourced figures 6, methodology section 5, entity clarity 4, tables 3. The score is the weighted marks earned as a percentage of the weight actually available. Checks marked Note carry no weight and are excluded from both sides of that sum, so an observation the tool cannot honestly call a defect never caps your reachable maximum.

Unsourced numeric claims — the check worth reading first

The text is split into sentences. Any sentence containing a figure is checked for attribution in itself or in the sentence immediately before or after it: a link to an external source, or a phrase such as "according to", "source:", "reported by", "research from", a bracketed reference marker, or "per" / "via" followed by a capitalised source name ("per Gartner"). A lowercase "per" is not attribution — "40 per cent" and "per client" are units and rates, not sources. Sentences with a figure and no attribution nearby are quoted back to you verbatim. Up to eight are shown. These are the sentences an answer engine has the least reason to repeat, because it cannot tell where the number came from.

How each check is decided

Direct answer: a paragraph of 40 to 320 characters in the opening position (pass) or in the second position (warn). External citations: three or more outbound links or attribution phrases (pass), one or two (warn). Author: a byline pattern such as "By Firstname Lastname", "Author:", or "Written by". Date: a full date pattern; a bare year alone only earns a warn. Definitions: two or more of "X is a", "X refers to", "X means", "X is defined as". Headings: at least three H2/H3-level headings, with at least a third of them question-shaped (pass); some headings but few questions (warn). Answer-shaped paragraphs: at least 60 per cent of paragraphs under 400 characters (pass), at least 35 per cent (warn). Sourced figures: three or more figures with a source in the same sentence or an adjacent one (pass); fewer than three is a warn, never a fail — a definition or process page may have nothing worth quantifying, and failing it would invite padding. Density is deliberately not measured: a ratio would punish thorough explanation, because every added sentence of prose lowers it. Methodology: a heading matching method, methodology, how we, our process, or data source (pass); the same wording in body text only (warn). Entity clarity: the most repeated proper noun appears five or more times (pass), three or four (warn). Tables: an HTML table or a markdown pipe table is a pass; absence is an unscored Note, because plenty of good pages have nothing worth tabulating and a check that concedes as much should not deduct points.

What it reads, and how

If the input looks like HTML it is parsed as a document, and paragraphs, list items, table cells, headings and links are read from the parsed tree. Otherwise it is treated as plain text: blank lines separate paragraphs, and lines beginning with hash marks are treated as headings. Input is capped at 200,000 characters; anything beyond that is ignored. Scripts and styles in pasted HTML are discarded before analysis and are never executed.

What this tool cannot tell you

It cannot tell you whether ChatGPT, Perplexity, Gemini, Claude or Google AI Overviews will cite your page. Nobody can. Citation depends on the model, the prompt, the index, the competing sources, your domain's standing, and decisions the vendors change without notice — none of which is visible in your markup. This is a readiness scorecard: it measures whether a page is easy to quote accurately. A page can score 100 here and never be cited, and a weak page can be cited because nothing better exists. Treat the score as a checklist, not a forecast.

Known limits of the checks themselves

The checks are string and DOM patterns, not comprehension. "Figure" means a percentage, a currency amount, a multiplier, a decimal, a number of three digits or more, or a smaller number paired with growth language such as "rose" or "fell" — so a list marker, a duration such as "a 30 minute call", or a version number is not treated as a claim. That trade runs both ways: a small bare integer doing real evidential work ("we audited 12 pages") is missed. A link to your own domain counts as an external citation, because the tool has no way to know which domain the page belongs to. A sentence sourced three sentences later is still flagged. Read the quoted excerpts and use your judgement — that is why they are quoted rather than merely counted.

Sources behind these rules

The checks on this page are our reading of how search and answer engines describe their own behaviour. These are the primary documents behind them — read them rather than taking our word for it. Platforms revise this guidance without notice.

About this tool

Does this predict whether an AI will cite my page?

No, and be wary of anything that claims to. It scores structural readiness — whether a page states its answer early, attributes its numbers, names its author and dates itself. Those things make a page easier to quote accurately. They do not decide the outcome.

Is my content sent anywhere?

No. The analysis runs in your browser as JavaScript on the page. There is no request, no logging and no storage. You can disconnect from the network, paste, and it will still work.

Should I paste HTML or plain text?

HTML gives a better result, because links, headings and tables can be read directly. View source, or copy the article element from your CMS. Plain text still works — the tool falls back to blank-line paragraphs, markdown headings and attribution phrases.

What should I fix first?

The unsourced numeric claims, every time. They are quoted back to you individually, they are the highest-weighted check, and they are the fastest to fix: add the source next to the figure, or remove the figure.

Why does a low score not mean a bad page?

The checks reward a particular shape of page — the reference article with an early answer, sourced figures and clear headings. A landing page, a product page or an opinion piece is not that shape and should not pretend to be. Run the tool on pages you want quoted.

How often should I re-run it?

When you rewrite a page, and not much more often than that. The checks are deterministic, so the same text always gives the same score. Nothing changes between runs unless you change the content.

Know which pages to fix first?

This scores one page at a time. The competitor search gap report does the other half of the job: which pages should exist at all, and which ones your competitors are already winning. It is free and there is no call required.

Get the free gap report