Skip links

Information Gain Score

Definition

Information Gain Score measures how much genuinely new, useful information a page adds compared with the content already ranking for the same query. It comes from a Google patent on scoring documents by novelty, first filed in 2018 and granted in 2022 (US11354342B2), with a continuation patent following in 2024.

The idea is simple: does this page say something the top results don’t already say, or does it just repeat them? A page can be well-written, comprehensive, and technically solid and still score low if it covers ground the SERP has already covered. What raises the score is originality, such as proprietary research, first-party data, a named framework, or expert findings a reader can’t get anywhere else.

Gestion des backlinks reçus par un site

Origin of the concept

The term borrows from decision-tree machine learning, where information gain identifies which feature most reduces uncertainty at each split of a dataset. Google adapted the same logic to documents: given what a user, or a competing SERP, has already seen, how much does a candidate page reduce the gap in their understanding? Content is turned into embeddings and compared against the existing ´already seen´ set, so the score is always relative to a specific query and its current competitors, never an absolute number.

The Patent Definition

The term Information Gain comes from a Google patent. The signal was always relational, meaning it only means something in comparison to a set of other documents. That is part of why it took years to actually put into practice: Google had to commit to re-ranking against a candidate set at scale, for every single query, rather than scoring a page in isolation.

How the underlying system works

Information Gain Score - how the underlying system works

An important nuance here: the patent describes this working at the level of an individual searcher’s viewing history, not just a fixed query. That means the underlying model isn’t necessarily limited to a single, static ranking for everyone typing the same search. It’s built to account for what a specific user has already looked at, which points toward a more personalised set of results rather than one identical SERP for every visitor.

Consensus content: the baseline before information gain

Before a page can compete on information gain, it first has to clear a lower bar: consensus. Consensus content is content that lines up with what’s widely accepted and verifiable on a topic, properly researched, factually correct, and backed by trustworthy sources.

Consensus and information gain aren’t in tension, even though they can sound like opposites. Consensus is the entry ticket; information gain is what happens after that. A page that contradicts well-established facts has no business ranking regardless of how original it is, and a page that only repeats the accepted facts, without adding anything on top, has no edge over the nine other pages doing exactly the same thing. Getting the facts right earns a seat at the table. Adding something the table doesn’t already have is what earns the top spot.

Why it matters for SEO

For search rankings

Because the score is query-dependent, it can’t be improved just by publishing more. Two pages can be equally long and equally optimised, but the one that introduces new entities, data or angles gets rewarded, while the one that mainly rewrites the top ten does not. This is part of why the Skyscraper technique, publishing a longer version of what already ranks, has become less reliable on its own. The goal has shifted from producing content that’s better than what’s already ranking to producing content that’s different from it. 

Following Google’s March 2026 core update, independent SEO trackers and analysts widely reported that pages built on proprietary research or first-hand case studies gained visibility, while templated or lightly rewritten content lost ground, with figures in the range of 15 to 25 percent visibility gain for original content and 30 to 80 percent losses for templated or AI-farmed content, depending on the tracker. Google has not confirmed any of this officially, so these figures should be treated as informed observation from the SEO community rather than a confirmed statement from Google.

For AI Search

Information Gain matters even more in AI Overviews, Perplexity or ChatGPT search. These systems retrieve a pool of documents and then decide which ones actually shape the generated answer. A page that repeats what another retrieved page already said usually doesn’t get cited; the page that adds something distinct does. In other words, ranking well on Google no longer guarantees visibility inside an AI-generated answer, because Google weighs authority, backlinks and domain age alongside originality, while AI answer engines lean much more heavily on originality alone.

Could it level the playing field?

Traditional ranking factors reward accumulation: the more backlinks, the more domain history, the more authority a site has built up, the harder it is for a newer or smaller site to compete, regardless of what its content actually says. Information Gain works on a different axis entirely. A five-person niche site with a genuinely original dataset can outscore a large publisher’s rewritten summary of that same topic, because the score isn’t asking who you are, it’s asking what you added that wasn’t already there. That doesn’t cancel out the advantage bigger sites still have elsewhere, backlinks and authority still count for plenty, but it does mean originality is one of the few levers a smaller site can pull that isn’t gated behind years of accumulated authority.

Information Gain Score

Higher originality means a higher Information Gain score (AI-Generated)

How does it compare to E-E-A-T, Helpful Content and GEO?

Information Gain doesn’t exist in a vacuum. It overlaps with three signals you’ve probably already heard about, but each of them is really asking a different question:

SignalWhat it is really asking
Information GainDoes this page add something new?
E-E-A-TCan this creator be trusted?
Helpful ContentWas this written for people or for search engines?
GEOWill an AI engine actually cite this?

Information Gain sits underneath the other three. E-E-A-T and Helpful Content are, in a sense, Google’s way of estimating whether a page is likely to be original without measuring originality directly, and GEO applies that same logic to getting picked up inside an AI-generated answer. None of the three replace it, because a page can tick every box on trust, tone and structure and still be a page an AI engine has no reason to cite, simply because nothing on it is new.

A quick example

Take the query “best budget espresso machine.” Most of the top 10 results cover the same five or six machines, list the same specs (bar pressure, water tank size, price), and repeat the same manufacturer claims. A new page that does the same thing has close to zero information gain, it’s not telling the reader anything they couldn’t already find on the first three results.
A page that instead runs its own test, pulling shots on each machine over a month, measuring actual extraction time, and noting which ones failed after 90 days of daily use, has high information gain. It’s not competing on being better-written; it’s competing on containing something the other nine results simply don’t have.

The scoring criteria

Several SEO teams have converged on a practical way to approximate the score manually, by rating a page on five dimensions. Four are scored from 0 to 2, and the fifth from 0 to 1, for a maximum of 9 points. Content scoring 7 or higher is generally considered ready to publish.

Proprietary data (0-2): Does the page contain a dataset, benchmark or statistic the author actually collected or measured, rather than data pulled from other publications?
First-hand evidence (0-2): Does the page describe something the author actually did, tested or experienced, rather than summarising other sources?
Original framework (0-2): Does the page introduce a named methodology, checklist or model, rather than an unlabelled list of general tips?
Expert attribution (0-2): Is the content credited to a specific, verifiable expert, with a track record that can be checked, rather than a generic team byline?
Freshness hook (0-1): Does the page reference something recent, a new release, deadline or data point, rather than only evergreen information?

Common mistakes that lower the score

Not every page that seems original actually adds new information. In many cases, information gain is reduced by patterns like:

  • Presenting third-party statistics without adding your own interpretation.
  • Crediting a named author with no verifiable expertise in the topic.
  • Calling a plain list of tips a framework without giving it a specific name.
  • Publishing evergreen-only content with no fresh data point, release or deadline attached.
  • Treating AI-generated numbers as proprietary. A dataset only counts if it was actually measured or collected.
  • Relying on an AI tool to draft the piece and expecting it to add information gain on its own: A model trained on existing content tends to reproduce the consensus it was trained on, not go beyond it, so anything an AI outputs by default is a reasonable stand-in for what already exists, not for what’s missing.

Signs your content is search-engine-first, not information-gain-first

Google’s own guidance on helpful content lists a set of warning signs for content made primarily to rank rather than to help. A few of them line up almost exactly with what tanks an information gain score: publishing across many different topics in the hope that some of it performs, mainly summarising what other sources already say without adding real value, and updating dates on a page to make it look fresh when nothing substantial actually changed. None of these are hard rules Google has confirmed as ranking factors on their own, but they’re a useful gut check: if a piece would only exist because it might rank, rather than because it says something worth saying, it’s unlikely to score well here either.

How to optimize for Information Gain

Pre-score the brief before writing starts. Three questions do most of the work: what is the reader actually trying to find out, what do the current top-ranking pages already give them, and what could you add that none of them do. That third question is where the actual content decisions get made, deciding what original data, first-hand experience or framework the piece will contribute. Identify the top ten to twenty ranking pages for the target query and treat them as the comparison set, not to match their structure, but to see exactly where they stop short.

Score the finished draft against the five dimensions before publishing. If it lands below 7, the two dimensions worth fixing first are expert attribution and first-hand evidence, since both cost time rather than data: a specific quote in someone’s own words or a detail that proves you actually did the thing rather than described it. Only touch the freshness hook if something genuinely changed since the draft was written; a swapped year in the title without new substance doesn’t move the score. After publishing, keep tracking backlinks, long-tail rankings, and engagement for four to eight weeks, since Information Gain is a moving target rather than a one-time score.

One practical shortcut worth trying: ask an AI model to write a piece on the same topic and intent, then compare it against your draft. Whatever the AI produced by default is a rough stand-in for consensus content. Whatever is left in your piece once you strip out everything that overlaps with the AI’s version is roughly what’s contributing to your information gain.

The limits of this metric

Information Gain isn’t a number you can pull up in a dashboard the way you can check a Domain Rating or a keyword’s search volume. Google never publishes it, so every five-dimension rubric floating around SEO circles is really an educated approximation, not the metric itself. It’s also entirely relative: a page can score high today against this week’s SERP and drop next month once three competitors publish their own original research on the same topic. Worth keeping in mind too, independent research scoring 150 top-ranking pages across ten verticals found that even top-3 results were only moderately original on average, and that longer pages weren’t reliably more original, so ranking well is never proof that a page is scoring well here.

In short

At its core, Information Gain is Google, and increasingly AI engines, asking one blunt question of every page: if I already read the other nine results, why would I need this one? A page that can’t answer that honestly, no matter how polished, is competing on a factor that stopped mattering the moment originality became the differentiator.

Most popular definitions

seo
seo amazon
seo camp
seo e-commerce
freelance seo
seo friendly
international seo
local seo
multilingual seo
off page seo
on page seo
technical seo

vocal seo
youtube seo

Notez ce page

Get your FREE GEO audit! 🚀