SEO Glossary · GEO

Co-occurrence

Two terms co-occur when they turn up in the same stretch of text more often than chance would allow. In 2026 that arithmetic is a large part of what decides whether a language model associates your brand with a category. It is not a factor Google publishes: it is the statistical residue of everything third parties write about you.

Key takeaways The essentials in 30 seconds
  • A co-occurrence is a text statistic, not a graph edge: no hyperlink is involved, which is exactly why it survives when the anchor is branded or naked.
  • Three parameters define any measurement (window, corpus, association measure) and most SEO discussion of the term never states any of them, which makes the claims unfalsifiable.
  • The strongest recent evidence is correlational: Ahrefs' 75,000-brand study found a 0.664 Spearman correlation between branded web mentions and AI Overviews visibility, against 0.218 for backlinks, with Ahrefs itself warning against reading causation into it.
  • Google has never announced a co-occurrence score. What it did ship, on 3 June 2026, is Search Console reporting for generative-AI surfaces, which is the first official way to correlate earned coverage with AI impressions.
  • Operationally, the paragraph around a link matters more than the anchor: a placement on a media whose vocabulary already sits in the category puts your name inside a document that co-occurs with the right terms by default.
  • Repetition across distinct third-party domains is what moves a distribution. One placement, however well written, does not register in any corpus large enough to matter.
3 questions to test your knowledge Read first, the quiz is waiting at the bottom.
Two-column figure contrasting co-occurrence, the raw symmetrical fact of two words sharing a window, with collocation, a co-occurrence that is statistically overrepresented and felt as such by speakers.
Every collocation is a co-occurrence, the reverse is false: the association measure sorts them out.

What co-occurrence actually measures

Two units of language co-occur when they appear inside the same stretch of text more often than random distribution would predict. That last clause carries the entire concept. Raw joint frequency tells you almost nothing, because the most frequent words in any corpus co-occur with everything. What matters is the gap between observed joint frequency and the frequency you would expect if the two terms were independent.

Three parameters define any co-occurrence measurement, and skipping them is where most SEO commentary on the term falls apart. The window: a sentence, a paragraph, a full document, or a fixed span of n tokens either side of the pivot. The corpus: the closed set of texts you are counting in. The association measure: the formula that turns counts into a score. Change any one of the three and the same pair of words can look tightly bound or completely unrelated. A claim about co-occurrence that names none of them is not a measurement, it is a vibe.

For a working SEO the operational translation is narrow and useful: how often does the brand name land in the same paragraph as the category, the use case, the competitor set and the buying language, across texts you do not own. No hyperlink is required. A co-occurrence is a text statistic, not a graph edge, and that single distinction separates it cleanly from a backlink, which is an explicit, directed, machine-readable pointer with its own accounting.

Google has never published a co-occurrence score and has never described one in its documentation. Anyone selling a « co-occurrence optimisation » line item is selling a proxy for editorial coverage under a more technical name. That does not empty the notion: it relocates it. Co-occurrence describes how retrieval systems and language models build associations. It is not a dial with a value you can set.

Checklist of the six configuration decisions behind a co-occurrence matrix: corpus, span, lemmatisation, stop words, association measure and minimum frequency threshold.
There is no universally right setting, only a setting consistent with the question asked.

Collocation, corpus and the linguistic root

The idea is old and it comes from linguistics, not from search. The distributional principle, that a word is characterised by the company it keeps, is usually traced to J. R. Firth in 1957, and it is the direct ancestor of every embedding model running in production today. French lexicometry built the applied machinery around it in the early 1980s, in the corpus work published in the journal Mots: count the forms in a closed textual corpus, compare observed frequencies against a theoretical distribution, and extract the pairs whose joint presence is statistically improbable. The vocabulary of that school (corpus, énoncé, unités lexicales, spécificités) is still the cleanest way to talk about the notion, and it predates SEO by three decades.

Collocation is not a synonym for co-occurrence, and treating them as interchangeable is a tell. Co-occurrence is the observation: two units appear together above chance in a corpus. Collocation is the subset that has crystallised in the language, the conventional pairing a native speaker produces without thinking, « heavy rain » rather than « strong rain ». Every collocation is a co-occurrence. Very few co-occurrences are collocations. In a netlinking context you are essentially never building collocations, you are shifting a distribution.

The spelling question deserves one sentence because it generates real search volume. English usage keeps the hyphen, co-occurrence. French lexicometry, and the post-1990 spelling norm, write cooccurrence as one word. Both label the same notion, search engines resolve them to a single concept, and building two pages to chase the two forms is exactly the thin duplication a content gap review flags first.

One more disambiguation. The « co-occurrence rule » that shows up in search results usually belongs to a different field: in statistics and epidemiology it describes conditions appearing together in the same subject more often than independent probability allows, which is the comorbidity literature. Same arithmetic, different unit of observation. Nothing transfers to search beyond the intuition that joint presence needs a null model to mean anything.

Why it became a citation signal in 2026

The term resurfaced because generative search changed what a mention is worth. Ahrefs' analysis of 75,000 brands found that branded web mentions showed a 0.664 Spearman correlation with brand visibility in AI Overviews, against 0.218 for backlinks, 0.295 for referring domains, 0.326 for Domain Rating and 0.527 for branded anchor text. The same study reported that roughly 26 % of brands had zero AI Overview mentions, and that brands in the top web-mention quartile averaged 169 AI Overview mentions against 14 in the next quartile. Ahrefs states explicitly that correlation does not establish causation, and that caveat is not decoration: mention volume and brand size are heavily confounded.

What the study does say, in language that maps directly onto this entry, is that models derive their understanding of a brand from words on the page, term prevalence, co-occurring terms and topics, and context. That is a description of a distributional representation, not of a link graph.

The surrounding numbers explain the commercial urgency. Semrush's refreshed AI Overviews study, published 15 December 2025 across more than 10 million keywords, tracked AI Overview presence rising from 6.49 % of queries in January 2025 to a peak of 24.61 % in July before settling at 15.69 % in November 2025, with commercial queries climbing from 8.15 % to 18.57 % and transactional from 1.98 % to 13.94 %. Seer Interactive's study of 3,119 informational queries across 42 organisations, relayed by Search Engine Land, measured organic CTR falling from 1.76 % to 0.61 % when an AI Overview was present. When the click is taken, being named inside the answer becomes an outcome in its own right, which is where unlinked brand mentions stopped being a curiosity and became a line in the plan. That is the whole rationale behind treating getting a brand named inside the right editorial context as deliverable work rather than a by-product.

Google's own direction is worth reading precisely. Its 15 May 2026 Search Central resource on appearing in generative AI features insists that conventional SEO practice remains the foundation and spends most of its length dispelling AEO and GEO misconceptions. On 3 June 2026 it began rolling out Search Console performance reports for generative-AI surfaces, exposing impressions, URLs, countries and devices. No co-occurrence score, no unlinked-mention report. What exists now is a measurement path, not a confirmed signal, and the honest position is to say so.

Four numbered steps: choose the window, fill the square matrix of counts, compare observed against expected, then cross log-likelihood with mutual information.
A count becomes a specificity only after passing through an association measure.

Where it fits in a netlinking operation

The practical consequence is that host selection and paragraph writing carry weight that the anchor alone never did. A placement on a media whose existing vocabulary already sits inside the category delivers a mention embedded in a document where the category terms are dense by construction. The same sentence dropped onto an unrelated site delivers a link and almost no distributional gain. This is one of the few arguments for owned editorial media that survives scrutiny: on the French titles we operate in-house, the surrounding text is written by people who cover the topic continuously, so the co-occurrence comes from the publication's normal output rather than from a paragraph engineered around a client name.

Machines do not read that paragraph, they count it. This short explainer walks through bag-of-words, TF-IDF and co-occurrence matrices, which is the actual representation a placement ends up living inside:

The industry has already moved, ahead of the evidence. Editorial.Link's survey of 518 SEO professionals, published 25 March 2026, found 48.6 % rating digital PR the most effective link-building tactic against 16 % for guest posting, and 80.9 % believing that unlinked brand mentions affect organic rankings. That second figure is a belief, not a finding, and it is worth naming as such: a large majority of practitioners holding an opinion has never once made that opinion true. The defensible reading is narrower. Mentions correlate with generative visibility, they cost roughly what editorial coverage costs, and they arrive as a by-product of work you are already paying for, which makes the case for spreading a campaign across several months instead of a single burst rather than for a separate mentions budget.

Concretely: choose hosts by the topical density of their archive, not by their metric card; write the paragraph so the brand sits next to the category and the use case in normal prose; accept branded and naked anchors more often, because the paragraph is doing work the anchor no longer needs to do alone. Publishers whose archives you can inspect before buying make this checkable, which is the point of browsing a media catalogue with its editorial history exposed rather than a spreadsheet of authority scores.

Measuring it: PMI, log-likelihood and matrices

The canonical structure is the co-occurrence matrix: vocabulary on both axes, each cell holding the joint count of two terms inside the chosen window. Every distributional representation, from the count-based models of the 1990s to the embeddings underpinning current retrieval, starts from that object or from a factorisation of it. Building one by hand once is the fastest way to stop hand-waving about the concept.

This walkthrough builds a small matrix step by step, which makes the window parameter concrete in a way prose does not:

Raw counts are unusable on their own, so you normalise. Pointwise mutual information, imported into corpus linguistics by the 1990 Computational Linguistics work on word association norms, divides observed joint probability by the product of the marginals. It is intuitive and it has a well-known defect: it inflates rare pairs violently, so a brand mentioned twice next to a niche term outranks a brand mentioned four hundred times next to the category. The log-likelihood ratio, standard for sparse text data since the early 1990s, behaves far better on the frequency profile you actually get from web corpora. If you build one score for a brand audit, build that one.

A workable protocol: assemble a corpus from the top twenty results for your ten money queries plus the trade press covering the category, strip navigation, set the window at the paragraph, count brand-to-term pairs, rank by log-likelihood, then repeat quarterly and read the delta rather than the absolute value. The absolute score is meaningless across corpora. The delta, on a stable corpus definition, is the only thing that tells you whether coverage moved. Network visualisation tools such as VOSviewer render the result readably when you need to show a client which terms sit adjacent to their name and which sit adjacent to a competitor's. On the outcome side, the Search Console generative-AI reports released on 3 June 2026 give the first official series to correlate against, so the loop closes: coverage in, AI impressions out.

What we see go wrong

The most common failure is measuring on your own site. Counting how often your brand sits beside your category on your own pages tells you what your copywriter did last week. Co-occurrence only carries information when the corpus is third-party, because the whole value of the signal is that you did not write it.

Second, the single-placement fallacy. Distributions move on repetition across distinct domains, and one article, however well written, contributes a rounding error to any corpus a model was trained or grounded on. If the plan does not include recurrence across several independent publishers over months, the co-occurrence rationale is decoration on a link buy.

Third, mechanical stuffing. Paragraphs engineered to jam the brand next to five category terms read as promotional to a human editor and offer no advantage over natural prose to a counting machine, since the window already covers the paragraph. Editorial.Link's respondents flagged low topical relevance as a placement red flag at 67.2 % and low-quality content at 86.3 %: the same texts that trip an editor's filter are the ones sitting on hosts that stop getting cited.

Fourth, treating the correlational evidence as settled. The 0.664 figure from the Ahrefs study is a correlation with AI Overviews visibility, on a sample filtered to domains above DR 40 with a top keyword above 800 monthly searches. It is not a ranking finding, it does not cover conventional organic results, and large brands generate mentions because they are large. Quote it accurately or leave it out.

Fifth, and this one is quiet: auditing a corpus that is mostly your own press releases and syndicated copies of them. Deduplicate near-identical documents before counting, otherwise a single successful press push shows up as forty independent observations and the report says the distribution moved when nothing did.

Put it into practice?

Nautilinks operates an owned network of editorial media. In-house written articles, transparency disclosures respected, anchor mix calibrated.

See pricing → Buy backlinks service
BD
Benoit Demonchaux Founder · Nautilinks

Founder and operator of Nautilinks. Edits and writes the site's editorial glossary, as well as the content published across the Nautilinks network of editorial media.

Frequently asked questions

Does an unlinked mention pass ranking equity the way a link does?

There is no evidence that it does, and Google has never said it does. Editorial.Link's March 2026 survey found 80.9 % of 518 SEO professionals believe unlinked mentions affect organic rankings, which measures industry belief and nothing else. What the data supports is narrower: mentions correlate strongly with visibility in AI Overviews, per Ahrefs' 75,000-brand study. Treat mentions as a generative-visibility play and links as a link-graph play, and do not let one budget justify itself with the other's evidence.

Co-occurrence or cooccurrence: does the hyphen matter for targeting?

Not enough to build around. English usage keeps the hyphen, French lexicometry and the post-1990 norm write it as one word. Search engines resolve both to the same concept and serve largely the same results, so a single page covering both spellings in its body is the correct move. Creating two pages to capture the two forms produces thin near-duplicates that compete with each other and get consolidated anyway, usually at the expense of the weaker one.

What window size should I use for a brand-level co-occurrence audit?

The paragraph, in almost every case. Sentence windows are too tight: brand and category rarely share a clause in natural editorial prose. Document windows are too loose: on a 2,000-word article, everything co-occurs with everything and the score collapses into topical relevance. A paragraph, or a fixed window of roughly 30 to 50 tokens, matches how a mention is actually embedded in coverage. State the choice in the report, because a score without its window is not comparable to anything.

Which association measure should I trust, PMI or log-likelihood?

Log-likelihood, for web corpora. PMI is intuitive but it overrates rare pairs badly, so a brand appearing twice beside an obscure term will outrank a brand appearing hundreds of times beside its core category. That is the opposite of what you want to read in an audit. The log-likelihood ratio handles the sparse, skewed frequency profile of scraped text far better. Raw joint counts are unusable on their own: without a null model, frequent function words dominate every ranking.

Can I actually measure any of this in Google's own tools?

Partially, since 3 June 2026. Search Console now reports on generative-AI surfaces, exposing impressions, URLs, countries, devices and date breakdowns for AI Overviews and AI Mode, with rollout initially limited to a subset of sites. There is no co-occurrence report and no unlinked-mention report. The workable setup is external: build your own third-party corpus, score it quarterly, and correlate the delta against the AI-surface impressions Search Console now gives you.

Quiz

Test your knowledge

Quiz: Co-occurrence

1/3

What separates a collocation from a plain co-occurrence?

Newsletter

GEO + SEO analyses and network case studies, in your inbox

Once or twice a month at most. No filler. One-click unsubscribe.

By subscribing you agree to receive our emails. See our privacy policy.