CC Backlink Checker

Free domain authority & referring domains data from Common Crawl

121M domains · 3.9B domain-to-domain links · live

Check any domain's referring domains & authority for free

A free referring domains & domain authority index built on Common Crawl's open web graph — 121 million domains connected by 3.9 billion domain-to-domain links. Harmonic Centrality ≈ trust/authority, PageRank ≈ link popularity — no signup, no limits.

Enter a root domain (no https:// or www needed). Results are ranked by each referring domain's own authority.
Finds high-authority domains that link to your competitors but not to you — the best outreach targets to raise your own Harmonic Rank / PageRank.

How to raise your CC PageRank and Harmonic Centrality

Simple, practical steps based on how the Common Crawl web graph is actually built.
🔗

Get links from strong domains

One link from a domain with a high Harmonic Rank or PageRank Rank can help more than 100 links from small sites. Open the Referring Domains tab for your own domain and look at the Authority Score of your current referring domains. Then try to get more links from similar or stronger sites.

🌐

Spread links across many domains

This tool counts referring domains, not raw link count. Ten links from ten different domains help more than fifty links from one domain. Aim for variety, not just volume.

🎯

Use "Outbound Links" as an opportunity list

In the Referring Domains and Link Gap Finder tables, check the Outbound Links column. A domain with a very high outbound link count is already linking out to many other sites. That is often a sign of a directory, resource page, or partner list. These sites are usually more open to adding one more link, including yours.

🔄

Turn competitor links into your own targets

Run Link Gap Finder with your competitors. The domains it finds already link out to sites in your niche, and many of them have a high Outbound Links count too. That makes them good outreach targets: they clearly link out often, and they are already linked to sites like yours.

🚀

Strong links can raise your crawl priority too

Common Crawl does not crawl the whole web. It picks pages based on prior crawls, sitemaps, and outside link signals. A link from a domain with strong Harmonic Rank or PageRank Rank works like a vote. It can help your pages get found and picked up in a future crawl, not just help your score.

⚠️

Important: Common Crawl does not run JavaScript

The crawler only reads static HTML. It does not execute JavaScript. If your links are added by React, Vue, or other client side code, Common Crawl will never see them, and they will not count here either. Make sure your important links exist directly in the raw server rendered HTML.

🚫

Do not block crawlers with robots.txt

Common Crawl fully respects robots.txt. If a page you want counted is disallowed there, it gets skipped completely, even if real users and other crawlers can see it fine.

🛠️

Keep your own site easy to crawl

Fix broken links and long redirect chains. Keep your internal linking simple, so a crawler can reach your important pages in as few hops as possible from your homepage.

Quick checklist

  • Get at least one link from a domain with a high Authority Score, not only many links from small sites.
  • Spread your outreach across many different domains, not just one.
  • Find domains in your niche with a high Outbound Links count and reach out for a listing or link.
  • Check your Link Gap Finder results for domains linking to competitors but not to you.
  • Make sure your links appear in the raw HTML, not only after JavaScript runs.
  • Check robots.txt does not block pages you want crawled and counted.
  • Fix broken redirects and keep internal links working.

Frequently Asked Questions

What this data actually is, how the scores are calculated, and why the numbers won't match paid tools 1:1.

What is CC Backlink Checker, and what data does it use?
This is a free referring domains & domain authority index built entirely on the Common Crawl Host/Domain-level Web Graph — a public dataset released by the non-profit Common Crawl Foundation. We combined three monthly crawls (April, May and June 2026) into a single domain-to-domain link graph with 121 million domains (the nodes) connected by 3.9 billion domain-to-domain links (the edges — i.e. "domain A links to domain B" at least once), then computed Harmonic Centrality and PageRank across the whole graph. Nothing is scraped live — every number comes from that public dataset, queried on demand from our backend.
What do "Harmonic Rank" and "Authority Score" mean?
Harmonic Centrality is a standard graph-theory metric (used widely in academic web-graph research) that measures how well-connected a domain is, based on its average distance to every other domain in the graph. It plays a similar role to Majestic's "Trust Flow." Authority Score is our own 0–100 scale derived from that rank on a log curve (similar in spirit to Ahrefs' DR or Moz's DA) — much easier to read at a glance than a raw rank out of 121 million.
What do "PageRank Rank" and "Citation Score" mean?
This is the original Google PageRank algorithm computed over the same domain graph: a link is a "vote," and votes from higher-PageRank domains count for more. It plays a similar role to Majestic's "Citation Flow" — raw link popularity, independent of topical trust. Citation Score is the 0–100 scale version of that raw rank.
Is "Referring Domains" the same thing as "backlinks"?
No. Referring Domains counts distinct domains that link to the target at least once — not the number of links or URLs. If one domain links to you from 500 different pages, it's still counted once. This is domain-graph granularity (like Majestic's "Referring Domains" metric), not URL-level or anchor-text-level — Common Crawl's public web graph is only released at the domain and host level.
Why don't these numbers match Ahrefs, Semrush, or Majestic?
Those tools run dedicated backlink crawlers that continuously fetch billions of URLs specifically to discover links, and they merge years of historical crawl data together. Common Crawl is a general-purpose, once-a-month research crawl, not a specialized backlink index — its crawl targets, budget, and depth per domain are all different, so its link graph is a large but incomplete sample of the real web. Treat this tool as a free, open-data approximation, not a 1:1 replacement for paid tools.
Why does a domain like G2.com show fewer referring domains / outbound links than I'd expect? Did Common Crawl just not crawl it?
Usually, yes — incomplete crawl coverage is the main reason. A few things stack up:
  • Domain-level de-duplication: even if a site links to G2 from 10,000 different pages, it's counted once, as a single referring domain — so counts here always look far lower than page-level backlink tools.
  • Snapshot coverage gap: we only use 3 monthly crawls (Apr–Jun 2026), not Common Crawl's full multi-year archive. If a page linking to G2 was crawled in a different month — or never selected for any crawl — that link isn't in this graph, even though it exists on the live web.
  • JavaScript-rendered links aren't captured: Common Crawl's main crawler fetches static HTML and does not execute JavaScript. Links injected client-side (review badges, embedded comparison widgets, SPA navigation) are invisible to the crawl even though a browser renders them fine.
  • robots.txt compliance: Common Crawl honors robots.txt. If a linking page or section of a site is disallowed, it's skipped entirely.
  • Crawl seed/budget bias: Common Crawl prioritizes a mix of prior crawls, sitemaps, and external seed lists; long-tail or newly published pages may not be selected for a given month's crawl at all.
A low number here doesn't mean "this domain has few real backlinks" — it means "this many were visible in the specific Common Crawl snapshots we indexed." Treat it as a large free sample, not a hard ceiling.
How fresh is the data, and how far back does it go?
The graph is a union of the CC-MAIN-2026 April, May, and June crawls — one recent snapshot period, not a rolling history. It isn't updated in real time; moving to a newer crawl requires us to rebuild and re-upload the whole graph.
What is the Link Gap Finder, and what counts as "niche-relevant"?
It compares your domain against one or more competitors and finds every domain that links to at least one competitor but not to you — high-value outreach targets. We flag a subset as niche-relevant & attainable: domains linking to 2+ of your competitors, with a Harmonic Rank between 300 and 300,000 (clearly authoritative, but not an unreachable top-300 mega-site like Google or Wikipedia).
Is this free? Are there limits? Is there an API?
Yes, completely free, no sign-up. It's a thin Cloudflare Worker in front of a Hugging Face Space backend, so please be reasonable with request volume. The JSON API behind this page is public: /api/report?domain=example.com and /api/gap?target=you.com&competitors=comp1.com,comp2.com.
Who built this, and why is peec.ai pre-filled in the Link Gap Finder?
This tool was personally built and published by Metehan Yeşilyurt as part of his AI Visibility Research work at peec.ai. The Link Gap Finder's competitor field is pre-filled with peec.ai as a working example so you can try the feature immediately — feel free to clear it and enter your own competitors instead.
Why did I get a "Request failed" / timeout error for a domain like stripe.com?
A handful of domains (typically massive platforms like Google, Cloudflare, or Stripe) have hundreds of thousands of referring domains. Returning that entire result set in one request can exceed our backend's processing time and hit an infrastructure timeout, which currently surfaces as a generic request-failed error instead of a friendly message. We're aware of this and are actively working on a performance improvement (pagination/streaming for very large result sets) so massive domains load reliably too. In the meantime, most domains — including the vast majority of real-world SEO/AEO use cases — work instantly.