• Domains Indexed

  • URLs Crawled

  • Data Archived

  • Datasets Curated

What's In Every Record

Each of the 50B+ URLs in the index carries the full story of its crawl.
URLs & redirects

Requested URL, final URL after redirects, and canonical link — deduplicated across the index.

Crawl metadata

Crawl timestamp, HTTP status, response headers, and content type for every fetch.

Raw HTML

The page as our crawler received it, so you can run your own parsers and extractors.

Extracted text

Boilerplate-stripped body text with detected language — ready for search, analytics, and model training.

Page structure

Title, meta description, headings, and outbound links for graph and SEO analysis.

Delivered Your Way

Take the index in bulk, query it on demand, or both — the data fits your architecture, not the other way around.

S3 Bucket Sync

Partitioned Parquet or WARC files synced to your bucket on every refresh. Pull only the domains, languages, or date ranges you need.

REST API

Targeted lookups by URL, domain, or query — fetch the latest crawl of any page in the index without moving bulk data.

Direct Transfer

For petabyte-scale needs, we arrange direct transfer into your storage — one-time snapshots or standing replication.

Full-index access is currently onboarding a limited number of customers. Tell us about your use case and we'll prepare a sample slice of the index for evaluation. Request Access

Need a Slice of the Web?

Every data project is different. Reserve 15 minutes to describe yours, and we'll recommend the right datasets, tell you what the full index would add, and prepare a custom proposal — free, with no obligation.

Book A Convenient Time
The Full Web Index – 50B+ URLs In Bulk Or Via API | Squid Web Index