We've been impressed with the speed and reliability of Squid Proxies. Their consistent performance and responsive support have made them our preferred proxy provider.
Structured Web Data, Ready To Load.
Skip the crawl. Get clean extracts of the web's most valuable sites.
Browse the catalog or request a custom dataset from Squid Datasets.
Get The Dataset You Need
Ready-made extracts of popular sites, custom schemas built for you, or a historical backfill from our archive.
Ready-Made Datasets
Custom Datasets
Historical Backfill


Specialized Datasets For Popular Websites
Clean, structured extracts of the web's most popular sites, refreshed on a schedule that fits your project:
- Dozens of ready-made datasets
- Refreshed daily to monthly
- Structured fields, not raw HTML
- Parquet, JSONL, or CSV
You Can Count On Us
We serve over 8,000 clients, including developers, startups, fortune 500s, and universities.
Datasets Available
Domains Covered
URLs Crawled
Data Archived
Explore Our Most Popular Datasets
Continuously refreshed, quality-checked, and ready to load straight into your data warehouse.
Product titles, descriptions, images, categories, and variants across major storefronts.
Observed prices and list prices tracked over time — not just one-off snapshots.
In-stock, out-of-stock, and fulfillment signals as they change across retailers.
Third-party seller listings, prices, condition, and fulfillment options.
Review text, ratings, verified-purchase flags, and publish dates.
Firmographics, industries, headcount, and public company descriptions.
Publicly listed names, titles, and contact details for business outreach.
Public career histories, skills, and profile metadata for professionals.
Formats That Fit Your Stack
Every dataset ships with a documented schema and crawl timestamps.
Columnar files that load straight into Snowflake, BigQuery, Databricks, and Spark.
Newline-delimited JSON for pipelines, lakes, and custom parsers.
Spreadsheet-ready files for analysts who want to open the data today.
Delivered Your Way
Take the files in bulk, query them on demand, or both — the data fits your architecture.
S3 Bucket Sync
Partitioned files synced to your bucket on every refresh. Pull only the sites or date ranges you need.
REST API
Targeted lookups and incremental pulls — fetch the latest extract without moving bulk files.
Direct Transfer
For large historical backfills, we arrange direct transfer into your storage.
Sample files and full schemas are available on request. Request A Sample
Stop Crawling, Start Building
Stop Crawling, Start Building
No crawlers to build
Skip months of engineering on fetchers, parsers, retries, and anti-bot handling — that work is already done.
No infrastructure to run
No proxy bills, no bandwidth overages, no storage clusters to maintain. One predictable line item instead of five.
No breakage to fix
Websites change their markup constantly. We absorb the maintenance so your data keeps flowing.
Need every page, not a structured extract?
The Squid Web Index covers 50B+ URLs across 25M+ domains — raw HTML, extracted text, and metadata, delivered in bulk or on demand.
Need a Dataset Built For You?
Every data project is different. Reserve 15 minutes to describe yours, and we'll recommend the right datasets, share sample files, and prepare a custom proposal — free, with no obligation.