
QuanticData — Web Scraping MCP ServerVerifiedFeatured
@quantumproxies
About QuanticData — Web Scraping MCP Server
Web scraping MCP server for AI agents: scrape any page to clean Markdown, SERP search (Google, Bing, DuckDuckGo), site crawl and map, 74 ready-made collectors and residential proxies. 25 tools, pay per success.

Config
Add this server to your MCP-compatible client using the configuration below.
{
"mcpServers": {
"quanticdata": {
"command": "npx",
"args": [
"-y",
"quanticdata-mcp"
],
"env": {
"QUANTICDATA_API_KEY": "qd_live_your_key_here"
}
}
}
}Tools
25Scrape a single web page through a residential proxy and return it as clean Markdown (or HTML/text). Uses a real Chrome TLS fingerprint by default and only spins up a headless browser if the page is bot-challenged. Optionally run structured extraction (CSS selectors) or AI extraction (natural-language prompt). Markdown keeps the complete page by default (content_mode 'smart': everything except nav/footer/cookie chrome, with GFM tables and absolutized links); to inspect a page's raw no-JS/SEO fallback use format 'html'.
Look at a page ONCE with an LLM and get back CSS selectors that extract the fields you asked for. Pass the returned `parser` as the `extract` argument on every later scrape of that same layout and no AI runs again — it becomes a plain, free, deterministic extraction. Use this instead of ai_prompt whenever you will scrape more than a couple of pages of the same shape. Every selector is run against the page before being returned, so `report`/`coverage` tell you which fields are actually reliable.
Store a generated parser under a name so it can be reused by id. Scrape later with scrape's `preset_id` instead of repeating the selectors, and every run is scored per field — when the recent success rate decays (the site redesigned), the preset regenerates itself from `source_url` and bumps a version. Give it a source_url whenever you can: without one it can never self-heal.
List your stored parser presets with their version, health stats and changelog.
How well a stored parser is still working: success rate per field, mean coverage over the recent runs, and whether it now counts as decayed (i.e. the site probably changed).
Regenerate a preset's selectors now (the manual trigger for the automatic repair). Refetches the source page and adopts new selectors ONLY if they extract more than the current ones — a heal that finds nothing better leaves the preset untouched and is not billed.
Audit a URL's SEO in one call: fetches it twice — as a pure HTTP bot (no JS) and fully rendered — and returns both views (title, description, canonical, h1, word count) plus the diff (JS-only content, changed title/description, canonical missing without JS) and bot-facing meta (robots, Open Graph, JSON-LD types). Use this instead of scraping manually when checking how a page indexes.
Run structured Google, Bing or DuckDuckGo searches through a residential proxy. Bing supports web, shopping, images, news, videos, places/maps and autocomplete over HTTP, including Copilot AI answers and citations when Bing returns them. Google web search also parses rich blocks directly from its HTTP response.
Search the live web, fetch the top organic pages as clean Markdown, and return citation-ready numbered sources plus one token-bounded `context` string ready for an AI prompt. Use this when the goal is answering/researching, and use `search` when raw SERP structure or a specialized vertical is needed.
Discover a site's URLs fast (robots.txt sitemaps + /sitemap.xml + homepage links) without a full crawl. Returns up to `limit` URLs (default 100) plus the site-wide `total` and a per-section `summary` (e.g. '/blog': 1988) so you see the site's shape without the full list. Narrow with `search` (substring filter — the primary way to find specific pages) or set group_by 'path' for the path tree with counts instead of URLs.
Start an asynchronous BFS crawl of a site from a seed URL, converting each page to Markdown. Returns a job id — poll with crawl_status.
Poll a crawl job for progress and the pages crawled so far. Polls are incremental: pass the previous response's `nextCursor` as `since` to receive only the pages crawled since your last poll. Pages omit their content by default — set include_content true only when you actually need the text (a large crawl's full content can be hundreds of KB).
Scrape many URLs asynchronously with shared options. Returns a job id — poll with batch_status. For SEO/status audits over many pages set mode 'summary': items carry metadata only (title, description, canonical, contentLength) instead of full page content.
Poll a batch job for progress and per-URL results. Polls are incremental: pass the previous response's `nextCursor` as `since` to receive only the items completed after your last poll. Items omit page content by default — set include_content true only when you actually need the text.
Paginate ONE search query asynchronously and merge deduplicated organic results. Page-one AI Overview/PAA/Knowledge Graph/answer enrichments are retained; set render:true to request those Google JS blocks. Billed per page actually fetched, with unavailable pages refunded.
Poll a bulk search job for progress and merged organic results. Polls are incremental: pass the previous response's `nextCursor` as `since` to receive only the organic results gathered after your last poll.
Build a structured dataset from a plain-language prompt. Quantic AI plans the search queries, searches Google/Bing/DuckDuckGo, maps the sites it finds and scrapes them into validated rows (CSV/JSON). Returns a job id — poll with dataset_status. Billed per delivered, validated record (email/phone fields cost extra, only when found); the run never exceeds limits.max_cost_usd, and the unspent budget is refunded.
Poll a dataset job for progress, the collection trace (steps) and the rows so far. Polls are incremental: pass the previous response's `nextCursor` as `since` to receive only rows delivered after your last poll. Set mode 'summary' to omit rows and get only progress + steps (light poll). When status is completed, the response includes signed CSV/JSON download URLs.
List the account's proxy services of every type — Residential Basic/Premium/Private, Mobile, Mobile V2, Datacenter (static or traffic-based), ISP, IPv6 — with plan type, bandwidth left, expiry, whitelisted IPs and the orderId to pass to generate_proxies. Call this first to see which proxy plans are available.
Generate ready-to-use proxy endpoint strings (credentials included) from one of the account's active proxy services — any type: residential, mobile, datacenter, ISP, IPv6. Supports geo targeting (country/state/city, ISP or ASN where the plan allows it), rotating or sticky sessions, HTTP or SOCKS5, and several output formats. Use list_proxies first to get the orderId, and proxy_locations for valid targeting codes. The returned strings plug straight into any HTTP client, e.g. curl -x.
Discover valid geo-targeting values for a proxy plan type before calling generate_proxies: countries, states, cities, ASNs, or the full location tree (countries → regions → cities → ISPs). Use level 'tree' for Residential Premium / Mobile V2 slugs and ISP codes, or for the static datacenter gateway list; note the tree can be large.
Manage IP-auth whitelisting on a proxy service (Residential Basic, Datacenter, ISP, IPv6, Mobile): add or remove an IP, or list the current entries. A whitelisted machine uses the proxies without username/password — required for the Mobile V2 IP-auth proxy list. Residential Premium/Private use user:pass auth and don't need this.
List the ready-made Collectors: paid, versioned scrapers you run with a semantic input (keyword + location, place id, product id, domain…) instead of URLs — e.g. web_search, search_images, search_videos, keyword_ideas, amazon_search, amazon_product, ebay_search, aliexpress_search, linkedin_jobs, indeed_jobs, reddit_posts, youtube_search, youtube_channel, instagram_profile, tiktok_profile, tiktok_video, linkedin_profile, linkedin_company, zillow_search, zillow_property, app_store_apps, app_store_reviews, google_play_apps, google_maps_places, place_reviews, google_jobs, google_news, google_shopping, product_offers, hotels, google_flights, google_events, google_trends, google_autocomplete, google_lens, youtube_video, ebay_product, flipkart_search, idealista_search, kleinanzeigen_search, autotrader_search, github_repos, hacker_news, coingecko_coins, wikipedia_articles, yahoo_finance, stackoverflow, steam, npm_packages, sec_filings, defillama, wayback_machine, clinical_trials, certificate_transparency, wikidata, nvd_cve, openfda, openalex, pypi_packages, exchange_rates, gleif_lei, docker_hub, crates_io, world_bank, openlibrary_books, arxiv_papers, weather_forecast, whois_domain, dns_records, itunes_search, local_business_leads, site_contacts, company_profile, business_directory. Returns each collector's slug, input/output schema, example input, price per delivered result and current health. Billing is pay-per-success: only delivered rows are charged.
Run a Collector by slug with a semantic input (see list_collectors for each collector's inputSchema and example). Short runs return the rows inline; long runs return 202 with a run_id + statusUrl — poll with collector_run_status. Results are billed per delivered row (never for failures). Set `async` true to force background execution.
Fetch a Collector run by run_id: status (queued|running|done|failed), result count, cost, partial flag and the result rows. Use after run_collector returned 202/async. Pass format 'csv' to get the rows as CSV text.
Overview
What is the QuanticData MCP server?
An MCP server that gives Claude, Cursor and any MCP client live web access through QuanticData's residential proxy network. 25 tools: page scraping to clean Markdown (with CSS and AI extraction), structured SERP search across Google, Bing and DuckDuckGo in 17 verticals, site crawl and map, SEO audits, 74 ready-made collectors (Amazon, Google Maps, LinkedIn jobs, app stores…), datasets from a plain-language prompt, and proxy generation of every type. Billing is pay-per-success: blocked pages, captchas and failed calls cost nothing. Every account includes $2 of free usage monthly, no card required.
How to use it
Claude Code — one command:
claude mcp add quanticdata -e QUANTICDATA_API_KEY=qd_live_your_key_here -- npx -y quanticdata-mcp
Claude Desktop and Cursor: paste the config block from the Config tab into claude_desktop_config.json or .cursor/mcp.json, restart the client, and the 25 tools appear. Get a free API key at quanticdata.io; full parameter reference in the docs.
Key features
- Scrape any URL to Markdown, HTML or text — real-browser TLS fingerprint, automatic browser escalation when a page fights back, CSS and AI extraction to structured JSON
- Search: SerpApi-compatible results from Google, Bing and DuckDuckGo — web, news, shopping, maps, jobs, flights, hotels and more
- Crawl and map whole sites to Markdown, async with incremental polling
- 74 ready-made collectors driven by semantic inputs (keyword + location, product id, domain) instead of URLs, billed per delivered row
- Self-healing parsers: learn CSS selectors once with an LLM, reuse them free forever; presets regenerate when a site redesigns
- Proxy tools: list plans and generate residential, mobile, datacenter, ISP or IPv6 endpoints with geo targeting down to city, ISP or ASN
Example prompts
- "Scrape the pricing page at example.com and give me the plans and prices as a table."
- "Search Google Shopping for 'nintendo switch oled' in the US and list the 5 cheapest offers."
- "Map docs.example.com, then crawl only the /guides/ pages and summarize each one."
- "Run the google_maps_places collector for dentists in Austin, TX with email and phone."
- "Generate 5 sticky US residential proxies as socks5 URLs and whitelist my server IP."
FAQ
Is the MCP server free?
The package is free and open (MIT). You pay only the underlying API calls — from $0.0002 per scraped page — and every account includes $2 of free monthly usage with no card.
Which clients are supported?
Any stdio MCP client: Claude Desktop, Claude Code, Cursor, Windsurf, VS Code, Zed, Cline, and agent frameworks like LangChain or CrewAI via their MCP adapters.
Does it run locally?
Yes — a small package launched with npx that talks to the hosted API with your key. No internal secrets, nothing to deploy: proxies, browsers, retries and parsing run on QuanticData's infrastructure.
Frequently asked questions
Is the MCP server free?
The package is free and open (MIT). You pay only the underlying API calls — from $0.0002 per scraped page — and every account includes $2 of free monthly usage with no card.
Which clients are supported?
Any stdio MCP client: Claude Desktop, Claude Code, Cursor, Windsurf, VS Code, Zed, Cline, and agent frameworks like LangChain or CrewAI via their MCP adapters.
Does it run locally?
Yes — a small package launched with npx that talks to the hosted API with your key. No internal secrets, nothing to deploy: proxies, browsers, retries and parsing run on QuanticData's infrastructure.
Basic information
More Search MCP servers
Search Console Mcp
saurabhsharma2uSearch & analytics data as infrastructure — MCP server for Google Search Console, Bing Webmaster Tools, and GA4, designed for AI agents and automation.
Google Search Tool
web-agent-masterA Playwright-based Node.js tool that bypasses search engine anti-scraping mechanisms to execute Google searches. Local alternative to SERP APIs with MCP server integration.
Perplexity MCP Server
wysh3MCP web search using perplexity without any API KEYS
Everything Search MCP Server
mamertofabian
Webz.io News Search
Webz.ioSearch global news using natural language. Webz.io News Search API returns the most relevant articles and content, with filters for source, country, language, date, sentiment, and category.
Comments