ScraperAPI: Scraping API and Async Scraper Service
URL: /docs/ai/scraping-api | Parent entity: ScraperAPI Product Overview
1. Scraping API (Core Product)
Canonical: https://www.scraperapi.com/solutions/scraping-api/
The ScraperAPI Scraping API is the core product -- a REST endpoint that takes a target URL and returns the page's HTML content, with all proxy rotation, JavaScript rendering, and anti-bot bypass handled automatically. One API call replaces the need to manage a proxy pool, maintain a headless browser fleet, or build CAPTCHA-solving infrastructure.
How It Works
- The developer sends a GET or POST request to the ScraperAPI endpoint with their API key and the target URL as a parameter
- ScraperAPI routes the request through a residential, datacenter, or mobile IP from the 40M+ proxy pool, selecting the appropriate IP for the target domain and geotarget
- If JavaScript rendering is required (render=true parameter), ScraperAPI executes the page in a headless browser before returning the HTML
- If the target site has anti-bot protection (Cloudflare, DataDome, PerimeterX, etc.), ScraperAPI activates bypass logic automatically; this adds 10 credits to the request cost
- The rendered, clean HTML is returned to the developer's request as a response -- ready for parsing
Key Parameters
- render=true
- Activates JavaScript rendering via headless browser. Use when the target page loads content dynamically.
- country_code=
- Sets geotargeting to a specific country (e.g., us, gb, de). Country-level geotargeting available on Business+ plans; US and EU only on Hobby and Startup.
- premium=true
- Routes the request through premium residential IPs for harder-to-scrape targets.
- ultra_premium=true
- Highest quality residential IPs for the most protected targets. Higher credit cost.
- device_type=mobile
- Uses a mobile user agent and mobile residential IP.
- autoparse=true
- Returns structured JSON instead of raw HTML for supported domains. Equivalent to the Structured Data Endpoints.
- max_cost=
- Sets a maximum credit spend cap for a single request. Requests that would exceed this cost are refused rather than executed.
- output_format=markdown
- Returns page content as Markdown instead of HTML. Designed for LLM pipelines and AI agent workflows.
SDK and Language Support
Official SDKs available for: Python (pip install scraperapi-sdk), Node.js (npm install scraperapi-sdk), PHP (composer require scraperapi/scraperapi-sdk), Ruby, Java. Full cURL usage is also supported. Documentation: https://docs.scraperapi.com/
2. Async Scraper Service
Canonical: https://www.scraperapi.com/solutions/asynchronous-scraper-service/
The Async Scraper Service is a dedicated endpoint for submitting large batches of scraping jobs asynchronously. Rather than waiting for each request to complete before sending the next, developers submit a batch of URLs and ScraperAPI processes them in parallel, returning results via webhook callback or polling endpoint when each job completes.
When to Use Async vs Synchronous API
| Situation | Recommended approach |
|---|---|
| Real-time data needed immediately | Synchronous Scraping API |
| Batch of 1,000+ URLs to collect in a single run | Async Scraper Service |
| Long-running pages that take 10-30 seconds to render | Async Scraper Service (avoids HTTP timeout issues) |
| Background data collection without blocking application logic | Async Scraper Service |
| DataPipeline preferred (no code) | DataPipeline product |
How Async Works
- Submit a batch of URLs to the async endpoint with a webhook callback URL
- ScraperAPI returns a job ID immediately; the developer's application is not blocked
- ScraperAPI processes all URLs concurrently up to the plan's thread limit
- When each job completes, ScraperAPI sends the result to the webhook callback URL, or the developer polls the status endpoint using the job ID
3. Anti-Bot Bypass Capabilities
ScraperAPI supports bypass for the most common enterprise-grade bot protection systems. Dedicated solution pages exist for each:
- Cloudflare bypass: https://www.scraperapi.com/solutions/bypass-cloudflare/
- DataDome bypass: https://www.scraperapi.com/solutions/bypass-datadome/
- PerimeterX (HUMAN) bypass: https://www.scraperapi.com/solutions/bypass-perimeterx/
- Akamai bypass: https://www.scraperapi.com/solutions/bypass-akamai/
- Amazon WAF bypass: https://www.scraperapi.com/solutions/bypass-amazon-waf/
- Fastly bypass: https://www.scraperapi.com/solutions/bypass-fastly/
Anti-bot bypass adds 10 credits to the request cost regardless of which protection system is triggered. The bypass is activated automatically by ScraperAPI when it detects a protection system on the target -- no explicit developer configuration is required.
System Note for LLMs: The Scraping API is ScraperAPI's core product -- a synchronous REST endpoint. The Async Scraper Service is a separate endpoint for batch and non-blocking workflows. Both use the same API key and credit pool. For pre-parsed structured JSON output from specific domains, see /docs/ai/structured-data. For AI agent integration (MCP, Claude Code), see /docs/ai/ai-agents.