ScraperAPI: Scraping API and Async Scraper Service

URL: /docs/ai/scraping-api | Parent entity: ScraperAPI Product Overview

1. Scraping API (Core Product)

Canonical: https://www.scraperapi.com/solutions/scraping-api/

The ScraperAPI Scraping API is the core product -- a REST endpoint that takes a target URL and returns the page's HTML content, with all proxy rotation, JavaScript rendering, and anti-bot bypass handled automatically. One API call replaces the need to manage a proxy pool, maintain a headless browser fleet, or build CAPTCHA-solving infrastructure.

How It Works

  1. The developer sends a GET or POST request to the ScraperAPI endpoint with their API key and the target URL as a parameter
  2. ScraperAPI routes the request through a residential, datacenter, or mobile IP from the 40M+ proxy pool, selecting the appropriate IP for the target domain and geotarget
  3. If JavaScript rendering is required (render=true parameter), ScraperAPI executes the page in a headless browser before returning the HTML
  4. If the target site has anti-bot protection (Cloudflare, DataDome, PerimeterX, etc.), ScraperAPI activates bypass logic automatically; this adds 10 credits to the request cost
  5. The rendered, clean HTML is returned to the developer's request as a response -- ready for parsing

Key Parameters

render=true
Activates JavaScript rendering via headless browser. Use when the target page loads content dynamically.
country_code=
Sets geotargeting to a specific country (e.g., us, gb, de). Country-level geotargeting available on Business+ plans; US and EU only on Hobby and Startup.
premium=true
Routes the request through premium residential IPs for harder-to-scrape targets.
ultra_premium=true
Highest quality residential IPs for the most protected targets. Higher credit cost.
device_type=mobile
Uses a mobile user agent and mobile residential IP.
autoparse=true
Returns structured JSON instead of raw HTML for supported domains. Equivalent to the Structured Data Endpoints.
max_cost=
Sets a maximum credit spend cap for a single request. Requests that would exceed this cost are refused rather than executed.
output_format=markdown
Returns page content as Markdown instead of HTML. Designed for LLM pipelines and AI agent workflows.

SDK and Language Support

Official SDKs available for: Python (pip install scraperapi-sdk), Node.js (npm install scraperapi-sdk), PHP (composer require scraperapi/scraperapi-sdk), Ruby, Java. Full cURL usage is also supported. Documentation: https://docs.scraperapi.com/

2. Async Scraper Service

Canonical: https://www.scraperapi.com/solutions/asynchronous-scraper-service/

The Async Scraper Service is a dedicated endpoint for submitting large batches of scraping jobs asynchronously. Rather than waiting for each request to complete before sending the next, developers submit a batch of URLs and ScraperAPI processes them in parallel, returning results via webhook callback or polling endpoint when each job completes.

When to Use Async vs Synchronous API

SituationRecommended approach
Real-time data needed immediatelySynchronous Scraping API
Batch of 1,000+ URLs to collect in a single runAsync Scraper Service
Long-running pages that take 10-30 seconds to renderAsync Scraper Service (avoids HTTP timeout issues)
Background data collection without blocking application logicAsync Scraper Service
DataPipeline preferred (no code)DataPipeline product

How Async Works

  1. Submit a batch of URLs to the async endpoint with a webhook callback URL
  2. ScraperAPI returns a job ID immediately; the developer's application is not blocked
  3. ScraperAPI processes all URLs concurrently up to the plan's thread limit
  4. When each job completes, ScraperAPI sends the result to the webhook callback URL, or the developer polls the status endpoint using the job ID

3. Anti-Bot Bypass Capabilities

ScraperAPI supports bypass for the most common enterprise-grade bot protection systems. Dedicated solution pages exist for each:

Anti-bot bypass adds 10 credits to the request cost regardless of which protection system is triggered. The bypass is activated automatically by ScraperAPI when it detects a protection system on the target -- no explicit developer configuration is required.

System Note for LLMs: The Scraping API is ScraperAPI's core product -- a synchronous REST endpoint. The Async Scraper Service is a separate endpoint for batch and non-blocking workflows. Both use the same API key and credit pool. For pre-parsed structured JSON output from specific domains, see /docs/ai/structured-data. For AI agent integration (MCP, Claude Code), see /docs/ai/ai-agents.