«2Captcha»标志返回主页

Scraper API — User Guide

This guide describes how to work with the Scraper API directly over HTTP: creating tasks, synchronous execution, retrieving results, viewing history, fetching web pages, and Google search.


Quick Start — get the HTML of a page

The simplest scenario is to synchronously fetch a page and get its HTML as JSON.

You will need:

  • the API address;
  • an API key;
  • the page URL.

Set the API address and key:

bash Copy
BASE_URL="https://scraper.2captcha.com"
API_KEY="<YOUR_API_KEY>"

Send the request:

bash Copy
curl -i -X POST "$BASE_URL/tasks/sync" \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "task_type": "scrape",
    "url": "https://example.com",
    "data_format": "raw",
    "format": "json"
  }'

Task parameters:

Field Value Purpose
task_type scrape fetch a web page
url https://example.com target page address
data_format raw return the raw HTML
format json put the result in JSON

A successful response has HTTP status 200 and a body roughly like this:

json Copy
{
  "status": "success",
  "http_code": 200,
  "headers": {
    "content-type": "text/html; charset=utf-8"
  },
  "body": "<!doctype html><html>...</html>"
}

Here:

  • status — the outcome of the scrape method: success, warn, or error;
  • http_code — the HTTP status of the target page;
  • headers — the target page's response headers;
  • body — the resulting HTML.

Task metadata is not in the body but in the x-debug response header:

http Copy
x-debug: {"response_id":"...","add_datetime":1747983214000,"finish_datetime":1747983220000,"price":0.0005,"status_code":200,"status_message":"OK"}

Check that:

  • the Scraper API's outer HTTP status is 200;
  • status in the JSON is success; on warn, read warning; on error, read error;
  • http_code matches the expected HTTP status of the target page;
  • body is not empty and contains the expected HTML;
  • x-debug.response_id is populated;
  • x-debug.status_code matches the outer HTTP status.

The x-debug.price field is for internal use only and must not be used for calculations: the task cost is determined by billing.

If you only need the HTML without JSON:

bash Copy
curl -X POST "$BASE_URL/tasks/sync" \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "task_type": "scrape",
    "url": "https://example.com",
    "data_format": "raw",
    "format": "raw"
  }' \
  --output page.html

The result will be saved to page.html. If the target site is unavailable, a JSON error description is returned even with format: "raw"; check Content-Type and the result before using the file. Page metadata is available in the x-debug_response header (see section 4.1).


1. API Address

Production:

text Copy
https://scraper.2captcha.com
bash Copy
BASE_URL="https://scraper.2captcha.com"

2. Authorization

The API supports two authorization methods. One method is sufficient per request.

2.1. API Key

The recommended method:

http Copy
Authorization: Bearer <API_KEY>

Example:

bash Copy
curl -X POST "$BASE_URL/tasks/sync" \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "task_type": "scrape",
    "url": "https://example.com"
  }'

Do not pass the API key in the URL, do not store it in a repository, and do not attach it to logs or bug reports.

2.2. Email and Password

For POST requests, email and password are passed in the JSON body:

json Copy
{
  "email": "user@example.com",
  "password": "secret"
}

For GET requests, they are passed in the query string:

text Copy
GET /task_history?email=user@example.com&password=secret

This method is less secure for GET requests: the URL can end up in the client's history, proxy logs, and server logs. Use an API key whenever possible.

2.3. Per-Endpoint Requirements

Endpoint Authorization
POST /tasks/request required
POST /tasks/sync required
GET /task_history required
GET /tasks/result/:response_id not required

With invalid or missing credentials, or a blocked account, the API returns:

http Copy
401 Unauthorized
json Copy
{
  "error": "Unauthorized",
  "status": "error"
}

If the credentials are valid but the balance is depleted (≤ 0), the response also has HTTP status 401, with a different body:

json Copy
{
  "error": "Balance depleted — top up to continue",
  "status": "error"
}

This differs from 402 Payment Required: with 402, the balance is positive but not enough for a new task, taking already reserved funds into account. An unconfirmed email does not by itself block access to the API.


3. General Request Rules

  • Use Content-Type: application/json for POST requests.
  • Method parameters are passed flat in the JSON body, without a nested params object.
  • task_type can be passed in the body or in the query string. If specified in both places, the body takes priority.
  • The JSON body size must not exceed 10,000 bytes.
  • The email and password fields are used only for authorization and are not stored with the task.
  • The task cost is determined by billing. The price field in the x-debug header is for internal use only; do not use it for calculations.

Correct:

json Copy
{
  "task_type": "scrape",
  "url": "https://example.com",
  "data_format": "raw"
}

Do not use a nested params object unless a separate API version explicitly requires it:

json Copy
{
  "task_type": "scrape",
  "params": {
    "url": "https://example.com"
  }
}

4. The x-debug Header

All API responses contain the x-debug header. Its value is always JSON, regardless of format. It holds metadata for the task and the API response itself:

http Copy
x-debug: {"response_id":"0193f2a4-1b2c-7d3e-8f4a-5b6c7d8e9f0a","add_datetime":1747983214000,"finish_datetime":1747983220000,"price":0.0005,"status_code":200,"status_message":"OK"}
Field Type Description
response_id string or null task identifier; null if the task was not created
add_datetime number or null task creation time, Unix milliseconds
finish_datetime number or null completion time; null while the task is running
price number internal field; do not use for calculations, the task cost is determined by billing
status_code number Scraper API HTTP status
status_message string HTTP status text

The API's HTTP status, the method's outcome, and the target page's HTTP status are different values:

text Copy
HTTP 200 from Scraper API — the task is complete
├── status: "success" — the scrape method returned the page without issues
└── http_code: 404 — the target site responded with 404

4.1. The x-debug_response Header

This header contains the service fields of a specific method's result. It is available with both format: "json" and format: "raw":

http Copy
x-debug_response: {"http_code":404,"status":"warn","warning":"waitFor not met within 30s (element=#comments) — page returned as is"}

For scrape, it includes status, http_code, and, if present, warning; if the site is unavailable — status, http_code, and error. The target page headers (headers) and the error address (url) are not included. The header contains only some of the result fields: the full set is available in the JSON body.

x-debug_response may be present on 200 and 422 responses if the worker provided service fields. If it did not, the header is absent — this is not an error in itself. The header is never present on 202 and 410. Its presence or absence does not indicate whether the request succeeded.

The value is always ASCII JSON: non-Latin characters are escaped as \uXXXX. Standard JSON parsing restores the original text.

4.2. Service status vs. Method status

In JSON responses that the service itself generates for task creation, waiting, or errors, the status field describes the task state:

Value When
pending the task has been created, is running, or did not finish within the synchronous wait but is still running
error the request was rejected or the task failed
expired the result's retention period has expired

On HTTP 200, the execution and result endpoints return the method's result as is, without a service wrapper. The method's own status field describes the result. For scrape, it is success, warn, or error (see sections 7.2 and 7.11); for google_search, it is success, no_results, or error (see section 8.3).

For example, HTTP 408 with status: "pending" means the synchronous wait has ended but the task is still running. HTTP 200 with status: "error" in a scrape result means the task is complete but the target site is unavailable. This is different from an HTTP 422 service error.

For service responses of POST endpoints with format: "raw", a plain string without a status field is usually returned: the task ID or the error text. GET /tasks/result/:response_id has no format parameter, and its 202, 404, 410, and 422 responses are always JSON. Site unavailability in scrape is also always described in JSON, including with format: "raw".

In task history, status has separate values: done and error (see section 9).


5. Synchronous Execution: POST /tasks/sync

The endpoint creates a task and waits for it to complete within a single HTTP request.

5.1. General Parameters

Field Type Required Default Description
task_type string yes — method: scrape or google_search
format json or raw no json format of the method result
timeout number no 60 wait time in seconds, allowed range 1–120
email string when authorizing by password — user email
password string when authorizing by password — user password
uule string no — Canonical Name or lat,lon[,radius]; radius in meters, default 200; for google_search, also a ready-made Google value (see section 8.7). scrape also accepts this parameter; its effect on the browser's geolocation is not documented
cdpurl string no — WebSocket URL of your own Chrome/CDP for the scrape method; if the connection fails, the task ends with a 422 error
other fields depends on the method depends on the method — scrape or google_search parameters

5.2. Successful Response

On successful completion, the API returns:

  • HTTP 200;
  • the task result in the body;
  • task metadata in x-debug;
  • result service fields in x-debug_response, if the worker provided them.

For scrape with format: "json", the actual structure used is:

json Copy
{
  "status": "success",
  "http_code": 200,
  "headers": {
    "content-type": "text/html; charset=utf-8"
  },
  "body": "<!DOCTYPE html>..."
}

For scrape with format: "raw", the contents of body are returned directly: HTML, Markdown, or a binary PNG. The status, http_code, and warning fields are available in x-debug_response. If the site is unavailable, the response is always JSON with status: "error", even with format: "raw" (see section 7.11).

google_search result formats are described in sections 8.3, 8.6, and 8.9.

5.3. Timeout

If the task does not complete within the specified time:

http Copy
408 Request Timeout

With format: "json":

json Copy
{
  "error": "timeout",
  "response_id": "0193f2a4-1b2c-7d3e-8f4a-5b6c7d8e9f0a",
  "status": "pending"
}

With format: "raw", the 408 response body contains only the task ID:

text Copy
0193f2a4-1b2c-7d3e-8f4a-5b6c7d8e9f0a

The task continues running. Use the received response_id for a subsequent /tasks/result/:response_id request.

5.4. Execution Error

If the task finished with an error:

http Copy
422 Unprocessable Entity

For GoogleParser, ScrapeParser, and CDP connection errors, the body depends on the endpoint and the format of the original task:

Request Task format 422 Response Content-Type Body
POST /tasks/sync json application/json; charset=utf-8 JSON {"error":"…","status":"error"}
POST /tasks/sync raw text/plain error text without a JSON wrapper
GET /tasks/result/:response_id json or raw application/json; charset=utf-8 JSON {"status":"error","error":"…"}

For example, an invalid waitFor returns ScrapeParser: params.waitFor must be an object. For GoogleParser, the text starts with GoogleParser:, and a failed cdpurl connection starts with CDP connect failed. The max_restarts_exceeded error means the task has used up the available number of execution attempts.


6. Asynchronous Execution

The asynchronous scenario consists of two steps:

  1. create a task via POST /tasks/request;
  2. get the result via GET /tasks/result/:response_id.

6.1. Creating a Task: POST /tasks/request

Example:

bash Copy
curl -X POST "$BASE_URL/tasks/request" \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "task_type": "scrape",
    "url": "https://example.com",
    "data_format": "raw",
    "format": "json"
  }'

Successful response:

http Copy
201 Created
json Copy
{
  "response_id": "0193f2a4-1b2c-7d3e-8f4a-5b6c7d8e9f0a",
  "status": "pending"
}

With format: "raw", the body contains only the ID:

text Copy
0193f2a4-1b2c-7d3e-8f4a-5b6c7d8e9f0a

6.2. Webhook

For an asynchronous task, you can specify:

Field Type Description
webhook_url string notification URL
webhook_method GET or POST defaults to GET
webhook_data string or object custom webhook data

webhook_url is the address the notification is sent to. The request_url field in the notification contains the URL of the task's target page.

Example:

json Copy
{
  "task_type": "scrape",
  "url": "https://example.com",
  "webhook_url": "https://client.example.com/scrape-callback",
  "webhook_method": "POST",
  "webhook_data": {
    "order_id": "A-10042"
  }
}

For POST, the API sends:

json Copy
{
  "status": 200,
  "response_id": "0193f2a4-1b2c-7d3e-8f4a-5b6c7d8e9f0a",
  "request_url": "https://example.com",
  "status_message": "OK",
  "webhook_data": {
    "order_id": "A-10042"
  }
}

For GET (the default webhook_method), the response_id, status, request_url, and status_message fields are appended to the webhook URL as query parameters. For example:

text Copy
https://client.example.com/scrape-callback?response_id=0193f2a4-1b2c-7d3e-8f4a-5b6c7d8e9f0a&status=200&request_url=https%3A%2F%2Fexample.com&status_message=OK

Do not confuse the numeric status field of the notification with the string status in the method result.

Webhook delivery errors are logged but do not change the task result.

6.3. Retrieving the Result

bash Copy
curl "$BASE_URL/tasks/result/$RESPONSE_ID"

The endpoint does not require authorization. On 200, it returns the stored result content as is, without an additional wrapper; it has no format parameter of its own. On 202, 404, 410, and 422, the service returns JSON regardless of the format specified when the task was created.

Status Meaning Action
200 OK task completed successfully read the result
202 Accepted task is still running retry in a few seconds
404 Not Found ID not found check response_id
410 Gone result deleted after retention period expired create the task again
422 Unprocessable Entity task finished with an error read the error field in the JSON

With 202:

json Copy
{
  "status": "pending"
}

With 404:

json Copy
{
  "status": "error",
  "error": "Not found"
}

With 410 (the task record is kept, but the result body has been deleted):

json Copy
{
  "status": "expired",
  "error": "Result expired"
}

With 422 (including when the original request had format: "raw"), JSON is returned with Content-Type: application/json; charset=utf-8:

json Copy
{
  "status": "error",
  "error": "ScrapeParser: params.waitFor must be an object"
}

Besides x-debug, this endpoint also returns x-debug_request with the parameters of the found task, and, when available, the task data. Example for a scrape task:

http Copy
x-debug_request: {"response_id":"0193f2a4-1b2c-7d3e-8f4a-5b6c7d8e9f0a","ip":"1.2.3.4","add_datetime":1747983214000,"finish_datetime":1747983220000,"price":0.0005,"task_type":1,"params":{"task_type":"scrape","url":"https://example.com"}}
Field Type Description
response_id string requested task ID
ip string client IP address
add_datetime number or null creation time, Unix milliseconds
finish_datetime number or null completion time, Unix milliseconds
price number internal field; do not use for calculations, the task cost is determined by billing
task_type number numeric task type ID: 1 — scrape, 2 — google_search
params object or string task parameters

The numeric task_type is used only in this header. In requests, the method is still passed as a string, for example "task_type": "google_search". For a completed task, the header contains the fields from the table; there is no time field in this structure.

Task fields are present on 200, 410, and 422. On 202 and 404, the header contains only response_id and ip. The x-debug_response header may be present on 200 and 422; it is absent on 202 and 410 (see section 4.1).


7. The scrape Method

The method fetches a single web page and returns HTML, Markdown, or a screenshot.

7.1. Parameters

Field Type Required Default Description
task_type string yes — always scrape
url string yes — full URL with http:// or https://
data_format string no raw raw, markdown, or screenshot
format string no json json or raw
waitFor JSON object no — condition to wait for before capturing the result; 30 seconds maximum
fullPage boolean no false full-page screenshot; applies to screenshot
cdpurl string no — your own Chrome via CDP

An address without a scheme, such as example.com, is not accepted: the task ends with HTTP 422 and ScrapeParser: invalid URL "example.com". An address like //example.com also results in 422. The data: and about: schemes are not suitable for fetching a web page: the result may be empty with http_code: 0. An ftp:// address may end in a loading error. Use http:// or https:// to fetch a page.

The uule field is accepted in a scrape request, but its effect on the browser's geolocation is not documented. Do not rely on it as the only way to set the location for this method.

waitFor is passed as a JSON object:

json Copy
{
  "waitFor": {
    "text": "Loaded"
  }
}

A string, even one containing valid JSON, is not supported and causes the task to fail.

7.2. data_format

Value format: "json" format: "raw"
raw JSON with status, http_code, headers, HTML in body HTML directly
markdown JSON with Markdown in body Markdown directly
screenshot JSON with Base64 PNG in body binary PNG

A normal successful response with format: "json" has Content-Type: application/json; charset=utf-8. With format: "raw", the type depends on data_format:

data_format Content-Type of a successful raw response
raw text/html
markdown text/markdown
screenshot image/png

If the site is unavailable, a raw response changes its type to application/json; charset=utf-8 (see section 7.11). Check the actual header before processing the body.

With format: "json", the fetched page result contains:

Field Type Description
status string success — no issues; warn — warning is present; error — the site is unavailable (see section 7.11)
http_code integer HTTP code of the target site; 0 on a network error
headers object response headers of the target site
body string HTML, Markdown, or Base64 PNG according to data_format
warning string only if the waitFor condition was not met within 30 seconds; in that case status: "warn"

status: "success" means the method retrieved the page without issues and does not require HTTP 2xx from the site. For example, a site response of 404 is returned with HTTP 200 from the API, status: "success", and http_code: 404. Site 4xx codes, including 403 and 404, are not considered unavailability.

With format: "raw", the body contains HTML, Markdown, or binary PNG. The status, http_code, and warning metadata are passed in x-debug_response, and the site's headers are not available in this mode. If the site is unavailable, JSON is returned instead of the content (see section 7.11).

7.3. Get HTML as JSON

bash Copy
curl -X POST "$BASE_URL/tasks/sync" \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "task_type": "scrape",
    "url": "https://example.com",
    "data_format": "raw",
    "format": "json"
  }'

Response:

json Copy
{
  "status": "success",
  "http_code": 200,
  "headers": {
    "content-type": "text/html; charset=utf-8"
  },
  "body": "<!DOCTYPE html>..."
}

7.4. Get HTML Directly

bash Copy
curl -X POST "$BASE_URL/tasks/sync" \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "task_type": "scrape",
    "url": "https://example.com",
    "data_format": "raw",
    "format": "raw"
  }'

The response begins with:

html Copy
<!DOCTYPE html>
<html>

7.5. Get Markdown

JSON:

json Copy
{
  "task_type": "scrape",
  "url": "https://example.com",
  "data_format": "markdown",
  "format": "json"
}

Direct:

json Copy
{
  "task_type": "scrape",
  "url": "https://example.com",
  "data_format": "markdown",
  "format": "raw"
}

Markdown example:

md Copy
# Example Domain

[More information](https://www.iana.org/domains/example)

7.6. Take a Viewport Screenshot

Base64 in JSON:

json Copy
{
  "task_type": "scrape",
  "url": "https://example.com",
  "data_format": "screenshot",
  "format": "json",
  "fullPage": false
}

Response:

json Copy
{
  "status": "success",
  "http_code": 200,
  "headers": {
    "content-type": "text/html; charset=utf-8"
  },
  "body": "iVBORw0KGgoAAAANSUhEUg..."
}

The body field contains Base64 without the data:image/png;base64, prefix.

7.7. Get a Binary PNG

bash Copy
curl -X POST "$BASE_URL/tasks/sync" \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "task_type": "scrape",
    "url": "https://example.com",
    "data_format": "screenshot",
    "format": "raw"
  }' \
  --output screenshot.png

The first eight bytes of the file:

text Copy
89 50 4E 47 0D 0A 1A 0A

7.8. Take a Full-Page Screenshot

json Copy
{
  "task_type": "scrape",
  "url": "https://example.com/long-page",
  "data_format": "screenshot",
  "format": "json",
  "fullPage": true
}

For data_format: "raw", the fullPage parameter is effectively ignored and does not change the HTML result.

7.9. waitFor

The condition is specified as a JSON object with text, element, or state. For the CSS selector element, checkVisible defaults to false.

Wait for Text

json Copy
{
  "task_type": "scrape",
  "url": "https://example.com/dynamic",
  "waitFor": {"text": "Content loaded"}
}

The page is suitable for this test only if the string is absent from the initial HTML and is added later.

Wait for an Element

json Copy
{
  "task_type": "scrape",
  "url": "https://example.com",
  "waitFor": {"element": "#content", "checkVisible": false}
}

Wait for a Visible Element

json Copy
{
  "task_type": "scrape",
  "url": "https://example.com",
  "waitFor": {"element": "#content", "checkVisible": true}
}

Wait for Full Load

json Copy
{
  "task_type": "scrape",
  "url": "https://example.com",
  "waitFor": {"state": "load"}
}

Wait for DOM Construction

json Copy
{
  "task_type": "scrape",
  "url": "https://example.com",
  "waitFor": {"state": "domcontentloaded"}
}

Supported state values:

Value Description
load the page and its dependent resources have loaded
domcontentloaded the DOM is built, resources may still be loading

Wait Limit and Warnings

waitFor waiting is limited to 30 seconds. Increasing the overall timeout does not extend this limit; timeout in /tasks/sync sets how long the HTTP request waits for the task to complete.

If the condition is not met within 30 seconds, the task does not fail: the fetched page is returned as is. The result then contains status: "warn" and warning:

json Copy
{
  "status": "warn",
  "http_code": 200,
  "headers": {
    "content-type": "text/html; charset=utf-8"
  },
  "body": "<!DOCTYPE html>...",
  "warning": "waitFor not met within 30s (element=#comments) — page returned as is"
}

If the condition is met or waitFor is not set, there is no warning field. With format: "raw", the warning is available in x-debug_response. The warning value uses an em dash — (U+2014) before page returned as is; in the header's JSON value it is transmitted as —. When handling warnings, do not compare the entire string against a hard-coded pattern.

waitFor Errors

An invalid value ends the task immediately, without retries and without contacting the target site. The error text comes with the ScrapeParser: … prefix:

Error Text Without Prefix Cause
params.waitFor must be an object a string (including one with JSON inside), number, array, or true was passed
params.waitFor must have "element", "text" or "state" the object has no known fields or element is empty
params.waitFor.state must be one of: load, domcontentloaded unknown state value

7.10. Your Own Chrome via cdpurl

json Copy
{
  "task_type": "scrape",
  "url": "https://example.com/account",
  "cdpurl": "wss://browser.example.com/devtools/browser/9b2c1f0a-..."
}

Requirements:

  • ws:// or wss:// is supported;
  • the browser must be reachable by the worker for the entire execution;
  • the task uses that browser's cookies, sessions, profile, fingerprint, and proxy;
  • do not publish the CDP URL: access to it effectively grants access to the browser session.

If cdpurl is not specified, the worker picks the browser. If it is specified but the connection fails after two attempts, the task ends with HTTP 422, with no further retries and without substituting the worker's own browser for yours. The error text starts with:

text Copy
CDP connect failed (user cdpurl) after 2 attempts: Timeout 12000ms exceeded.
CDP connect failed (user cdpurl) after 2 attempts: WebSocket error: <endpoint> 401 Unauthorized deny_no_user
CDP connect failed (user cdpurl) after 2 attempts: WebSocket error: <endpoint> 500 Internal Server Error profile_locked

In the error text, the CDP address is replaced with <endpoint> so that credentials from the URL do not end up in the response or task history. Check that the browser is running and reachable from outside, that the link has not become stale after a Chrome restart, and that the profile is not in use by another connection.

7.11. Site Unavailability

On a network error (DNS, connection refused, TLS error, load timeout) or a 5xx response from the site, the task is considered successfully completed at the API level: the HTTP response code is 200, and the method result contains status: "error" and the reason:

json Copy
{
  "status": "error",
  "http_code": 0,
  "error": "net::ERR_NAME_NOT_RESOLVED at https://example.invalid",
  "url": "https://example.invalid"
}
Field Description
status error — the site is unavailable
http_code 0 on a network error, or the site's actual 5xx HTTP code
error reason text
url the address where the error occurred

In this case the response is always JSON, including with format: "raw". On a network error, Content-Type is application/json; charset=utf-8 with both format: "raw" and format: "json"; the outer HTTP status is 200. The x-debug_response header duplicates status, http_code, and error; the url field is available only in the body. Do not save such a response as HTML or PNG without checking Content-Type and the result.

Site 4xx responses are returned as a normal result with the corresponding http_code. A site load timeout is different from the synchronous endpoint's HTTP 408: with 408, the task itself keeps running (see section 5.3).


8. The google_search Method

The method performs a Google query and returns structured results grouped by type, or a Markdown document. Authorization, synchronous and asynchronous execution, webhooks, and result retrieval follow the general rules in sections 2–6.

8.1. Parameters

All fields are passed flat in the JSON body, without a nested params.

Field Type Required Default Description
task_type string yes — always google_search
url string yes — full search URL; the domain must be google.com or one of its subdomains; the q parameter is required
format json or raw no json body format; with data_format: "markdown", raw returns the document directly. Without Markdown, raw returns the full JSON result (see section 8.9)
data_format string no standard grouped results markdown — a document; any other value, including json, — standard grouped results
page_num or pages integer no 1 number of result pages, an integer from 1 to 10
gl string no — two-letter country code; overrides gl inside url
hl string no — two-letter language code; overrides hl inside url
with_ai boolean no true collect the AI answer on standard results
with_preview boolean no false add previews to news, video, and books items
with_html boolean no false add the HTML of each page; only with format: "json"
saveScreenshot boolean no false add a screenshot of each page; only with format: "json"
uule string no — Canonical Name, lat,lon[,radius], or a ready-made Google value; radius in meters, default 200

To set the number of pages, use one of the fields: page_num or pages. The timeout parameter of the synchronous request is described in section 5.1.

8.2. Parameters Inside url

The method keeps the following Google parameters and removes all others, because an unknown parameter in the request increases the likelihood of being blocked. The uule geolocation is handled separately: it can be passed in the URL or in the request body (see section 8.7).

Parameter Example Description
q q=fastify+nodejs required search query
oq oq=fastify original query before autocomplete
hl hl=en results page language
gl gl=us search country; if gl is in the URL and no geolocation is set explicitly, the API adds the country's uule. A gl field in the body alone does not trigger this (see section 8.7)
start start=20 results offset; non-negative, a multiple of 10, no more than 90
tbm tbm=isch search type from the table below
udm udm=2 search type from the table below
tbs tbs=qdr:d Google filters, including time filters
lr, cr lr=lang_en restrict documents by language and country
safe safe=active SafeSearch
pws pws=0 0 disables results personalization
nfpr nfpr=1 search without typo autocorrection

Spaces in URL parameter values are encoded as %20 or +.

Search Types

For standard web results, do not set tbm or udm. In rows with two options, one of them is enough.

Results Parameter layout Main Group
Standard web results no tbm or udm organic organic
Places udm=1 or tbm=lcl places places
Images udm=2 or tbm=isch images images
Videos udm=7 or tbm=vid video video
Jobs udm=8 jobs jobs
News udm=12 or tbm=nws news news
Books udm=36 or tbm=bks books books
Shorts udm=39 shorts video
AI Mode udm=50 ai_mode ai_overview

Other tbm and udm values are not supported and cause the task to fail. In particular, udm=14, udm=18, and udm=56 are not on the allowed list. Shopping (udm=28, tbm=shop) is temporarily unavailable.

Results Depth

A results page contains up to 10 results. For multiple pages, set page_num from 1 to 10; the num and filter parameters are not supported.

start=20 starts collection from the third page. Page and position numbers are absolute in this case (see section 8.8).

json Copy
{
  "task_type": "google_search",
  "url": "https://www.google.com/search?q=fastify+nodejs&hl=en&gl=us&start=20",
  "page_num": 2,
  "format": "json"
}

To get only the second page, specify start=10 without page_num:

json Copy
{
  "task_type": "google_search",
  "url": "https://www.google.com/search?q=fastify+nodejs&hl=en&gl=us&start=10"
}
bash Copy
curl -X POST "$BASE_URL/tasks/sync" \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "task_type": "google_search",
    "url": "https://www.google.com/search?q=fastify+nodejs&hl=en&gl=us",
    "page_num": 1,
    "format": "json",
    "data_format": "json",
    "with_html": false
  }'

Example response structure with a single item:

json Copy
{
  "status": "success",
  "results": 484000000,
  "layout": "organic",
  "task_param": {
    "url": "https://www.google.com/search?q=fastify+nodejs&hl=en&gl=us",
    "pages": 1
  },
  "real_location": "United States",
  "groups": {
    "organic": {
      "count": 1,
      "items": [
        {
          "url": "https://fastify.dev/",
          "title": "Fastify",
          "description": "Fast and low overhead web framework for Node.js",
          "position": 1,
          "page_position": {
            "page": 1,
            "position": 1
          }
        }
      ]
    }
  }
}
Field Type Description
status string success — results collected; no_results — no items, this is not an error; error — an error, its text may be in the error field
results number number of results according to Google; -1 if the counter could not be read
layout string main results type from section 8.2; unknown if the layout was not recognized
task_param object parameters received by the handler, including geolocation added by the server
real_location string or null region actually applied by Google; null if the label could not be read
groups object results by group; a regular group contains count and an items array, ai_overview has a separate structure (see section 8.8)
error string optional error text when status: "error"
html array HTML of each page; only with with_html: true and format: "json"
screenshots array page screenshots in Base64; only with saveScreenshot: true and format: "json"

Organic items are in groups.organic.items. The organic group may be absent, for example in image search. A single response may contain groups of different types.

results: 0 and results: -1 mean different things: 0 means Google reported no results; -1 means the counter is unavailable. Jobs and AI Mode results have no counter, so they always return -1. Check for items using status and groups, not results alone. Google's result count may vary between sessions for the same query.

With with_html: false, the html field is not added. To get page HTML and screenshots along with the result:

json Copy
{
  "task_type": "google_search",
  "url": "https://www.google.com/search?q=fastify+nodejs&hl=en&gl=us",
  "format": "json",
  "with_html": true,
  "saveScreenshot": true
}

To skip collecting the AI answer on standard results, pass with_ai: false. Even with collection enabled, whether an AI Overview appears depends on whether Google showed one.

json Copy
{
  "task_type": "google_search",
  "url": "https://www.google.com/search?q=modern+furniture&hl=en&gl=us&tbm=isch",
  "format": "json",
  "data_format": "json"
}

Alternatively:

text Copy
https://www.google.com/search?q=modern+furniture&hl=en&gl=us&udm=2

The main group is groups.images, and layout is images. The item's url field points to the source page, and origin_image_url to the original image. The preview is available in image_url or image_base64.

8.5. News

json Copy
{
  "task_type": "google_search",
  "url": "https://www.google.com/search?q=technology&hl=en&gl=us&tbm=nws",
  "format": "json",
  "data_format": "json"
}

The main group is groups.news, and layout is news. You can use udm=12 instead of tbm=nws.

News from the last 24 hours, with previews:

json Copy
{
  "task_type": "google_search",
  "url": "https://www.google.com/search?q=technology&hl=en&gl=us&udm=12&tbs=qdr:d",
  "format": "json",
  "with_preview": true
}

The time field contains the age as text, as Google displays it. Without with_preview, news items have no image_url or image_base64 fields.

8.6. Result in Markdown

With data_format: "markdown", the results are assembled into a single document.

In JSON:

json Copy
{
  "task_type": "google_search",
  "url": "https://www.google.com/search?q=fastify+nodejs&hl=en&gl=us",
  "format": "json",
  "data_format": "markdown"
}

The response contains status, results, layout, task_param, real_location, and a body string with the document. There is no groups field in this mode. With with_html: true, page HTML is added to the JSON result in the html field. With saveScreenshot: true, a screenshots field is added to the Markdown JSON result; for a single page, it contains one Base64 PNG.

Directly:

bash Copy
curl -X POST "$BASE_URL/tasks/sync" \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "task_type": "google_search",
    "url": "https://www.google.com/search?q=fastify+nodejs&hl=en&gl=us",
    "format": "raw",
    "data_format": "markdown"
  }' \
  --output search.md

When the document is retrieved successfully, the body is Markdown with Content-Type: text/markdown; the html and screenshots attachments are not included. Check the HTTP status and Content-Type before using the file: a service error is not a Markdown result.

The document consists of ## <group> sections with numbered items; images are embedded in the document as pictures. The ## Task information section and the yaml block may be absent. For task parameters, use task_param with format: "json"; with format: "raw", do not rely on them being present in the document text.

8.7. Geo-Targeting

Geolocation can be set via uule in the request body or inside url. The value from the URL is extracted automatically; there is no need to duplicate it in a separate field. If different values are passed in both places, no error is returned: the value from the URL takes precedence. Check the applied geolocation in task_param and real_location.

Canonical Name

json Copy
{
  "task_type": "google_search",
  "url": "https://www.google.com/search?q=pizza&hl=en&gl=us",
  "uule": "New York,New York,United States"
}

A Canonical Name is a location name from the Google Ads Geo Targets database. It must match the database entry exactly. The number of parts varies by city: for example, the canonical name for Paris has four parts.

text Copy
Paris,Paris,Ile-de-France,France
New York,New York,United States
London,England,United Kingdom

Inside a URL, spaces are encoded as usual:

text Copy
https://www.google.com/search?q=pizza&gl=us&uule=New+York,New+York,United+States

Coordinates

For coordinates, use a uule string in the lat,lon or lat,lon,radius format:

json Copy
{
  "task_type": "google_search",
  "url": "https://www.google.com/search?q=pizza&hl=en&gl=us",
  "uule": "40.7128,-74.0060,5000"
}

Latitude and longitude are specified in decimal degrees, and the radius in meters. If no radius is specified, 200 meters is used. In the example, the radius is 5 km; a value of 10 would mean 10 meters.

Ready-Made Google Value

google_search also accepts a ready-made uule value copied from the address bar of a Google results page. It is passed to Google unchanged and is not checked against the region directory, so it can specify a country or a region at any level.

Prefix Content
w+ region name
a+ latlng coordinates and radius in meters

For example, a search with a ready-made value for the United Kingdom:

json Copy
{
  "task_type": "google_search",
  "url": "https://www.google.com/search?q=coffee+shops&hl=en",
  "uule": "w+CAIQICIOVW5pdGVkIEtpbmdkb20"
}

The ready-made value can also be passed inside url; the server extracts it automatically.

Geolocation by gl

If uule is not set explicitly, automatic addition is triggered by gl in the URL. Both gl=gb and gl=uk result in the same uule for the United Kingdom. For the ISO country code, use gb; the API also accepts uk. An explicitly set uule takes precedence over automatic addition by gl.

A gl field only in the request body does not trigger automatic uule addition. In that case, the resulting real_location may differ from the expected country. For a reproducible region, pass uule explicitly or put gl in the URL, and check real_location. Without gl or geolocation, Google chooses the region by IP; real_location: null means the region label could not be read.

8.8. Result Groups and Positions

Regular groups have the structure { "count": N, "items": [...] }. The ai_overview group is structured differently.

Group Item Fields
organic url, title, description
news url, title, description, source, time; image_url, image_base64 — with with_preview
video url, title, description, is_shorts; image_url, image_base64 — with with_preview
images url (source page), title, image_url, image_base64, origin_image_url, origin_image_width, origin_image_height
books url, title, description; image_url, image_base64 — with with_preview
ads url, title, advertiser, display_url
places title, url, place_id, rating, reviews, category, price_level, address, distance, hours, review_quote, image_url, image_base64
jobs title, company, location, source, tags, url, job_id, image_url, image_base64
ai_overview an object with text and a sources array; a source contains url, label, and, if available, domain

Standard results may include news, videos, places, and ads alongside organic results. In that case, layout remains organic. For example, the presence of groups.places does not by itself mean layout: "places".

Numbering

Field Description
position running item number across the collected pages, taking start into account
page_position an object with the absolute page number and the position on it, for example { "page": 3, "position": 1 }

When multiple pages are collected, items are merged into a single list for the corresponding group. The start offset applies only to the group being paginated: for standard search, that is organic. News, video, and places blocks are numbered from one regardless of start.

Previews and Missing Data

For previews, either image_url (a link) or image_base64 (an image embedded by Google in the page) is filled in. Without with_preview, these fields are absent in news, video, and books.

In places, jobs, and books, some fields may be null if the data is not on the page. For example, this can happen with rating, number of reviews, price level, distance, opening hours, review quote, and place images, a job logo, or a book cover. Consider missing previews in books together with the with_preview setting.

null is not the same as zero or an empty string: rating: null means there is no rating. A special case is places.url: if the place has no ID in the Google Knowledge Graph, the field may be an empty string.

AI Overview

ai_overview.text contains Markdown. The sources field is an array of sources with addresses and labels. If a source link could not be resolved, url points to a Google redirect, and there is no domain field.

Google does not show an overview for every query. If there is no overview, the corresponding group is absent from the result. The with_ai parameter controls AI answer collection on standard results; the separate AI Mode is selected with udm=50.

8.9. The x-debug_response Header and format: "raw"

Task metadata is in x-debug. Search results service fields are available in x-debug_response with both format values, subject to the general conditions in section 4.1:

http Copy
x-debug_response: {"status":"success","results":484000000,"layout":"organic","real_location":"United States"}

Items and groups are not included in the header. With format: "raw" without data_format: "markdown" (either without data_format or with data_format: "json"), the API returns the full JSON result: status, results, layout, task_param, real_location, and groups. Content-Type is application/json; charset=utf-8. To get Markdown directly, use format: "raw" together with data_format: "markdown" (see section 8.6); the body is then the document text.

8.10. Parameter Errors and Retries

Parameter errors end the task immediately with HTTP 422, without automatic retries. The error text has the GoogleParser: … prefix.

Error Text Without Prefix Cause
params.url required the url field was not passed
invalid URL "…" the value cannot be parsed as a URL, for example a search phrase or an address without a scheme was passed
URL domain must be google.com, got "…" the domain is not google.com or one of its subdomains
URL missing required parameter "q" the q search query is missing
params.pages must be an integer 1..10 the number of pages is fractional, non-numeric, or outside the 1–10 range
"start" must be a non-negative multiple of 10 the offset is negative or not a multiple of 10
"start" must be <= 90 the offset is greater than 90
unsupported search type message tbm or udm is not on the allowed list; the allowed values are listed in the error text

For parameter errors, the method passes status: "error" and a description in x-debug_response. Header example:

http Copy
x-debug_response: {"status":"error","error":"params.url required"}

Besides parameter errors, a task may end with 422 and the text max_restarts_exceeded: it could not be completed within the allotted number of attempts. In this case, repeat the request. Handle it as a task error, based on the HTTP status and the actual Content-Type.

The general 422 response formats are described in sections 5.4 and 6.3. When parsing the response, consider the HTTP status, Content-Type, and the error description in x-debug_response.

A CAPTCHA, a Google sign-in page, a dropped connection, and a timeout are not parameter errors: such tasks are retried automatically. The synchronous wait timeout, HTTP 408, is handled according to the general rules — save the response_id and request the result later, without creating a duplicate task.


9. Task History

bash Copy
curl "$BASE_URL/task_history?task_type=scrape&from=2026-01-01&to=2026-12-31&limit=50&offset=0" \
  -H "Authorization: Bearer $API_KEY"
Parameter Type Default Description
task_type string — filter by method
from DateTime — lower bound of add_datetime
to DateTime — upper bound of add_datetime
limit integer 100 number of records
offset integer 0 offset

Response:

json Copy
[
  {
    "response_id": "0193f2a4-1b2c-7d3e-8f4a-5b6c7d8e9f0a",
    "add_datetime": 1747983214000,
    "finish_datetime": 1747983220000,
    "task_type": "scrape",
    "params": {
      "task_type": "scrape",
      "url": "https://example.com",
      "format": "json"
    },
    "price": 0.0005,
    "mime_type": "application/json",
    "error": "",
    "status": "done"
  }
]

The params field contains the request parameters, including task_type and method fields; the email and password authorization fields are not stored in the task.

The history contains only completed tasks. Its status field is derived from the error field: an empty string means done, and error text means error. There is no pending value in the history. This status refers to task completion and does not replace the status inside a scrape or google_search result.


10. Common Errors

HTTP Status Cause
400 Bad Request invalid parameters, unknown or disabled task_type, body larger than 10,000 bytes
401 Unauthorized invalid or missing credentials, blocked account, or depleted balance (≤ 0); the reason is given in error
402 Payment Required the balance is positive but insufficient for a new task, taking reserves into account; the task is not created
404 Not Found the requested task ID was not found
408 Request Timeout synchronous wait ended; the task keeps running, the service status is pending
410 Gone result deleted after retention period
422 Unprocessable Entity task finished with an error
503 Service Unavailable acceptance of tasks of this type is temporarily suspended — Service overloaded, retry after Retry-After; on /health/ready — infrastructure not ready

Example error in JSON:

json Copy
{
  "error": "Insufficient balance",
  "status": "error"
}

With format: "raw", the API may return plain text only:

text Copy
Insufficient balance

On 400, the JSON response also contains status: "error". Example error values:

  • params exceeds 10 000 bytes;
  • Unknown task type: unknown_name;
  • Task type temporarily disabled.

For GET /tasks/result/:response_id, waiting and error responses are always JSON (see section 6.3).

When a task is created, the balance is checked taking already reserved funds and the cost of the new task into account. On 402, the task is not created. The actual charge is made after the task completes successfully; the cost is determined by billing, and x-debug.price is not intended for calculations.

Service Overload

If workers cannot keep up with the queue, acceptance of tasks of a specific type via /tasks/request or /tasks/sync may be temporarily suspended. When creating a task, the API returns:

http Copy
503 Service Unavailable
Retry-After: 60
json Copy
{
  "error": "Service overloaded",
  "status": "error"
}

With format: "raw", the body contains Service overloaded with Content-Type: text/plain; with format: "json", the response contains JSON with Content-Type: application/json; charset=utf-8. The task is not created, and no balance is reserved. Retry the request after the number of seconds in Retry-After (60 in the example; use the actual value from the response). The restriction is lifted automatically as the queue is processed.

This response is different from the 503 of the /health/ready endpoint, which reports that an infrastructure dependency is unavailable.

Always check:

  1. the API's HTTP status;
  2. the Content-Type;
  3. x-debug.status_code;
  4. x-debug.response_id;
  5. the service status and error, if the response was generated by the service;
  6. for a scrape result — status, http_code, and, if present, warning or error;
  7. for google_search — status, error if present, and groups or body according to data_format; a results: -1 value alone does not mean there are no results.

With format: "raw", use the available x-debug_response fields, keeping in mind that the header is not guaranteed and its presence does not mean success. If the site is unavailable, read the JSON body.


11. Utility Endpoints

No authorization required.

GET /health

bash Copy
curl "$BASE_URL/health"
json Copy
{
  "status": "ok",
  "ts": 1747983214000
}

GET /health/ready

bash Copy
curl "$BASE_URL/health/ready"

Successful response:

json Copy
{
  "status": "ready",
  "deps": {
    "redis_streams": true,
    "redis_cache": true,
    "clickhouse": true,
    "s3": true
  }
}

If a dependency is unavailable, the endpoint returns 503 and status: "degraded".


12. Integration Recommendations

  • Use the synchronous endpoint for short tasks and the asynchronous one for long or bulk tasks. If a task may take longer than half a minute, the asynchronous mode is more reliable: you don't have to keep the connection open.
  • When polling GET /tasks/result/:response_id, retry every few seconds while you receive 202.
  • On 408, do not immediately create a duplicate task: save the response_id and request the result later.
  • Do not treat an outer HTTP 200 as confirmation that the target site also responded with 200.
  • The task cost is determined by billing; do not use the internal x-debug.price field for calculations.
  • For scrape, check the result in this order: first status: "error" and the error field, then status: "warn" and warning, then http_code >= 400, and only after that process body. An outer HTTP 200 does not cancel method or target site errors.
  • On 503 Service overloaded, retry task creation after the interval in Retry-After.
  • For PNGs with format: "raw", check Content-Type before saving: if the site is unavailable, a JSON error is returned instead of a PNG.
  • For PNGs with format: "json", decode the Base64 from the top-level body.
  • Pass waitFor as a JSON object. For its text field, use a string that is absent from the original HTML and appears without user action.
  • Keep in mind the 30-second waitFor limit: when it expires, the page is returned with status: "warn" and a warning field.
  • For google_search with standard structured results, read groups from the JSON. With format: "raw" without Markdown, the result also contains groups; rely on the actual Content-Type. With data_format: "markdown", read the document from body in JSON mode, or get the text directly with format: "raw".
  • Do not treat google_search.status: "no_results" as an error; Google's result count in results is not the number of collected items and may be unavailable (-1).
  • For reproducible search results, set uule explicitly or put gl inside the URL; a gl field in the body alone does not trigger automatic uule addition. Check the actual region in real_location.
  • Do not use external demo sites as permanent fixtures: their content and availability may change.
  • Protect API keys, passwords, webhook URLs, and CDP URLs.