Scraper API — User Guide
This guide describes how to work with the Scraper API directly over HTTP: creating tasks, synchronous execution, retrieving results, viewing history, fetching web pages, and Google search.
Quick Start — get the HTML of a page
The simplest scenario is to synchronously fetch a page and get its HTML as JSON.
You will need:
- the API address;
- an API key;
- the page URL.
Set the API address and key:
bash
BASE_URL="https://scraper.2captcha.com"
API_KEY="<YOUR_API_KEY>"
Send the request:
bash
curl -i -X POST "$BASE_URL/tasks/sync" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"task_type": "scrape",
"url": "https://example.com",
"data_format": "raw",
"format": "json"
}'
Task parameters:
| Field | Value | Purpose |
|---|---|---|
task_type |
scrape |
fetch a web page |
url |
https://example.com |
target page address |
data_format |
raw |
return the raw HTML |
format |
json |
put the result in JSON |
A successful response has HTTP status 200 and a body roughly like this:
json
{
"status": "success",
"http_code": 200,
"headers": {
"content-type": "text/html; charset=utf-8"
},
"body": "<!doctype html><html>...</html>"
}
Here:
status— the outcome of thescrapemethod:success,warn, orerror;http_code— the HTTP status of the target page;headers— the target page's response headers;body— the resulting HTML.
Task metadata is not in the body but in the x-debug response header:
http
x-debug: {"response_id":"...","add_datetime":1747983214000,"finish_datetime":1747983220000,"price":0.0005,"status_code":200,"status_message":"OK"}
Check that:
- the Scraper API's outer HTTP status is
200; statusin the JSON issuccess; onwarn, readwarning; onerror, readerror;http_codematches the expected HTTP status of the target page;bodyis not empty and contains the expected HTML;x-debug.response_idis populated;x-debug.status_codematches the outer HTTP status.
The x-debug.price field is for internal use only and must not be used for calculations: the task cost is determined by billing.
If you only need the HTML without JSON:
bash
curl -X POST "$BASE_URL/tasks/sync" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"task_type": "scrape",
"url": "https://example.com",
"data_format": "raw",
"format": "raw"
}' \
--output page.html
The result will be saved to page.html. If the target site is unavailable, a JSON error description is returned even with format: "raw"; check Content-Type and the result before using the file. Page metadata is available in the x-debug_response header (see section 4.1).
1. API Address
Production:
text
https://scraper.2captcha.com
bash
BASE_URL="https://scraper.2captcha.com"
3. General Request Rules
- Use
Content-Type: application/jsonfor POST requests. - Method parameters are passed flat in the JSON body, without a nested
paramsobject. task_typecan be passed in the body or in the query string. If specified in both places, the body takes priority.- The JSON body size must not exceed 10,000 bytes.
- The
emailandpasswordfields are used only for authorization and are not stored with the task. - The task cost is determined by billing. The
pricefield in thex-debugheader is for internal use only; do not use it for calculations.
Correct:
json
{
"task_type": "scrape",
"url": "https://example.com",
"data_format": "raw"
}
Do not use a nested params object unless a separate API version explicitly requires it:
json
{
"task_type": "scrape",
"params": {
"url": "https://example.com"
}
}
4. The x-debug Header
All API responses contain the x-debug header. Its value is always JSON, regardless of format. It holds metadata for the task and the API response itself:
http
x-debug: {"response_id":"0193f2a4-1b2c-7d3e-8f4a-5b6c7d8e9f0a","add_datetime":1747983214000,"finish_datetime":1747983220000,"price":0.0005,"status_code":200,"status_message":"OK"}
| Field | Type | Description |
|---|---|---|
response_id |
string or null |
task identifier; null if the task was not created |
add_datetime |
number or null |
task creation time, Unix milliseconds |
finish_datetime |
number or null |
completion time; null while the task is running |
price |
number | internal field; do not use for calculations, the task cost is determined by billing |
status_code |
number | Scraper API HTTP status |
status_message |
string | HTTP status text |
The API's HTTP status, the method's outcome, and the target page's HTTP status are different values:
text
HTTP 200 from Scraper API — the task is complete
├── status: "success" — the scrape method returned the page without issues
└── http_code: 404 — the target site responded with 404
4.1. The x-debug_response Header
This header contains the service fields of a specific method's result. It is available with both format: "json" and format: "raw":
http
x-debug_response: {"http_code":404,"status":"warn","warning":"waitFor not met within 30s (element=#comments) — page returned as is"}
For scrape, it includes status, http_code, and, if present, warning; if the site is unavailable — status, http_code, and error. The target page headers (headers) and the error address (url) are not included. The header contains only some of the result fields: the full set is available in the JSON body.
x-debug_response may be present on 200 and 422 responses if the worker provided service fields. If it did not, the header is absent — this is not an error in itself. The header is never present on 202 and 410. Its presence or absence does not indicate whether the request succeeded.
The value is always ASCII JSON: non-Latin characters are escaped as \uXXXX. Standard JSON parsing restores the original text.
4.2. Service status vs. Method status
In JSON responses that the service itself generates for task creation, waiting, or errors, the status field describes the task state:
| Value | When |
|---|---|
pending |
the task has been created, is running, or did not finish within the synchronous wait but is still running |
error |
the request was rejected or the task failed |
expired |
the result's retention period has expired |
On HTTP 200, the execution and result endpoints return the method's result as is, without a service wrapper. The method's own status field describes the result. For scrape, it is success, warn, or error (see sections 7.2 and 7.11); for google_search, it is success, no_results, or error (see section 8.3).
For example, HTTP 408 with status: "pending" means the synchronous wait has ended but the task is still running. HTTP 200 with status: "error" in a scrape result means the task is complete but the target site is unavailable. This is different from an HTTP 422 service error.
For service responses of POST endpoints with format: "raw", a plain string without a status field is usually returned: the task ID or the error text. GET /tasks/result/:response_id has no format parameter, and its 202, 404, 410, and 422 responses are always JSON. Site unavailability in scrape is also always described in JSON, including with format: "raw".
In task history, status has separate values: done and error (see section 9).
5. Synchronous Execution: POST /tasks/sync
The endpoint creates a task and waits for it to complete within a single HTTP request.
5.1. General Parameters
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
task_type |
string | yes | — | method: scrape or google_search |
format |
json or raw |
no | json |
format of the method result |
timeout |
number | no | 60 |
wait time in seconds, allowed range 1–120 |
email |
string | when authorizing by password | — | user email |
password |
string | when authorizing by password | — | user password |
uule |
string | no | — | Canonical Name or lat,lon[,radius]; radius in meters, default 200; for google_search, also a ready-made Google value (see section 8.7). scrape also accepts this parameter; its effect on the browser's geolocation is not documented |
cdpurl |
string | no | — | WebSocket URL of your own Chrome/CDP for the scrape method; if the connection fails, the task ends with a 422 error |
| other fields | depends on the method | depends on the method | — | scrape or google_search parameters |
5.2. Successful Response
On successful completion, the API returns:
- HTTP
200; - the task result in the body;
- task metadata in
x-debug; - result service fields in
x-debug_response, if the worker provided them.
For scrape with format: "json", the actual structure used is:
json
{
"status": "success",
"http_code": 200,
"headers": {
"content-type": "text/html; charset=utf-8"
},
"body": "<!DOCTYPE html>..."
}
For scrape with format: "raw", the contents of body are returned directly: HTML, Markdown, or a binary PNG. The status, http_code, and warning fields are available in x-debug_response. If the site is unavailable, the response is always JSON with status: "error", even with format: "raw" (see section 7.11).
google_search result formats are described in sections 8.3, 8.6, and 8.9.
5.3. Timeout
If the task does not complete within the specified time:
http
408 Request Timeout
With format: "json":
json
{
"error": "timeout",
"response_id": "0193f2a4-1b2c-7d3e-8f4a-5b6c7d8e9f0a",
"status": "pending"
}
With format: "raw", the 408 response body contains only the task ID:
text
0193f2a4-1b2c-7d3e-8f4a-5b6c7d8e9f0a
The task continues running. Use the received response_id for a subsequent /tasks/result/:response_id request.
5.4. Execution Error
If the task finished with an error:
http
422 Unprocessable Entity
For GoogleParser, ScrapeParser, and CDP connection errors, the body depends on the endpoint and the format of the original task:
| Request | Task format |
422 Response Content-Type |
Body |
|---|---|---|---|
POST /tasks/sync |
json |
application/json; charset=utf-8 |
JSON {"error":"…","status":"error"} |
POST /tasks/sync |
raw |
text/plain |
error text without a JSON wrapper |
GET /tasks/result/:response_id |
json or raw |
application/json; charset=utf-8 |
JSON {"status":"error","error":"…"} |
For example, an invalid waitFor returns ScrapeParser: params.waitFor must be an object. For GoogleParser, the text starts with GoogleParser:, and a failed cdpurl connection starts with CDP connect failed. The max_restarts_exceeded error means the task has used up the available number of execution attempts.
6. Asynchronous Execution
The asynchronous scenario consists of two steps:
- create a task via
POST /tasks/request; - get the result via
GET /tasks/result/:response_id.
6.1. Creating a Task: POST /tasks/request
Example:
bash
curl -X POST "$BASE_URL/tasks/request" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"task_type": "scrape",
"url": "https://example.com",
"data_format": "raw",
"format": "json"
}'
Successful response:
http
201 Created
json
{
"response_id": "0193f2a4-1b2c-7d3e-8f4a-5b6c7d8e9f0a",
"status": "pending"
}
With format: "raw", the body contains only the ID:
text
0193f2a4-1b2c-7d3e-8f4a-5b6c7d8e9f0a
6.2. Webhook
For an asynchronous task, you can specify:
| Field | Type | Description |
|---|---|---|
webhook_url |
string | notification URL |
webhook_method |
GET or POST |
defaults to GET |
webhook_data |
string or object | custom webhook data |
webhook_url is the address the notification is sent to. The request_url field in the notification contains the URL of the task's target page.
Example:
json
{
"task_type": "scrape",
"url": "https://example.com",
"webhook_url": "https://client.example.com/scrape-callback",
"webhook_method": "POST",
"webhook_data": {
"order_id": "A-10042"
}
}
For POST, the API sends:
json
{
"status": 200,
"response_id": "0193f2a4-1b2c-7d3e-8f4a-5b6c7d8e9f0a",
"request_url": "https://example.com",
"status_message": "OK",
"webhook_data": {
"order_id": "A-10042"
}
}
For GET (the default webhook_method), the response_id, status, request_url, and status_message fields are appended to the webhook URL as query parameters. For example:
text
https://client.example.com/scrape-callback?response_id=0193f2a4-1b2c-7d3e-8f4a-5b6c7d8e9f0a&status=200&request_url=https%3A%2F%2Fexample.com&status_message=OK
Do not confuse the numeric status field of the notification with the string status in the method result.
Webhook delivery errors are logged but do not change the task result.
6.3. Retrieving the Result
bash
curl "$BASE_URL/tasks/result/$RESPONSE_ID"
The endpoint does not require authorization. On 200, it returns the stored result content as is, without an additional wrapper; it has no format parameter of its own. On 202, 404, 410, and 422, the service returns JSON regardless of the format specified when the task was created.
| Status | Meaning | Action |
|---|---|---|
200 OK |
task completed successfully | read the result |
202 Accepted |
task is still running | retry in a few seconds |
404 Not Found |
ID not found | check response_id |
410 Gone |
result deleted after retention period expired | create the task again |
422 Unprocessable Entity |
task finished with an error | read the error field in the JSON |
With 202:
json
{
"status": "pending"
}
With 404:
json
{
"status": "error",
"error": "Not found"
}
With 410 (the task record is kept, but the result body has been deleted):
json
{
"status": "expired",
"error": "Result expired"
}
With 422 (including when the original request had format: "raw"), JSON is returned with Content-Type: application/json; charset=utf-8:
json
{
"status": "error",
"error": "ScrapeParser: params.waitFor must be an object"
}
Besides x-debug, this endpoint also returns x-debug_request with the parameters of the found task, and, when available, the task data. Example for a scrape task:
http
x-debug_request: {"response_id":"0193f2a4-1b2c-7d3e-8f4a-5b6c7d8e9f0a","ip":"1.2.3.4","add_datetime":1747983214000,"finish_datetime":1747983220000,"price":0.0005,"task_type":1,"params":{"task_type":"scrape","url":"https://example.com"}}
| Field | Type | Description |
|---|---|---|
response_id |
string | requested task ID |
ip |
string | client IP address |
add_datetime |
number or null |
creation time, Unix milliseconds |
finish_datetime |
number or null |
completion time, Unix milliseconds |
price |
number | internal field; do not use for calculations, the task cost is determined by billing |
task_type |
number | numeric task type ID: 1 — scrape, 2 — google_search |
params |
object or string | task parameters |
The numeric task_type is used only in this header. In requests, the method is still passed as a string, for example "task_type": "google_search". For a completed task, the header contains the fields from the table; there is no time field in this structure.
Task fields are present on 200, 410, and 422. On 202 and 404, the header contains only response_id and ip. The x-debug_response header may be present on 200 and 422; it is absent on 202 and 410 (see section 4.1).
7. The scrape Method
The method fetches a single web page and returns HTML, Markdown, or a screenshot.
7.1. Parameters
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
task_type |
string | yes | — | always scrape |
url |
string | yes | — | full URL with http:// or https:// |
data_format |
string | no | raw |
raw, markdown, or screenshot |
format |
string | no | json |
json or raw |
waitFor |
JSON object | no | — | condition to wait for before capturing the result; 30 seconds maximum |
fullPage |
boolean | no | false |
full-page screenshot; applies to screenshot |
cdpurl |
string | no | — | your own Chrome via CDP |
An address without a scheme, such as example.com, is not accepted: the task ends with HTTP 422 and ScrapeParser: invalid URL "example.com". An address like //example.com also results in 422. The data: and about: schemes are not suitable for fetching a web page: the result may be empty with http_code: 0. An ftp:// address may end in a loading error. Use http:// or https:// to fetch a page.
The uule field is accepted in a scrape request, but its effect on the browser's geolocation is not documented. Do not rely on it as the only way to set the location for this method.
waitFor is passed as a JSON object:
json
{
"waitFor": {
"text": "Loaded"
}
}
A string, even one containing valid JSON, is not supported and causes the task to fail.
7.2. data_format
| Value | format: "json" |
format: "raw" |
|---|---|---|
raw |
JSON with status, http_code, headers, HTML in body |
HTML directly |
markdown |
JSON with Markdown in body |
Markdown directly |
screenshot |
JSON with Base64 PNG in body |
binary PNG |
A normal successful response with format: "json" has Content-Type: application/json; charset=utf-8. With format: "raw", the type depends on data_format:
data_format |
Content-Type of a successful raw response |
|---|---|
raw |
text/html |
markdown |
text/markdown |
screenshot |
image/png |
If the site is unavailable, a raw response changes its type to application/json; charset=utf-8 (see section 7.11). Check the actual header before processing the body.
With format: "json", the fetched page result contains:
| Field | Type | Description |
|---|---|---|
status |
string | success — no issues; warn — warning is present; error — the site is unavailable (see section 7.11) |
http_code |
integer | HTTP code of the target site; 0 on a network error |
headers |
object | response headers of the target site |
body |
string | HTML, Markdown, or Base64 PNG according to data_format |
warning |
string | only if the waitFor condition was not met within 30 seconds; in that case status: "warn" |
status: "success" means the method retrieved the page without issues and does not require HTTP 2xx from the site. For example, a site response of 404 is returned with HTTP 200 from the API, status: "success", and http_code: 404. Site 4xx codes, including 403 and 404, are not considered unavailability.
With format: "raw", the body contains HTML, Markdown, or binary PNG. The status, http_code, and warning metadata are passed in x-debug_response, and the site's headers are not available in this mode. If the site is unavailable, JSON is returned instead of the content (see section 7.11).
7.3. Get HTML as JSON
bash
curl -X POST "$BASE_URL/tasks/sync" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"task_type": "scrape",
"url": "https://example.com",
"data_format": "raw",
"format": "json"
}'
Response:
json
{
"status": "success",
"http_code": 200,
"headers": {
"content-type": "text/html; charset=utf-8"
},
"body": "<!DOCTYPE html>..."
}
7.4. Get HTML Directly
bash
curl -X POST "$BASE_URL/tasks/sync" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"task_type": "scrape",
"url": "https://example.com",
"data_format": "raw",
"format": "raw"
}'
The response begins with:
html
<!DOCTYPE html>
<html>
7.5. Get Markdown
JSON:
json
{
"task_type": "scrape",
"url": "https://example.com",
"data_format": "markdown",
"format": "json"
}
Direct:
json
{
"task_type": "scrape",
"url": "https://example.com",
"data_format": "markdown",
"format": "raw"
}
Markdown example:
md
# Example Domain
[More information](https://www.iana.org/domains/example)
7.6. Take a Viewport Screenshot
Base64 in JSON:
json
{
"task_type": "scrape",
"url": "https://example.com",
"data_format": "screenshot",
"format": "json",
"fullPage": false
}
Response:
json
{
"status": "success",
"http_code": 200,
"headers": {
"content-type": "text/html; charset=utf-8"
},
"body": "iVBORw0KGgoAAAANSUhEUg..."
}
The body field contains Base64 without the data:image/png;base64, prefix.
7.7. Get a Binary PNG
bash
curl -X POST "$BASE_URL/tasks/sync" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"task_type": "scrape",
"url": "https://example.com",
"data_format": "screenshot",
"format": "raw"
}' \
--output screenshot.png
The first eight bytes of the file:
text
89 50 4E 47 0D 0A 1A 0A
7.8. Take a Full-Page Screenshot
json
{
"task_type": "scrape",
"url": "https://example.com/long-page",
"data_format": "screenshot",
"format": "json",
"fullPage": true
}
For data_format: "raw", the fullPage parameter is effectively ignored and does not change the HTML result.
7.9. waitFor
The condition is specified as a JSON object with text, element, or state. For the CSS selector element, checkVisible defaults to false.
Wait for Text
json
{
"task_type": "scrape",
"url": "https://example.com/dynamic",
"waitFor": {"text": "Content loaded"}
}
The page is suitable for this test only if the string is absent from the initial HTML and is added later.
Wait for an Element
json
{
"task_type": "scrape",
"url": "https://example.com",
"waitFor": {"element": "#content", "checkVisible": false}
}
Wait for a Visible Element
json
{
"task_type": "scrape",
"url": "https://example.com",
"waitFor": {"element": "#content", "checkVisible": true}
}
Wait for Full Load
json
{
"task_type": "scrape",
"url": "https://example.com",
"waitFor": {"state": "load"}
}
Wait for DOM Construction
json
{
"task_type": "scrape",
"url": "https://example.com",
"waitFor": {"state": "domcontentloaded"}
}
Supported state values:
| Value | Description |
|---|---|
load |
the page and its dependent resources have loaded |
domcontentloaded |
the DOM is built, resources may still be loading |
Wait Limit and Warnings
waitFor waiting is limited to 30 seconds. Increasing the overall timeout does not extend this limit; timeout in /tasks/sync sets how long the HTTP request waits for the task to complete.
If the condition is not met within 30 seconds, the task does not fail: the fetched page is returned as is. The result then contains status: "warn" and warning:
json
{
"status": "warn",
"http_code": 200,
"headers": {
"content-type": "text/html; charset=utf-8"
},
"body": "<!DOCTYPE html>...",
"warning": "waitFor not met within 30s (element=#comments) — page returned as is"
}
If the condition is met or waitFor is not set, there is no warning field. With format: "raw", the warning is available in x-debug_response. The warning value uses an em dash — (U+2014) before page returned as is; in the header's JSON value it is transmitted as —. When handling warnings, do not compare the entire string against a hard-coded pattern.
waitFor Errors
An invalid value ends the task immediately, without retries and without contacting the target site. The error text comes with the ScrapeParser: … prefix:
| Error Text Without Prefix | Cause |
|---|---|
params.waitFor must be an object |
a string (including one with JSON inside), number, array, or true was passed |
params.waitFor must have "element", "text" or "state" |
the object has no known fields or element is empty |
params.waitFor.state must be one of: load, domcontentloaded |
unknown state value |
7.10. Your Own Chrome via cdpurl
json
{
"task_type": "scrape",
"url": "https://example.com/account",
"cdpurl": "wss://browser.example.com/devtools/browser/9b2c1f0a-..."
}
Requirements:
ws://orwss://is supported;- the browser must be reachable by the worker for the entire execution;
- the task uses that browser's cookies, sessions, profile, fingerprint, and proxy;
- do not publish the CDP URL: access to it effectively grants access to the browser session.
If cdpurl is not specified, the worker picks the browser. If it is specified but the connection fails after two attempts, the task ends with HTTP 422, with no further retries and without substituting the worker's own browser for yours. The error text starts with:
text
CDP connect failed (user cdpurl) after 2 attempts: Timeout 12000ms exceeded.
CDP connect failed (user cdpurl) after 2 attempts: WebSocket error: <endpoint> 401 Unauthorized deny_no_user
CDP connect failed (user cdpurl) after 2 attempts: WebSocket error: <endpoint> 500 Internal Server Error profile_locked
In the error text, the CDP address is replaced with <endpoint> so that credentials from the URL do not end up in the response or task history. Check that the browser is running and reachable from outside, that the link has not become stale after a Chrome restart, and that the profile is not in use by another connection.
7.11. Site Unavailability
On a network error (DNS, connection refused, TLS error, load timeout) or a 5xx response from the site, the task is considered successfully completed at the API level: the HTTP response code is 200, and the method result contains status: "error" and the reason:
json
{
"status": "error",
"http_code": 0,
"error": "net::ERR_NAME_NOT_RESOLVED at https://example.invalid",
"url": "https://example.invalid"
}
| Field | Description |
|---|---|
status |
error — the site is unavailable |
http_code |
0 on a network error, or the site's actual 5xx HTTP code |
error |
reason text |
url |
the address where the error occurred |
In this case the response is always JSON, including with format: "raw". On a network error, Content-Type is application/json; charset=utf-8 with both format: "raw" and format: "json"; the outer HTTP status is 200. The x-debug_response header duplicates status, http_code, and error; the url field is available only in the body. Do not save such a response as HTML or PNG without checking Content-Type and the result.
Site 4xx responses are returned as a normal result with the corresponding http_code. A site load timeout is different from the synchronous endpoint's HTTP 408: with 408, the task itself keeps running (see section 5.3).
8. The google_search Method
The method performs a Google query and returns structured results grouped by type, or a Markdown document. Authorization, synchronous and asynchronous execution, webhooks, and result retrieval follow the general rules in sections 2–6.
8.1. Parameters
All fields are passed flat in the JSON body, without a nested params.
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
task_type |
string | yes | — | always google_search |
url |
string | yes | — | full search URL; the domain must be google.com or one of its subdomains; the q parameter is required |
format |
json or raw |
no | json |
body format; with data_format: "markdown", raw returns the document directly. Without Markdown, raw returns the full JSON result (see section 8.9) |
data_format |
string | no | standard grouped results | markdown — a document; any other value, including json, — standard grouped results |
page_num or pages |
integer | no | 1 |
number of result pages, an integer from 1 to 10 |
gl |
string | no | — | two-letter country code; overrides gl inside url |
hl |
string | no | — | two-letter language code; overrides hl inside url |
with_ai |
boolean | no | true |
collect the AI answer on standard results |
with_preview |
boolean | no | false |
add previews to news, video, and books items |
with_html |
boolean | no | false |
add the HTML of each page; only with format: "json" |
saveScreenshot |
boolean | no | false |
add a screenshot of each page; only with format: "json" |
uule |
string | no | — | Canonical Name, lat,lon[,radius], or a ready-made Google value; radius in meters, default 200 |
To set the number of pages, use one of the fields: page_num or pages. The timeout parameter of the synchronous request is described in section 5.1.
8.2. Parameters Inside url
The method keeps the following Google parameters and removes all others, because an unknown parameter in the request increases the likelihood of being blocked. The uule geolocation is handled separately: it can be passed in the URL or in the request body (see section 8.7).
| Parameter | Example | Description |
|---|---|---|
q |
q=fastify+nodejs |
required search query |
oq |
oq=fastify |
original query before autocomplete |
hl |
hl=en |
results page language |
gl |
gl=us |
search country; if gl is in the URL and no geolocation is set explicitly, the API adds the country's uule. A gl field in the body alone does not trigger this (see section 8.7) |
start |
start=20 |
results offset; non-negative, a multiple of 10, no more than 90 |
tbm |
tbm=isch |
search type from the table below |
udm |
udm=2 |
search type from the table below |
tbs |
tbs=qdr:d |
Google filters, including time filters |
lr, cr |
lr=lang_en |
restrict documents by language and country |
safe |
safe=active |
SafeSearch |
pws |
pws=0 |
0 disables results personalization |
nfpr |
nfpr=1 |
search without typo autocorrection |
Spaces in URL parameter values are encoded as %20 or +.
Search Types
For standard web results, do not set tbm or udm. In rows with two options, one of them is enough.
| Results | Parameter | layout |
Main Group |
|---|---|---|---|
| Standard web results | no tbm or udm |
organic |
organic |
| Places | udm=1 or tbm=lcl |
places |
places |
| Images | udm=2 or tbm=isch |
images |
images |
| Videos | udm=7 or tbm=vid |
video |
video |
| Jobs | udm=8 |
jobs |
jobs |
| News | udm=12 or tbm=nws |
news |
news |
| Books | udm=36 or tbm=bks |
books |
books |
| Shorts | udm=39 |
shorts |
video |
| AI Mode | udm=50 |
ai_mode |
ai_overview |
Other tbm and udm values are not supported and cause the task to fail. In particular, udm=14, udm=18, and udm=56 are not on the allowed list. Shopping (udm=28, tbm=shop) is temporarily unavailable.
Results Depth
A results page contains up to 10 results. For multiple pages, set page_num from 1 to 10; the num and filter parameters are not supported.
start=20 starts collection from the third page. Page and position numbers are absolute in this case (see section 8.8).
json
{
"task_type": "google_search",
"url": "https://www.google.com/search?q=fastify+nodejs&hl=en&gl=us&start=20",
"page_num": 2,
"format": "json"
}
To get only the second page, specify start=10 without page_num:
json
{
"task_type": "google_search",
"url": "https://www.google.com/search?q=fastify+nodejs&hl=en&gl=us&start=10"
}
8.3. Regular Web Search
bash
curl -X POST "$BASE_URL/tasks/sync" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"task_type": "google_search",
"url": "https://www.google.com/search?q=fastify+nodejs&hl=en&gl=us",
"page_num": 1,
"format": "json",
"data_format": "json",
"with_html": false
}'
Example response structure with a single item:
json
{
"status": "success",
"results": 484000000,
"layout": "organic",
"task_param": {
"url": "https://www.google.com/search?q=fastify+nodejs&hl=en&gl=us",
"pages": 1
},
"real_location": "United States",
"groups": {
"organic": {
"count": 1,
"items": [
{
"url": "https://fastify.dev/",
"title": "Fastify",
"description": "Fast and low overhead web framework for Node.js",
"position": 1,
"page_position": {
"page": 1,
"position": 1
}
}
]
}
}
}
| Field | Type | Description |
|---|---|---|
status |
string | success — results collected; no_results — no items, this is not an error; error — an error, its text may be in the error field |
results |
number | number of results according to Google; -1 if the counter could not be read |
layout |
string | main results type from section 8.2; unknown if the layout was not recognized |
task_param |
object | parameters received by the handler, including geolocation added by the server |
real_location |
string or null |
region actually applied by Google; null if the label could not be read |
groups |
object | results by group; a regular group contains count and an items array, ai_overview has a separate structure (see section 8.8) |
error |
string | optional error text when status: "error" |
html |
array | HTML of each page; only with with_html: true and format: "json" |
screenshots |
array | page screenshots in Base64; only with saveScreenshot: true and format: "json" |
Organic items are in groups.organic.items. The organic group may be absent, for example in image search. A single response may contain groups of different types.
results: 0 and results: -1 mean different things: 0 means Google reported no results; -1 means the counter is unavailable. Jobs and AI Mode results have no counter, so they always return -1. Check for items using status and groups, not results alone. Google's result count may vary between sessions for the same query.
With with_html: false, the html field is not added. To get page HTML and screenshots along with the result:
json
{
"task_type": "google_search",
"url": "https://www.google.com/search?q=fastify+nodejs&hl=en&gl=us",
"format": "json",
"with_html": true,
"saveScreenshot": true
}
To skip collecting the AI answer on standard results, pass with_ai: false. Even with collection enabled, whether an AI Overview appears depends on whether Google showed one.
8.4. Image Search
json
{
"task_type": "google_search",
"url": "https://www.google.com/search?q=modern+furniture&hl=en&gl=us&tbm=isch",
"format": "json",
"data_format": "json"
}
Alternatively:
text
https://www.google.com/search?q=modern+furniture&hl=en&gl=us&udm=2
The main group is groups.images, and layout is images. The item's url field points to the source page, and origin_image_url to the original image. The preview is available in image_url or image_base64.
8.5. News
json
{
"task_type": "google_search",
"url": "https://www.google.com/search?q=technology&hl=en&gl=us&tbm=nws",
"format": "json",
"data_format": "json"
}
The main group is groups.news, and layout is news. You can use udm=12 instead of tbm=nws.
News from the last 24 hours, with previews:
json
{
"task_type": "google_search",
"url": "https://www.google.com/search?q=technology&hl=en&gl=us&udm=12&tbs=qdr:d",
"format": "json",
"with_preview": true
}
The time field contains the age as text, as Google displays it. Without with_preview, news items have no image_url or image_base64 fields.
8.6. Result in Markdown
With data_format: "markdown", the results are assembled into a single document.
In JSON:
json
{
"task_type": "google_search",
"url": "https://www.google.com/search?q=fastify+nodejs&hl=en&gl=us",
"format": "json",
"data_format": "markdown"
}
The response contains status, results, layout, task_param, real_location, and a body string with the document. There is no groups field in this mode. With with_html: true, page HTML is added to the JSON result in the html field. With saveScreenshot: true, a screenshots field is added to the Markdown JSON result; for a single page, it contains one Base64 PNG.
Directly:
bash
curl -X POST "$BASE_URL/tasks/sync" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"task_type": "google_search",
"url": "https://www.google.com/search?q=fastify+nodejs&hl=en&gl=us",
"format": "raw",
"data_format": "markdown"
}' \
--output search.md
When the document is retrieved successfully, the body is Markdown with Content-Type: text/markdown; the html and screenshots attachments are not included. Check the HTTP status and Content-Type before using the file: a service error is not a Markdown result.
The document consists of ## <group> sections with numbered items; images are embedded in the document as pictures. The ## Task information section and the yaml block may be absent. For task parameters, use task_param with format: "json"; with format: "raw", do not rely on them being present in the document text.
8.7. Geo-Targeting
Geolocation can be set via uule in the request body or inside url. The value from the URL is extracted automatically; there is no need to duplicate it in a separate field. If different values are passed in both places, no error is returned: the value from the URL takes precedence. Check the applied geolocation in task_param and real_location.
Canonical Name
json
{
"task_type": "google_search",
"url": "https://www.google.com/search?q=pizza&hl=en&gl=us",
"uule": "New York,New York,United States"
}
A Canonical Name is a location name from the Google Ads Geo Targets database. It must match the database entry exactly. The number of parts varies by city: for example, the canonical name for Paris has four parts.
text
Paris,Paris,Ile-de-France,France
New York,New York,United States
London,England,United Kingdom
Inside a URL, spaces are encoded as usual:
text
https://www.google.com/search?q=pizza&gl=us&uule=New+York,New+York,United+States
Coordinates
For coordinates, use a uule string in the lat,lon or lat,lon,radius format:
json
{
"task_type": "google_search",
"url": "https://www.google.com/search?q=pizza&hl=en&gl=us",
"uule": "40.7128,-74.0060,5000"
}
Latitude and longitude are specified in decimal degrees, and the radius in meters. If no radius is specified, 200 meters is used. In the example, the radius is 5 km; a value of 10 would mean 10 meters.
Ready-Made Google Value
google_search also accepts a ready-made uule value copied from the address bar of a Google results page. It is passed to Google unchanged and is not checked against the region directory, so it can specify a country or a region at any level.
| Prefix | Content |
|---|---|
w+ |
region name |
a+ |
latlng coordinates and radius in meters |
For example, a search with a ready-made value for the United Kingdom:
json
{
"task_type": "google_search",
"url": "https://www.google.com/search?q=coffee+shops&hl=en",
"uule": "w+CAIQICIOVW5pdGVkIEtpbmdkb20"
}
The ready-made value can also be passed inside url; the server extracts it automatically.
Geolocation by gl
If uule is not set explicitly, automatic addition is triggered by gl in the URL. Both gl=gb and gl=uk result in the same uule for the United Kingdom. For the ISO country code, use gb; the API also accepts uk. An explicitly set uule takes precedence over automatic addition by gl.
A gl field only in the request body does not trigger automatic uule addition. In that case, the resulting real_location may differ from the expected country. For a reproducible region, pass uule explicitly or put gl in the URL, and check real_location. Without gl or geolocation, Google chooses the region by IP; real_location: null means the region label could not be read.
8.8. Result Groups and Positions
Regular groups have the structure { "count": N, "items": [...] }. The ai_overview group is structured differently.
| Group | Item Fields |
|---|---|
organic |
url, title, description |
news |
url, title, description, source, time; image_url, image_base64 — with with_preview |
video |
url, title, description, is_shorts; image_url, image_base64 — with with_preview |
images |
url (source page), title, image_url, image_base64, origin_image_url, origin_image_width, origin_image_height |
books |
url, title, description; image_url, image_base64 — with with_preview |
ads |
url, title, advertiser, display_url |
places |
title, url, place_id, rating, reviews, category, price_level, address, distance, hours, review_quote, image_url, image_base64 |
jobs |
title, company, location, source, tags, url, job_id, image_url, image_base64 |
ai_overview |
an object with text and a sources array; a source contains url, label, and, if available, domain |
Standard results may include news, videos, places, and ads alongside organic results. In that case, layout remains organic. For example, the presence of groups.places does not by itself mean layout: "places".
Numbering
| Field | Description |
|---|---|
position |
running item number across the collected pages, taking start into account |
page_position |
an object with the absolute page number and the position on it, for example { "page": 3, "position": 1 } |
When multiple pages are collected, items are merged into a single list for the corresponding group. The start offset applies only to the group being paginated: for standard search, that is organic. News, video, and places blocks are numbered from one regardless of start.
Previews and Missing Data
For previews, either image_url (a link) or image_base64 (an image embedded by Google in the page) is filled in. Without with_preview, these fields are absent in news, video, and books.
In places, jobs, and books, some fields may be null if the data is not on the page. For example, this can happen with rating, number of reviews, price level, distance, opening hours, review quote, and place images, a job logo, or a book cover. Consider missing previews in books together with the with_preview setting.
null is not the same as zero or an empty string: rating: null means there is no rating. A special case is places.url: if the place has no ID in the Google Knowledge Graph, the field may be an empty string.
AI Overview
ai_overview.text contains Markdown. The sources field is an array of sources with addresses and labels. If a source link could not be resolved, url points to a Google redirect, and there is no domain field.
Google does not show an overview for every query. If there is no overview, the corresponding group is absent from the result. The with_ai parameter controls AI answer collection on standard results; the separate AI Mode is selected with udm=50.
8.9. The x-debug_response Header and format: "raw"
Task metadata is in x-debug. Search results service fields are available in x-debug_response with both format values, subject to the general conditions in section 4.1:
http
x-debug_response: {"status":"success","results":484000000,"layout":"organic","real_location":"United States"}
Items and groups are not included in the header. With format: "raw" without data_format: "markdown" (either without data_format or with data_format: "json"), the API returns the full JSON result: status, results, layout, task_param, real_location, and groups. Content-Type is application/json; charset=utf-8. To get Markdown directly, use format: "raw" together with data_format: "markdown" (see section 8.6); the body is then the document text.
8.10. Parameter Errors and Retries
Parameter errors end the task immediately with HTTP 422, without automatic retries. The error text has the GoogleParser: … prefix.
| Error Text Without Prefix | Cause |
|---|---|
params.url required |
the url field was not passed |
invalid URL "…" |
the value cannot be parsed as a URL, for example a search phrase or an address without a scheme was passed |
URL domain must be google.com, got "…" |
the domain is not google.com or one of its subdomains |
URL missing required parameter "q" |
the q search query is missing |
params.pages must be an integer 1..10 |
the number of pages is fractional, non-numeric, or outside the 1–10 range |
"start" must be a non-negative multiple of 10 |
the offset is negative or not a multiple of 10 |
"start" must be <= 90 |
the offset is greater than 90 |
| unsupported search type message | tbm or udm is not on the allowed list; the allowed values are listed in the error text |
For parameter errors, the method passes status: "error" and a description in x-debug_response. Header example:
http
x-debug_response: {"status":"error","error":"params.url required"}
Besides parameter errors, a task may end with 422 and the text max_restarts_exceeded: it could not be completed within the allotted number of attempts. In this case, repeat the request. Handle it as a task error, based on the HTTP status and the actual Content-Type.
The general 422 response formats are described in sections 5.4 and 6.3. When parsing the response, consider the HTTP status, Content-Type, and the error description in x-debug_response.
A CAPTCHA, a Google sign-in page, a dropped connection, and a timeout are not parameter errors: such tasks are retried automatically. The synchronous wait timeout, HTTP 408, is handled according to the general rules — save the response_id and request the result later, without creating a duplicate task.
9. Task History
bash
curl "$BASE_URL/task_history?task_type=scrape&from=2026-01-01&to=2026-12-31&limit=50&offset=0" \
-H "Authorization: Bearer $API_KEY"
| Parameter | Type | Default | Description |
|---|---|---|---|
task_type |
string | — | filter by method |
from |
DateTime | — | lower bound of add_datetime |
to |
DateTime | — | upper bound of add_datetime |
limit |
integer | 100 |
number of records |
offset |
integer | 0 |
offset |
Response:
json
[
{
"response_id": "0193f2a4-1b2c-7d3e-8f4a-5b6c7d8e9f0a",
"add_datetime": 1747983214000,
"finish_datetime": 1747983220000,
"task_type": "scrape",
"params": {
"task_type": "scrape",
"url": "https://example.com",
"format": "json"
},
"price": 0.0005,
"mime_type": "application/json",
"error": "",
"status": "done"
}
]
The params field contains the request parameters, including task_type and method fields; the email and password authorization fields are not stored in the task.
The history contains only completed tasks. Its status field is derived from the error field: an empty string means done, and error text means error. There is no pending value in the history. This status refers to task completion and does not replace the status inside a scrape or google_search result.
10. Common Errors
| HTTP Status | Cause |
|---|---|
400 Bad Request |
invalid parameters, unknown or disabled task_type, body larger than 10,000 bytes |
401 Unauthorized |
invalid or missing credentials, blocked account, or depleted balance (≤ 0); the reason is given in error |
402 Payment Required |
the balance is positive but insufficient for a new task, taking reserves into account; the task is not created |
404 Not Found |
the requested task ID was not found |
408 Request Timeout |
synchronous wait ended; the task keeps running, the service status is pending |
410 Gone |
result deleted after retention period |
422 Unprocessable Entity |
task finished with an error |
503 Service Unavailable |
acceptance of tasks of this type is temporarily suspended — Service overloaded, retry after Retry-After; on /health/ready — infrastructure not ready |
Example error in JSON:
json
{
"error": "Insufficient balance",
"status": "error"
}
With format: "raw", the API may return plain text only:
text
Insufficient balance
On 400, the JSON response also contains status: "error". Example error values:
params exceeds 10 000 bytes;Unknown task type: unknown_name;Task type temporarily disabled.
For GET /tasks/result/:response_id, waiting and error responses are always JSON (see section 6.3).
When a task is created, the balance is checked taking already reserved funds and the cost of the new task into account. On 402, the task is not created. The actual charge is made after the task completes successfully; the cost is determined by billing, and x-debug.price is not intended for calculations.
Service Overload
If workers cannot keep up with the queue, acceptance of tasks of a specific type via /tasks/request or /tasks/sync may be temporarily suspended. When creating a task, the API returns:
http
503 Service Unavailable
Retry-After: 60
json
{
"error": "Service overloaded",
"status": "error"
}
With format: "raw", the body contains Service overloaded with Content-Type: text/plain; with format: "json", the response contains JSON with Content-Type: application/json; charset=utf-8. The task is not created, and no balance is reserved. Retry the request after the number of seconds in Retry-After (60 in the example; use the actual value from the response). The restriction is lifted automatically as the queue is processed.
This response is different from the 503 of the /health/ready endpoint, which reports that an infrastructure dependency is unavailable.
Always check:
- the API's HTTP status;
- the
Content-Type; x-debug.status_code;x-debug.response_id;- the service
statusanderror, if the response was generated by the service; - for a
scraperesult —status,http_code, and, if present,warningorerror; - for
google_search—status,errorif present, andgroupsorbodyaccording todata_format; aresults: -1value alone does not mean there are no results.
With format: "raw", use the available x-debug_response fields, keeping in mind that the header is not guaranteed and its presence does not mean success. If the site is unavailable, read the JSON body.
11. Utility Endpoints
No authorization required.
GET /health
bash
curl "$BASE_URL/health"
json
{
"status": "ok",
"ts": 1747983214000
}
GET /health/ready
bash
curl "$BASE_URL/health/ready"
Successful response:
json
{
"status": "ready",
"deps": {
"redis_streams": true,
"redis_cache": true,
"clickhouse": true,
"s3": true
}
}
If a dependency is unavailable, the endpoint returns 503 and status: "degraded".
12. Integration Recommendations
- Use the synchronous endpoint for short tasks and the asynchronous one for long or bulk tasks. If a task may take longer than half a minute, the asynchronous mode is more reliable: you don't have to keep the connection open.
- When polling
GET /tasks/result/:response_id, retry every few seconds while you receive202. - On
408, do not immediately create a duplicate task: save theresponse_idand request the result later. - Do not treat an outer HTTP
200as confirmation that the target site also responded with200. - The task cost is determined by billing; do not use the internal
x-debug.pricefield for calculations. - For
scrape, check the result in this order: firststatus: "error"and theerrorfield, thenstatus: "warn"andwarning, thenhttp_code >= 400, and only after that processbody. An outer HTTP200does not cancel method or target site errors. - On
503 Service overloaded, retry task creation after the interval inRetry-After. - For PNGs with
format: "raw", checkContent-Typebefore saving: if the site is unavailable, a JSON error is returned instead of a PNG. - For PNGs with
format: "json", decode the Base64 from the top-levelbody. - Pass
waitForas a JSON object. For itstextfield, use a string that is absent from the original HTML and appears without user action. - Keep in mind the 30-second
waitForlimit: when it expires, the page is returned withstatus: "warn"and awarningfield. - For
google_searchwith standard structured results, readgroupsfrom the JSON. Withformat: "raw"without Markdown, the result also containsgroups; rely on the actualContent-Type. Withdata_format: "markdown", read the document frombodyin JSON mode, or get the text directly withformat: "raw". - Do not treat
google_search.status: "no_results"as an error; Google's result count inresultsis not the number of collected items and may be unavailable (-1). - For reproducible search results, set
uuleexplicitly or putglinside the URL; aglfield in the body alone does not trigger automaticuuleaddition. Check the actual region inreal_location. - Do not use external demo sites as permanent fixtures: their content and availability may change.
- Protect API keys, passwords, webhook URLs, and CDP URLs.