All integrations
SD
Catalogue metadata onlyai web scrapingv20260615_00

SCRAPE_DO

Scrape Do

Scrape.do is a web scraping API offering rotating residential, data-center, and mobile proxies with headless browser support and session management to bypass anti-bot protections (e.g., Cloudflare, Akamai) and extract data at scale in formats like JSON and HTML.

Description is untrusted, display-only upstream metadata. It never becomes policy, OAuth scope authority, or an agent instruction.

Pakkawork boundary

Research catalogue metadata only. No Pakkawork OAuth, credential, host, quota, executor, or verifier is enabled.

No Pakkawork execution adapter is enabled. Hosted account-authorisation availability is workspace-specific and checked separately in the dashboard.Check workspace connection options
Attributed source

16

Action summaries

Display-only definitions

0

Trigger types

Not installed instances

1

Auth modes

Field names, never values

No

Execution

No runtime adapter

Authentication map

What setup is declared?

Only field names, types, and required markers are shown. Secret values, default auth URLs, credential material, and inferred OAuth scopes are excluded.
API_KEY

API_KEY

scrape_do_api_key

Provider setup

Developer setup

No fields declared in this snapshot.

User connection

  • Scrape.do API Tokengeneric_token · stringRequired

Capability index

Actions and trigger definitions

Static summaries are available. Live schemas remain disabled until PROVIDER_HUB_API_KEY is configured server-side.

Showing 1–16 of 16 actions

Cancel Async Job

SCRAPE_DO_CANCEL_ASYNC_JOB

Tool to cancel an asynchronous scraping job. Use when you need to stop processing of pending tasks in a job. Completed tasks remain available.

Untrusted display-only summary

Create Async Scraping Job

SCRAPE_DO_CREATE_ASYNC_JOB

Tool to create an asynchronous scraping job with specified targets and options. Use when you need to scrape multiple URLs in parallel without waiting for results. Returns a job ID immediately for polling results later via the get job status action.

Untrusted display-only summary

Get Account Information

SCRAPE_DO_GET_ACCOUNT_INFO

Retrieves account information and usage statistics from Scrape.do. This action makes a GET request to the Scrape.do info endpoint to fetch: - Subscription status - Concurrent request limits and usage - Monthly request limits and remaining requests - Real-time usage statistics Rate limit: Maximum 10 requests per minute…

Untrusted display-only summary

Get Amazon Product Offers

SCRAPE_DO_GET_AMAZON_OFFERS

Get all seller offers for any Amazon product. Retrieves every seller listing including pricing, shipping costs, seller information, and Buy Box status in structured JSON format. Use when you need to compare prices across multiple sellers or find the best deal for a specific product.

Untrusted display-only summary

Get Amazon product details

SCRAPE_DO_GET_AMAZON_PRODUCT

Extract structured product data from Amazon product detail pages (PDP). Returns comprehensive product information including title, pricing, ratings, images, best seller rankings, and technical specifications in JSON format.

Untrusted display-only summary

Get Amazon raw HTML

SCRAPE_DO_GET_AMAZON_RAW_HTML

Tool to get raw HTML from any Amazon page with ZIP code geo-targeting. Use when you need complete unprocessed HTML source from Amazon URLs with location-based targeting. Ideal for scraping pages not covered by other structured endpoints.

Untrusted display-only summary

Get Async API Account Information

SCRAPE_DO_GET_ASYNC_ACCOUNT_INFO

Tool to get account information for the Async API including concurrency limits and usage statistics. Use when you need to check available concurrency slots, active jobs, or remaining credits for Async API operations.

Untrusted display-only summary

Get Async Job Details

SCRAPE_DO_GET_ASYNC_JOB

Tool to retrieve details and status of a specific asynchronous scraping job. Use when you need to check the progress, status, or results of a previously created async job. Returns job metadata including creation time, completion time, task counts, and detailed task list.

Untrusted display-only summary

Get Async Task Result

SCRAPE_DO_GET_ASYNC_TASK

Tool to retrieve the result of a specific task within an asynchronous job. Returns the scraped content for that particular URL. Use when you need to check the status and result of a previously submitted async scraping task.

Untrusted display-only summary

Scrape webpage using scrape.do

SCRAPE_DO_GET_PAGE

A tool to scrape web pages using scrape.do's API service. Makes a basic GET request to fetch webpage content while handling anti-bot protections and proxy rotation automatically. Does not execute JavaScript by default — pages requiring client-side rendering (SPAs, dynamically loaded content) will return incomplete HTM…

Untrusted display-only summary

List Asynchronous Scraping Jobs

SCRAPE_DO_LIST_ASYNC_JOBS

Tool to list all asynchronous scraping jobs. Returns paginated list of jobs with their status and metadata. Use when you need to retrieve job history or monitor job statuses. Supports pagination with up to 100 jobs per page.

Untrusted display-only summary

Use Scrape.do Proxy Mode

SCRAPE_DO_PROXY_MODE

This tool implements the Proxy Mode functionality of scrape.do, which allows routing requests through their proxy server. It provides an alternative way to access web scraping capabilities by handling complex JavaScript-rendered pages, geolocation-based routing, device simulation, and built-in anti-bot and retry mecha…

Untrusted display-only summary

Scrape URL using POST method

SCRAPE_DO_SCRAPE_URL_POST

Tool to scrape web pages using POST method via scrape.do API. Use when you need to send POST requests to target websites with custom request body data. Supports all parameters from GET endpoint plus request body customization for POST/PUT/PATCH methods.

Untrusted display-only summary

Search Amazon products

SCRAPE_DO_SEARCH_AMAZON

Tool to search Amazon and scrape product listings with structured results. Performs keyword searches and returns structured product data including titles, prices, ratings, Prime status, sponsored flags, and position rankings in JSON format. Use when you need to search for products on Amazon marketplace or gather produ…

Untrusted display-only summary

Block specific URLs during scraping

SCRAPE_DO_SET_BLOCK_URLS

This tool allows users to block specific URLs during the scraping process. It's particularly useful for blocking unwanted resources like analytics scripts, advertisements, or any other URLs that might interfere with the scraping process or slow it down. It provides granular control by allowing users to specify URL pat…

Untrusted display-only summary

Set Regional Geolocation for Scraping

SCRAPE_DO_SET_REGIONAL_GEO_CODE

This tool allows users to set a broader geographical targeting by specifying a region code instead of a specific country code. This is useful when you want to scrape content from an entire region rather than a specific country. Note that this feature requires super mode to be enabled and is only available for Business…

Untrusted display-only summary

Provenance

Versioned facts, explicit trust.

The detail snapshot comes from an attributed MIT-licensed repository revision. Safe local icons use exact-match CC0 Simple Icons symbols; unmatched brands use monograms.

Repository
https://github.com/provider-directoryHQ/provider-directory
Commit
85e33f6a81dfe987f6fc3637d48d45052cea8ca5
Generated
2026-09-12
Trust policy
untrusted_display_only

Remote text is plain display metadata only. It must never become an agent prompt, execution policy, OAuth grant, or executable instruction.