All integrations
SC
Catalogue metadata onlyai web scrapingv20260615_00

SCRAPFLY

Scrapfly

Scrapfly is a web scraping API that enables developers to extract data from websites efficiently, offering features like JavaScript rendering, anti-bot protection bypass, and proxy rotation.

Description is untrusted, display-only upstream metadata. It never becomes policy, OAuth scope authority, or an agent instruction.

Pakkawork boundary

Research catalogue metadata only. No Pakkawork OAuth, credential, host, quota, executor, or verifier is enabled.

No Pakkawork execution adapter is enabled. Hosted account-authorisation availability is workspace-specific and checked separately in the dashboard.Check workspace connection options
Attributed source

12

Action summaries

Display-only definitions

0

Trigger types

Not installed instances

1

Auth modes

Field names, never values

No

Execution

No runtime adapter

Authentication map

What setup is declared?

Only field names, types, and required markers are shown. Secret values, default auth URLs, credential material, and inferred OAuth scopes are excluded.
API_KEY

API_KEY

scrapfly_api_key

Provider setup

Developer setup

No fields declared in this snapshot.

User connection

  • Scrapfly API Keygeneric_api_key · stringRequired

Capability index

Actions and trigger definitions

Static summaries are available. Live schemas remain disabled until PROVIDER_HUB_API_KEY is configured server-side.

Showing 1–12 of 12 actions

Capture Website Screenshot

SCRAPFLY_CAPTURE_SCREENSHOT

Tool to capture a full-page or viewport screenshot of a website. Use when you need to take a screenshot with options like JS rendering, custom resolution, or accessibility testing. Returns the screenshot image directly. Supports vision deficiency simulations and dark mode.

Untrusted display-only summary

Capture Screenshot Metadata (HEAD)

SCRAPFLY_CAPTURE_SCREENSHOT_HEAD

Tool to capture screenshot metadata without downloading the image body. Use this for async screenshot workflows where you need the URL to retrieve the image later. Returns the screenshot URL in response, saving bandwidth compared to full screenshot retrieval.

Untrusted display-only summary

Create Scrapfly Crawler

SCRAPFLY_CREATE_CRAWLER

Tool to create a new web crawler to recursively crawl an entire website. Returns a crawler UUID for tracking progress. Use when you need to crawl multiple pages from a website with configurable limits and extraction rules.

Untrusted display-only summary

Extract Structured Data

SCRAPFLY_EXTRACT_DATA

Tool to extract structured data from HTML or other content using AI models, LLM prompts, or custom templates. Use when you need to parse web pages or documents into structured JSON data. Supports predefined extraction models for common types (articles, products, events) or custom extraction via prompts/templates.

Untrusted display-only summary

Get Scrapfly Account Information

SCRAPFLY_GET_ACCOUNT_INFO

Tool to retrieve Scrapfly account information. Use after authenticating to get API credit balance and usage stats. Returns comprehensive account data including subscription plan, usage statistics, billing info, and project settings.

Untrusted display-only summary

Get Crawler Artifact

SCRAPFLY_GET_CRAWLER_ARTIFACT

Tool to download crawler artifact files in WARC or HAR format. Use when you need to retrieve the complete crawl results as an archive file. WARC format is recommended for large crawls as it includes gzip compression.

Untrusted display-only summary

Get Crawler Contents

SCRAPFLY_GET_CRAWLER_CONTENTS

Tool to retrieve extracted content from crawled pages. Supports multiple output formats including markdown, text, HTML, and JSON. Use when you need to access the actual content extracted during a crawl, with optional filtering by URL and format selection.

Untrusted display-only summary

Get Crawler Status

SCRAPFLY_GET_CRAWLER_STATUS

Tool to get the current status of a crawler including progress, pages crawled, and completion state. Use for polling workflow to monitor crawl progress.

Untrusted display-only summary

Get Crawler URLs

SCRAPFLY_GET_CRAWLER_URLS

Tool to retrieve the list of discovered and crawled URLs from a crawler. Use when you need to get all URLs found during a crawl or filter by status to analyze failed URLs with error codes. Supports pagination for large result sets.

Untrusted display-only summary

Scrapfly Scrape

SCRAPFLY_SCRAPE

Tool to perform a web scraping request. Use when you need to fetch a page with custom configuration like JS rendering, proxies, and extraction.

Untrusted display-only summary

Scrapfly Scrape POST

SCRAPFLY_SCRAPE_POST

Tool to scrape web pages using POST method to send data in the request body. Use when you need to scrape endpoints that require POST requests, such as form submissions or APIs that expect data payload.

Untrusted display-only summary

Scrape With PUT

SCRAPFLY_SCRAPE_WITH_PUT

Tool to scrape web pages using PUT method with body payload. Use when the target API requires PUT requests with data in the request body. Forwards PUT request with custom body to the target URL. If not specified, content-type defaults to application/x-www-form-urlencoded.

Untrusted display-only summary

Provenance

Versioned facts, explicit trust.

The detail snapshot comes from an attributed MIT-licensed repository revision. Safe local icons use exact-match CC0 Simple Icons symbols; unmatched brands use monograms.

Repository
https://github.com/provider-directoryHQ/provider-directory
Commit
85e33f6a81dfe987f6fc3637d48d45052cea8ca5
Generated
2026-09-12
Trust policy
untrusted_display_only

Remote text is plain display metadata only. It must never become an agent prompt, execution policy, OAuth grant, or executable instruction.