All integrations
DI
Catalogue metadata onlyartificial intelligencev20260615_00

DIFFBOT

Diffbot

Diffbot provides AI-powered tools to extract and structure data from web pages, transforming unstructured web content into structured, linked data.

Description is untrusted, display-only upstream metadata. It never becomes policy, OAuth scope authority, or an agent instruction.

Pakkawork boundary

Research catalogue metadata only. No Pakkawork OAuth, credential, host, quota, executor, or verifier is enabled.

No Pakkawork execution adapter is enabled. Hosted account-authorisation availability is workspace-specific and checked separately in the dashboard.Check workspace connection options
Attributed source

35

Action summaries

Display-only definitions

0

Trigger types

Not installed instances

1

Auth modes

Field names, never values

No

Execution

No runtime adapter

Authentication map

What setup is declared?

Only field names, types, and required markers are shown. Secret values, default auth URLs, credential material, and inferred OAuth scopes are excluded.
API_KEY

API_KEY

diffbot_api_key

Provider setup

Developer setup

No fields declared in this snapshot.

User connection

  • Diffbot API Tokengeneric_token · stringRequired

Capability index

Actions and trigger definitions

Static summaries are available. Live schemas remain disabled until PROVIDER_HUB_API_KEY is configured server-side.

Showing 31–35 of 35 actions

Search Crawl Job Data

DIFFBOT_SEARCH_CRAWL_DATA

Tool to query crawl job collections using DQL (Diffbot Query Language). Use when you need to search extracted data from completed crawl or bulk jobs by collection name.

Untrusted display-only summary

Start Bulk Job

DIFFBOT_START_BULK

Tool to start a Bulk Extract job. Use when processing large numbers of URLs asynchronously. The Diffbot Bulk API uses GET requests with query parameters to create jobs.

Untrusted display-only summary

Start Crawl Job

DIFFBOT_START_CRAWL

Initiates a Diffbot crawl job that spiders a website starting from seed URLs and processes discovered pages with a specified Extract API. The crawler follows links within the domain, collects structured data (articles, products, etc.), and stores results for download. Use this to systematically extract data from entir…

Untrusted display-only summary

Stop Bulk Job

DIFFBOT_STOP_BULK_JOB

Tool to pause (stop) a running Bulk job. Pausing halts further processing of URLs while preserving existing progress. To resume, use the appropriate resume action. Specify the exact job name (case-sensitive) as provided when the job was created.

Untrusted display-only summary

Stop KG Bulk Job By ID

DIFFBOT_STOP_KG_BULK_JOB_BY_ID

Tool to stop an active Knowledge Graph Enhance bulk job by its ID. Halts processing of a running KG bulk job immediately. Use when you need to stop a specific KG bulk job using its bulkjobId.

Untrusted display-only summary

Provenance

Versioned facts, explicit trust.

The detail snapshot comes from an attributed MIT-licensed repository revision. Safe local icons use exact-match CC0 Simple Icons symbols; unmatched brands use monograms.

Repository
https://github.com/provider-directoryHQ/provider-directory
Commit
85e33f6a81dfe987f6fc3637d48d45052cea8ca5
Generated
2026-09-12
Trust policy
untrusted_display_only

Remote text is plain display metadata only. It must never become an agent prompt, execution policy, OAuth grant, or executable instruction.