Skip to main content

Crawl

Crawl — see the properties below.

This node is also exposed as an AI Agent tool.

This node has 1 input port and 1 output port.

Common Properties

Every node shares these.

  • Name — the node's display name on the canvas.
  • Color — the node's colour on the canvas.
  • Delay Before (sec) — wait this long before the node runs.
  • Delay After (sec) — wait this long after the node runs.
  • Continue On Error — carry on instead of failing the flow. Defaults to false.
info

When Continue On Error is true the error is not raised at all, so a Catch node will not see it either.

Inputs

PropertyFieldDescription
Client IDinClientIDClient ID from Connect node (optional if credentials provided)
URLinURLRoot URL to start crawling from

Outputs

PropertyFieldDescription
ResultsoutResultsArray of crawled pages with URL and content

Options

PropertyFieldDescription
Allow ExternaloptAllowExternalInclude links to external domains
CredentialsoptCredentialsTavily API key (optional if Connect node is used)
Extract DepthoptExtractDepthAdvanced extracts tables and embedded content One of: basic, advanced.
FormatoptFormatOutput format for extracted content One of: markdown, text.
LimitoptLimitTotal maximum number of pages to crawl
Max BreadthoptMaxBreadthMaximum number of links to follow per level (1-500)
Max DepthoptMaxDepthHow many levels deep to crawl from the root URL (1-5)
Timeout (seconds)optTimeoutMaximum time in seconds to wait

Requirements

  • A Client ID from this package's Connect node, unless you set credentials directly on this node.
  • Tavily API Key — API key for Tavily AI search
    • value — API Key (password, required). Your Tavily API key

Store these in a Vault and reference the vault item from the node, rather than typing the secret into the property.

Connect · Disconnect · Search · Extract · Map · Toolkit

See also