Tellagent
Collect from a web page
The Collect from a web page block brings what a page holds into a project's Data Management: the page's text, filed as a page, or the documents the page links to. It is a block in Tellagent's AI Workflows, so it runs when you press Run, or on a schedule.
Set it up
- In an AI workflow, click Collect from a web page under Sources in the Blocks list, and connect it after a trigger.
- Enter the Web page address.
- Under Collect, choose The page's text, as a page or The files it links to.
- Under File into project, choose the project, then a Data Management folder. Folders are listed by their full path. The project's default folder files them in the project's Uncategorised folder.
- If you like, say What to look for. It is handed to the steps after this one, with what was collected; it does not change what is collected.
- Press Run.
Anyone who can open the project can add the block and run the workflow. AI Workflows
What it collects
| The page's text | Filed as a page named after the page's title and the day it was collected. Run it again on the same day and that page is left as it is, and the run says so. |
|---|---|
| The files it links to | The PDF, Word, Excel, PowerPoint, CSV, text and RTF files the page links to: up to 12 a run, skipping any over 20 MB. Each is read like an upload, so Chat & Research and Odin can answer from it. What the AI can read |
A file already in the folder under the same name is left as it is, and the run says so: delete it and run again to replace it. A run stops early if the account's storage is full, and says so. A run you watch spends up to about 15 seconds collecting files; a scheduled or background run, up to a minute.
The steps after it get what it collected: an Agent / LLM after it can summarise the page, or list the documents it filed, and an Output to project can save that.
It respects the site
The site's robots.txt is checked first, and a page it disallows is not read: the step fails and says why. If a site blocks a plain fetch, the page is read with a headless browser. That costs a small amount, shown in Observability as Tellagent — reading a blocked page with a headless browser (Collect from a web page).
Crawls you already had
Each crawl in a project is a paused workflow in that project's Tellagent, named Crawl: and the crawl's name, with this block set up as the crawl was. Open it and press Run to collect again, or set it to Active with a schedule. The page text a crawl collected is a page in Data Management, in a folder named after the crawl; the documents it collected are in the folder it filed them in. A chart made from a crawl keeps its data. Data Visualizations
Signed in to VeraGen? Press Guide me on any screen for a walkthrough, or ask Odin.