ELTRACT COLLECT
Public web data collection without a one-off scraping workflow.
ElTract Collect organizes public-web extraction around projects, sources and runs. You can analyse an authorized public source, compare detected extraction approaches, configure the fields you need, start a crawl when the collection runtime is available and review the results inside the same workspace.
CAPABILITIES
What ElTract does here
Projects and source configuration
Create, edit, archive, restore and manage projects that define the public source and structured fields you want to collect.
Source analysis before a crawl
Inspect an authorized public source for extraction approaches, repeated records, pagination signals, detail pages and detected fields without starting a full crawl.
Tracked extraction runs
Start runs, view progress and events, inspect details and cancel active work through the product run model.
Structured results and review
Search, filter, sort and review project results, including source links and data-quality metadata.
CSV/XLSX export
Export authorized project results after collection instead of manually rebuilding the dataset from pages.
WORKFLOW
From input to working data
- 1. Define the project
Choose a public source and the fields the resulting records should contain.
- 2. Analyse the source
Review detected extraction approaches and source signals before starting the crawl.
- 3. Run collection
The crawl worker uses Playwright and Redis to execute public-web extraction when that runtime is available.
- 4. Review the records
Inspect, correct and export the structured results or move useful records into later ElTract workflows.
WHY ELTRACT
Keep the data connected to the work
Repeatable projects
Collection settings and run history stay attached to a project instead of living in an ad-hoc script.
Review before reuse
Results remain inspectable, searchable and reviewable before they feed another workflow.
A path beyond scraping
Collected records can continue into Data, Leads, Monitor and Web Memory instead of stopping at a download.
CURRENT BOUNDARIES
What not to assume
ElTract states availability directly so buyers do not have to infer capabilities from vague product language.
- Collection targets public websites. ElTract does not currently store authenticated browser sessions for private or logged-in sites.
- Actual public-web crawling requires the ElTract crawl worker, Playwright and Redis to be deployed and healthy.
- The source-analysis browser preview also requires Playwright Chromium on the API host.
THE WEB, STRUCTURED.