ELTRACT COLLECT

Public web data collection without a one-off scraping workflow.

ElTract Collect organizes public-web extraction around projects, sources and runs. You can analyse an authorized public source, compare detected extraction approaches, configure the fields you need, start a crawl when the collection runtime is available and review the results inside the same workspace.

Public-web collection · crawl runtime required for extraction

CAPABILITIES

What ElTract does here

01

Projects and source configuration

Create, edit, archive, restore and manage projects that define the public source and structured fields you want to collect.

02

Source analysis before a crawl

Inspect an authorized public source for extraction approaches, repeated records, pagination signals, detail pages and detected fields without starting a full crawl.

03

Tracked extraction runs

Start runs, view progress and events, inspect details and cancel active work through the product run model.

04

Structured results and review

Search, filter, sort and review project results, including source links and data-quality metadata.

05

CSV/XLSX export

Export authorized project results after collection instead of manually rebuilding the dataset from pages.

WORKFLOW

From input to working data

  1. 1. Define the project

    Choose a public source and the fields the resulting records should contain.

  2. 2. Analyse the source

    Review detected extraction approaches and source signals before starting the crawl.

  3. 3. Run collection

    The crawl worker uses Playwright and Redis to execute public-web extraction when that runtime is available.

  4. 4. Review the records

    Inspect, correct and export the structured results or move useful records into later ElTract workflows.

WHY ELTRACT

Keep the data connected to the work

Repeatable projects

Collection settings and run history stay attached to a project instead of living in an ad-hoc script.

Review before reuse

Results remain inspectable, searchable and reviewable before they feed another workflow.

A path beyond scraping

Collected records can continue into Data, Leads, Monitor and Web Memory instead of stopping at a download.

CURRENT BOUNDARIES

What not to assume

ElTract states availability directly so buyers do not have to infer capabilities from vague product language.

  • Collection targets public websites. ElTract does not currently store authenticated browser sessions for private or logged-in sites.
  • Actual public-web crawling requires the ElTract crawl worker, Playwright and Redis to be deployed and healthy.
  • The source-analysis browser preview also requires Playwright Chromium on the API host.

THE WEB, STRUCTURED.

Turn selected public web data into something your team can keep using.