A Fortune Global 500 electronics manufacturer
Runs web data programs across its markets on Import.io.
Train on it. Ground on it. Compete on it. The web as infrastructure for technology companies.
Import.io collects public pages, documentation, pricing and changelogs for training, retrieval and agent workflows. AI and technology teams receive structured data with provenance through the extraction API, Web Scraper MCP or a managed feed.
| Company | Page | Change | Detected |
|---|---|---|---|
| Acme Cloud | Pricing | Pro plan $49 → $59 per seat | 16:14 |
| Acme Cloud | Pricing | SSO moved from Pro to Business | 16:14 |
| Vectorly | Changelog | Released: batch embeddings API | 15:02 |
| Northwind AI | Docs | New endpoint: /v2/agents | 14:47 |
| Relaywise | Integrations | Listed in 3 new app marketplaces | 11:20 |
| Vectorly | Careers | +12 roles: GPU infrastructure | 09:05 |
Production examples from organisations using Import.io today and in recent years.
Runs web data programs across its markets on Import.io.
Has fed its models with structured web data from Import.io.
Collects the public web data behind its API on Import.io.
For a technology company the web does three jobs at once. It is training data for your models, the runtime your agents act in, and the market where competitors change pricing and ship features every week. Each one fails differently. Import.io is built for all three.
Corpora fail on what you can’t see: duplicates, boilerplate, missing licence signals and pages that asked not to be used. Agents fail on page text that looks like data and isn’t, and on sites that block them. Competitive intelligence fails on noise — A/B tests and personalised pricing pages that change for no one but the crawler.
Import.io answers each with the same discipline: provenance on every document, typed output against a schema, and change detection that separates a real price change from a test.
Six programs we run most often in this industry.
Related guidance: AI and LLM data workflows, MCP, Firecrawl and Playwright comparison, Import.io and Firecrawl comparison, training data.
Domain corpora built to spec — deduplicated, quality-filtered, licence- and opt-out-aware — plus held-out evaluation sets.
corporaSearch, fetch and extract as tools for Claude, Cursor and any MCP client, returning typed records rather than page text.
agentsCompetitor pricing pages diffed daily or faster: plan prices, seat minimums, feature moves between tiers.
pricingNew endpoints, deprecations and launches detected as they publish, not when the blog post lands.
productApp-store and integration-directory listings, ratings and reviews across the ecosystems you sell into.
ecosystemsRoles, locations and skills from careers pages, plus community questions and sentiment on public forums.
signals| field | type | example |
|---|---|---|
| doc_id | str | sha256:1a9e… |
| url | url | docs.example.com/v2/agents |
| domain | str | docs.example.com |
| lang | str | en |
| text | str | main content, boilerplate removed |
| tokens | int | 2,418 |
| licence_signal | str | CC-BY-4.0 |
| robots | enum | allowed |
| change_type | enum | added_endpoint |
| fetched_at | ts | 2026-09-25T14:47:10Z |
What breaks when this is done with scripts, and how Import.io handles it.
Near-duplicate removal, quality filtering and language identification, with licence signals, robots status and a content hash recorded for every document.
Extraction to a schema you define, so agents act on typed fields, not on paragraphs that happen to contain a price.
Diffs are checked across repeated captures and regions, so A/B tests and personalisation don’t show up as competitor moves.
Rendering, challenges and regional egress are handled server-side, and blocked calls are not billed.
A clear operating plan before collection begins.
Agree the public sites, pages and coverage the program needs.
Define the entities, fields and validation rules for each structured record.
Set the refresh schedule and choose API or managed-feed delivery.
Same capture engine underneath each one.
Model builders who need corpora, retrieval feeds or evaluation sets built to spec.
learn more →Agent teams who want structured web access in Claude, Cursor or their own MCP client, 10,000 calls free.
learn more →Product marketing and CI teams who want to set up their own page monitors and extractors.
learn more →Collection honours robots.txt and AI-crawler rules, records licence signals per document, and detects and removes personal data. What a corpus may be used for is agreed per program.
Straight answers.
Yes: domains, languages, date ranges, exclusions, deduplication and quality thresholds, delivered with provenance and a quality report per build.
Through the Web Scraper MCP server or the API. Agents call extraction tools and get typed records back; the first 10,000 successful calls are free.
Changes are confirmed across repeated captures and regions before they are reported, so a test variant shown to one visitor doesn’t become a price change.
Yes. Robots.txt and AI-crawler rules are honoured, and licence signals are recorded per document so your use can follow them.
Yes. The self-service platform lets your team set up page monitors and extractors, from the free trial upwards.