One of the largest US health-information publishers
Runs its public web data collection on the Import.io platform.
The public side of healthcare, structured. Trials, labels, prices and providers. Never patients.
Import.io structures public healthcare data from trial registries, regulator notices, product labels, pharmacy prices, transparency files and provider directories. Research, regulatory and commercial teams receive current, non-patient records through an API or managed feed.
| Event | Product or entity | Source | Detected |
|---|---|---|---|
| Trial status → recruiting | NCT0612XXXX · Phase 3 | trial registry | Mon 08:12 |
| Label update | New warning added to label | regulator | Mon 14:40 |
| Generic entry | List price −38% vs brand | pharmacy sites | Tue 06:05 |
| Recall · Class II | Infusion set, lot 24-118 | regulator | Tue 11:30 |
| Network change | 214 oncologists added | provider directory | Wed 09:02 |
| Price transparency | Negotiated rates updated | payer files | Thu 04:00 |
Production examples from organisations using Import.io today and in recent years.
Runs its public web data collection on the Import.io platform.
Banks, public-sector data companies and health publishers run on the same governed platform, with PII detection, audit trails and data processing agreements.
Processed on the same engine that healthcare programs run on.
Healthcare publishes an enormous amount of public data — trial registries, regulator notices, product labels, pharmacy prices, provider directories and price transparency files so large they are measured in terabytes. Almost none of it arrives in a form you can query. Import.io turns it into structured, current data, scoped strictly to public, non-patient information.
Much of healthcare’s most valuable public data exists because regulation requires it: trial registrations, safety communications, product labels, price transparency files. It is published to comply, not to be used — scattered across registries, PDFs and machine-readable files too large to open.
Commercial, market access and medical teams need it structured and current: which trials just started recruiting, which labels changed, where a generic undercut the brand, which providers joined a network. That is an extraction problem, and it is ours.
Six programs we run most often in this industry.
Related guidance: data governance, data quality, data schemas, managed web data.
Trials by indication, phase, status, sponsor and site, with status changes as events.
trialsApprovals, safety communications, recalls and label revisions across regulators, parsed from their documents.
regulatoryRetail pharmacy prices, generic entry and device distributor pricing across markets.
pricingNegotiated rates from hospital and payer machine-readable files, filtered to the codes and markets you need.
transparencyProviders, specialties, locations and network participation, kept current.
providersOTC and supplement prices, content, reviews and MAP across retailers and marketplaces.
consumer health| field | type | example |
|---|---|---|
| entity_type | enum | trial |
| identifier | str | NCT / NDC / NPI / UDI |
| name | str | Phase 3 study, indication X |
| event_type | enum | status_change |
| value | str | not yet recruiting → recruiting |
| jurisdiction | str | US |
| source | str | trial registry |
| document_url | url | registry record |
| effective_date | date | 2026-09-22 |
| retrieved_at | ts | 2026-09-22T08:12:40Z |
What breaks when this is done with scripts, and how Import.io handles it.
Labels, notices and filings arrive as PDFs; layout-aware parsing keeps tables and sections intact and cites the page.
Price transparency files run to gigabytes each; they are streamed and filtered to the codes, plans and markets you need.
Trials, drugs, devices and providers are mapped to their standard identifiers so sources join cleanly.
Programs are scoped to public, non-patient data; provider data is limited to professional information published for that purpose.
A clear operating plan before collection begins.
Agree the public sites, pages and coverage the program needs.
Define the entities, fields and validation rules for each structured record.
Set the refresh schedule and choose API or managed-feed delivery.
Same capture engine underneath each one.
Commercial, market access and medical teams who want trials, labels and prices delivered as structured events, under an SLA.
learn more →Consumer health and OTC brands who need prices, availability and MAP evidence across retailers.
learn more →Analysts who want to build their own registry or catalogue extractors.
learn more →Programs are scoped to public, non-patient information. We do not collect protected health information or patient-level data; provider information is limited to professional details published for that purpose.
Straight answers.
No. Programs are scoped to public, non-patient information such as registries, labels, prices and provider directories.
Yes. The files are streamed and filtered to the billing codes, plans and markets you specify, and delivered as structured tables.
Registries are covered per program, in the US and internationally; each trial keeps its registry identifiers so records join across sources.
Yes. Regulator sites and label repositories are monitored, and changes are delivered as events with the section that changed.
Yes, through Aperture or a managed feed: prices, sellers and dated evidence across retailers and marketplaces.