Media & Publishing web data

Every story, who ran it first, and where it ran next. Including where your content turns up without a licence.

Import.io collects public articles from publisher sites, news sitemaps, wire pages and web archives, then clusters them into stories. Newsroom, monitoring, research and rights teams receive timestamped records through managed feeds or on-demand MCP access.

Story tracker · “Harbor Bridge closure” · last 48 hclustered · syndication detected · sample
214articles
37outlets
9originals
3unlicensed reuse matches
OutletPublishedTypeWordsNote
Riverton Ledger03 Mar 06:12original1,240first to publish
Coast Dispatch03 Mar 07:40original1,680funding angle
Wire service03 Mar 08:05wire61042 pickups
Regional aggregator03 Mar 09:10syndicated610wire copy
Content farm03 Mar 11:32reuse1,19096% overlap with Ledger
Coast Dispatch04 Mar 10:02update420editorial
live events
60 schecks on breaking-news pages
190locales and languages
Thousandsof outlets, structured into stories
2012in production since

Publishers already on it.

Production examples from organisations using Import.io today and in recent years.

long-standing customer

One of the world’s largest business-news publishers

A long-standing Import.io customer for news and business-information collection.

platform

A streaming-guide company

Runs its web data collection on the Import.io platform.

data journalism

A major US newspaper

Published data journalism built on Import.io data.

Coverage is cheap. Understanding it isn’t.

News is infinite and mostly duplicated. The value is in knowing which articles are the same story, who broke it, how it spread, and where your own journalism is being republished or reused. Import.io captures coverage from thousands of outlets and structures it into stories, not just links.

Any feed can hand you ten thousand links. Editors, comms teams and rights holders need the structure underneath: which pieces are the same story, which are originals, which are wire copy, what changed in the update, and which sites lifted your reporting whole.

That same structure is now a licensing question. As AI systems read the news, publishers need evidence of where their content appears and how it is used. We capture it with timestamps, snapshots and similarity scores that stand up in a conversation with a platform.

who uses it
  • Newsrooms and editorial strategy
  • Media monitoring and PR
  • Rights, licensing and legal
  • Audience and product teams
  • Research and intelligence units

What media & publishing teams do with it.

Six programs we run most often in this industry.

Media monitoring

Mentions of companies, people and topics across outlets, clustered into stories with sentiment and reach.

monitoring

Competitive editorial intelligence

What rival outlets publish, how fast, and which stories they break first.

editorial

Content reuse and licensing

Where your articles are republished, rewritten or excerpted, with overlap scores and snapshots.

rights

Archive and metadata enrichment

Headlines, bylines, dates, sections and entities for back catalogues and syndication partners.

archive

Trend and topic detection

Emerging stories and topics across languages before they peak.

trends

Sponsored and branded content

Where brands appear in sponsored content and partnerships across publishers.

commercial

The data, field by field.

fieldtypeexample
article_idstrrl-2026-0303-0612
outletstrRiverton Ledger
urlurlledger.example/harbor-bridge
headlinestrHarbor Bridge closes after cracked welds found
bylinestrStaff
published_atts2026-03-03T06:12Z
updated_atts2026-03-03T09:40Z
word_countint1,240
languagestren
cluster_idstrstory-88121
is_syndicatedboolfalse
paywallenummetered
where it comes from
  • News sites and news sitemaps
  • Wire services’ public pages
  • Blogs and newsletters’ web archives
  • Press release distribution sites
  • Broadcasters’ web articles and transcripts pages
  • Aggregators and content sites

The hard parts, handled.

What breaks when this is done with scripts, and how Import.io handles it.

01

Story clustering

Near-duplicate detection and entity matching group articles into stories and separate originals from syndication.

02

First-publisher detection

Publication times are captured from the page, the sitemap and our first sighting, so “who was first” has evidence.

03

Paywall-aware capture

Only what an outlet shows publicly is captured; licensed full-text access uses credentials the customer holds.

04

Reuse detection

Your archive is fingerprinted and compared against the open web, with overlap scores and snapshots for every match.

Three ways to get it.

Same capture engine underneath each one.

data scope

Coverage is captured as outlets publish it publicly; paywalled text is not captured without licensed access. Full text is delivered only where your licence allows; otherwise metadata, summaries and links.

Media & Publishing questions.

Straight answers.

Can you deliver full article text?

Where your licence allows it, yes. Otherwise we deliver metadata, headlines, summaries, entities and links, which covers most monitoring needs.

How fast do you pick up breaking news?

Breaking-news and section pages can be checked every 60 seconds; the long tail of outlets is checked on a schedule set by how often they publish.

How do you know who published first?

We compare the published time on the page, the news sitemap and our own first sighting, and keep all three as evidence.

Can you find where our content is reused?

Yes. Your archive is fingerprinted and compared with what’s published across the web, with overlap scores and captured snapshots for each match.

Do you cover non-English outlets?

Yes. Outlets are covered across languages and markets, with language identification on every article.

Tell us the sources.
We’ll show you the data.