Data types · Product Details

Every product page, as a clean record. Identifiers, specs, variants, prices and images, field by field.

A product detail page carries more decisions than any other page in commerce: the price, the promise, the content that sells or doesn’t. Import.io turns those pages into typed records across every retailer and marketplace you care about, with variants split out and products matched to your own catalogue.

Product record · M18 FUEL Hammer Drill Kithomedepot.com · 32 fields · matched by GTIN · sample
32fields per record
98.6%field fill rate
2variants split out
1match to your SKU
FieldValueSource on page
gtin0045242566151JSON-LD + spec table
titleM18 FUEL ½ in. Hammer Drill Kith1
brandMilwaukeebreadcrumb + JSON-LD
price / was$159.00 / $179.00price block
variantsTool only · Kit (2 batteries)selector
specsChuck ½ in · 1,400 in-lbs · 4.4 lbspec table
images9 · 1600 pxgallery
availabilityIn stock · pickup todayfulfilment block
live events
32typical fields per product record
2.6Mproducts collected a day
190locales, prices as shoppers see them
2×complete records vs conventional scraping, in a published customer study

The page is the product.

For most shoppers, the product detail page is the product: its title, images, specifications and price are all they will ever see before buying. For retailers that page is a competitive surface; for brands it is a compliance problem; for data companies it is raw material.

The difficulty is completeness. Conventional scrapers miss fields on a large share of pages — lazy-loaded prices, variants behind selectors, specs rendered by scripts. Import.io captures the whole record, validates it against a schema, and matches it to your catalogue.

who uses it
  • Retail pricing and category teams
  • Brand ecommerce and content teams
  • Digital shelf and commerce-intelligence platforms
  • Marketplace operators
  • Data science and AI teams

What product details data answers.

Six questions teams use it for most.

Competitive product intelligence

Every competitor’s specs, prices and positioning for the products you compete on.

competition

Catalogue enrichment

Missing attributes, images and specifications filled from manufacturer and retailer pages.

enrichment

Content compliance

Titles, images and specs on every retailer checked against your master data.

compliance

Assortment matching

Product-to-product matches across retailers by identifier, attributes and image.

matching

New product detection

Launches, discontinuations and range changes as they appear.

launches

Training data for commerce AI

Structured product records with provenance for search, recommendation and agent use.

AI

The data, field by field.

fieldtypeexample
gtin / mpnstr0045242566151 / 2904-22
titlestrM18 FUEL ½ in. Hammer Drill Kit
brandstrMilwaukee
category_pathlistTools › Power Tools › Drills
pricenum159.00
was_pricenum179.00
currencystrUSD
variantslisttool_only · kit_2x5ah
specificationsobj{chuck: “½ in”, torque: 1400}
imageslist9 URLs
descriptionstrmain content, cleaned
retrieved_atts2026-09-25T16:14:02Z
where it comes from
  • Retailer product detail pages
  • Marketplace listings
  • Manufacturer and brand sites
  • Distributor and B2B catalogues
  • Product feeds and structured markup

The hard parts, handled.

What breaks when this is done with scripts, and how Import.io handles it.

01

Complete, not partial

Lazy-loaded prices, script-rendered specs and tabbed content are rendered and captured, and fill rates are checked per field on every run.

02

Variants split out

Size, colour and pack variants are captured as separate records with their own price and availability, not just the default.

03

Structured specs

Specification tables are parsed into typed attributes with units normalised, so 1,400 in-lbs and 158 N·m compare.

04

Matched to you

Records are matched to your catalogue by GTIN and model first, then attributes and image, with confidence scores.

Three ways to get it.

Same capture engine underneath each one.

data scope

Programs collect publicly displayed product information. Product images and descriptions remain the property of their owners; how they may be reused is agreed per program.

Product details data questions.

Straight answers.

Which fields can you extract from a product page?

Everything the page shows: identifiers, title, brand, category, price and was-price, variants, specifications, images, description, availability, delivery, seller, ratings and review counts — typically 25 to 40 fields.

How do you handle variants?

Each size, colour or pack variant is captured as its own record with its own price and availability, linked to its parent product.

Can you match products to our catalogue?

Yes: by GTIN and model number first, then attributes, text and image, each match with a confidence score and a review queue.

How complete is the data?

Fill rates are measured per field on every run and failures are held back in managed programs. In a published study, Import.io delivered twice as many complete product records as conventional scraping.

Which retailers can you cover?

Any retailer or marketplace you name, subject to feasibility per source, which is confirmed in writing before anything is quoted.

Tell us the sources.
We’ll show you the data.