The Client
A home decor and furnishings brand selling a large catalogue across several marketplaces and its own D2C storefront — lighting, soft furnishings, wall decor, and small furniture. Thousands of SKUs, sourced from many suppliers, each supplier supplying product data in a different, incomplete format.
Client details are anonymised. Figures are representative of the engagement.
The Problem: A Catalogue Full of Half-Built Pages
The brand's catalogue looked complete in a spreadsheet and was anything but on the page.
Because SKUs came from dozens of suppliers, the product data arrived inconsistent and incomplete. Some listings had one low-resolution image; competitors had eight. Some had a material and a dimension; competitors had material, dimension, weight, finish, care instructions, assembly requirement, and room recommendation. The brand's own storefront and its marketplace listings were, on a large share of the catalogue, thinner than the competition on exactly the attributes that decide a home decor purchase.
This showed up in two places that matter.
Search. On marketplaces, product discoverability depends heavily on structured attributes. A sofa with no "seating capacity" attribute does not appear when a shopper filters for a three-seater. The brand was invisible to filtered search on any attribute it had failed to populate — and it had failed to populate many.
Conversion. Home decor is a considered, visual purchase. A shopper deciding on a light fixture wants to see it from multiple angles, in a room, at a known size. A single image and a two-line spec loses to a competitor's eight images and full specification, even at the same price.
The catalogue lead's summary: we're not losing on price or product. We're losing on pages that aren't finished.
Why the Brand Could Not Fix This Manually
The obvious fix — populate the missing data — was the problem, because of scale and source fragmentation.
Scale. Tens of thousands of SKUs, each missing a different subset of attributes and images. Manually researching and populating them, SKU by SKU, was a multi-year project the brand did not have the staff for.
Source fragmentation. The missing data existed — on the manufacturers' own listings, on other marketplaces, on the brand's own better-populated SKUs — but it was scattered across dozens of sources in dozens of formats. There was no single place to copy it from.
Consistency. Even where a team member did populate a field, different people entered "Solid Wood," "solid-wood," and "Sheesham (solid wood)" for the same thing, which defeated the filtered search the exercise was meant to fix.
The Solution: Catalog Enrichment at Scale
Product Data Scrape built a home decor catalog enrichment pipeline that extracted, normalised, and mapped product images and specifications across marketplaces at scale.
Image extraction. For each SKU, all available product images were located across the manufacturer's listing and matched marketplace listings, captured at full resolution, deduplicated, and ordered — so a SKU that had one image on the brand's storefront could be enriched toward the eight the market expected.
Specification extraction. Every available structured attribute was captured — material, dimensions, weight, finish, colour, assembly requirement, care, room, style — from the richest available source per SKU.
Attribute normalisation. This was the step that made the rest usable. Captured attribute values were normalised to a single controlled vocabulary, so "Solid Wood," "solid-wood," and "Sheesham (solid wood)" all resolved to one canonical value that filtered search could act on.
Gap-first prioritisation. The pipeline scored each SKU by how incomplete it was and how much traffic it received, so enrichment effort concentrated where it moved the most revenue first.
Delivery into the catalogue. Enriched, normalised records were delivered in the brand's own schema, ready to load, rather than as a raw dump requiring manual reconciliation.
Sample Data: One SKU, Before and After
An illustrative field-completeness comparison for a single lighting SKU.
| Attribute |
Before Enrichment |
After Enrichment |
| Product images |
1 |
7 |
| Material |
— |
Brushed brass, glass |
| Dimensions |
Height only |
Height, width, drop, canopy diameter |
| Weight |
— |
2.4 kg |
| Bulb type / base |
— |
E27, max 40W |
| Finish |
— |
Antique brass |
| Room recommendation |
— |
Living room, dining |
| Style tag |
— |
Mid-century |
| Care instructions |
— |
Dry cloth, no solvents |
| Filterable attributes populated |
2 |
11 |
The enriched record, structured:
{
"sku": "LIGHT-PENDANT-4471",
"title": "Brass Pendant Lamp — Mid-Century Dome",
"images": [
{"url": "img_001.jpg", "type": "primary", "resolution": "2000x2000"},
{"url": "img_002.jpg", "type": "angle"},
{"url": "img_003.jpg", "type": "in_room"},
{"url": "img_004.jpg", "type": "scale_reference"},
{"url": "img_005.jpg", "type": "detail"},
{"url": "img_006.jpg", "type": "detail"},
{"url": "img_007.jpg", "type": "packaging"}
],
"attributes": {
"material_primary": "brass",
"material_secondary": "glass",
"finish": "antique_brass",
"height_cm": 34,
"width_cm": 28,
"drop_max_cm": 120,
"weight_kg": 2.4,
"bulb_base": "E27",
"wattage_max": 40,
"room": ["living_room", "dining"],
"style": "mid_century",
"care": "dry_cloth_no_solvents"
},
"attribute_completeness_pct": 92,
"normalisation_applied": true
}
Every value in attributes is drawn to a controlled vocabulary — antique_brass, not "Antique Brass" or "brass, antique" — which is what allows the marketplace filter to match it.
What Changed
The catalogue became searchable. Filterable-attribute coverage across the catalogue rose sharply, and SKUs began appearing in filtered searches — by material, by room, by style, by dimension — where they had previously been invisible.
Pages started converting. Enriched pages, with full imagery and complete specifications, converted at a materially higher rate than the thin pages they replaced. In home decor, the completeness of the page is a substantial part of the purchase decision, and the data made that visible.
One vocabulary, everywhere. Normalisation meant the same SKU carried the same canonical attributes on the D2C storefront and every marketplace, which fixed both filtered search and the internal reporting that had been fragmented by inconsistent values.
The Results
| Metric |
Before |
After Enrichment |
| SKUs with complete image sets (5+) |
Minority of catalogue |
Large majority |
| Median filterable attributes per SKU |
~3 |
~10 |
| SKUs appearing in filtered marketplace search |
Low |
Substantially higher |
| Conversion rate on enriched pages, indexed |
100 |
131 |
| Catalogue enrichment throughput |
SKUs per week (manual) |
50,000+ SKUs, structured |
| Attribute value consistency |
Fragmented |
Single controlled vocabulary |
Figures are representative of the engagement outcome.
The catalogue lead's follow-up: we didn't add products or cut prices. We finished the pages we already had, and the numbers moved.
The Lesson
In categories where the purchase is visual and considered — and home decor is the clearest example — the product page is the product. A shopper cannot touch the fabric or lift the lamp. All they have is the imagery and the specification, and if those are thinner than the competitor's, the competitor wins at the same price and the same quality.
The brand's catalogue was not bad. It was unfinished, and unfinished in a way that was invisible in a spreadsheet and decisive on the page. Catalog enrichment at scale did not require new products, new suppliers, or lower prices. It required completing what already existed — across tens of thousands of SKUs, to one consistent standard, faster than any manual effort could.
That is a data problem, and it has a data solution.
Work With Product Data Scrape
Product Data Scrape delivers home decor catalog enrichment at scale: full-resolution image extraction, complete specification capture, attribute normalisation to your controlled vocabulary, and gap-first prioritisation — delivered in your own schema as JSON, CSV, API, or straight into your catalogue system.
We capture publicly available product information only, structured and ready to load.
Ask us for an enrichment audit on a sample of your catalogue — we will show you exactly which SKUs are losing to unfinished pages, and by how much.
Product Data Scrape — turning marketplace complexity into decision-ready data.