Product Title Attribute Extraction for Accurate Similarity

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional computer systems are limited in accurately determining product similarity from online product listings due to the lack of universal product identifiers and the variability in product titles, leading to inaccurate comparisons of products that humans recognize as similar.

Innovation Solution

A computer-implemented system that retrieves and refines product titles, extracts attributes, and generates product identifiers by combining tags from historical data, allowing for accurate comparisons and identifications of products, even when titles include multiple options.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional computer systems compare product identifiers (UPC, EAN) to determine product similarity, then product comparison can be performed using standardized codes, but this approach becomes useless or impractical in areas where these identifiers are not widely used

Engineering Contradiction:
Improveproduct comparison capabilityVSAvoidaccuracy of product identification
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent introduces natural language processing as an intermediary layer between product titles and comparison logic. Instead of directly comparing product identifiers or raw titles, the system uses NLP to extract semantic attributes from product titles, creating a mediator representation that enables accurate comparison without relying on universal product identifiers.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces the mechanical system of identifier-based comparison with an intelligent system that uses natural language understanding. Instead of mechanically matching UPC/EAN codes or string titles, the system substitutes this with semantic analysis that extracts meaning from product titles, enabling comparison based on actual product attributes rather than formal identifiers.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Ease of operation

If conventional computer systems perform string comparison of product titles to determine similarity, then they can process titles without universal identifiers, but they fail to recognize products as similar when titles include different words, descriptors, and quantities

Engineering Contradiction:
Improveproduct title processingVSAvoidproduct similarity determination
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent extracts essential attributes from product titles by applying NLP techniques. Instead of processing entire titles as strings, the system takes out and isolates key semantic elements such as product type, brand, specifications, and quantities, creating a structured representation that enables precise similarity determination by comparing only the essential attributes.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent transforms product titles from unstructured text into structured attribute parameters. By changing the representation from raw strings to parameterized attribute-value pairs, the system enables precise comparison based on semantic meaning rather than literal string matching, allowing titles with different wording to be recognized as similar when they describe the same product.

Inventive Principle:
Principle #35Parameter changes

3Loss of information

If the system extracts attributes from product titles to enable accurate comparison, then product understanding improves, but the complexity of processing and analyzing unstructured text increases

Engineering Contradiction:
Improveproduct attribute extractionVSAvoidtext processing system
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent segments the product title processing into distinct stages: text preprocessing, entity recognition, attribute extraction, and identifier generation. By dividing the complex NLP task into manageable segments, the system reduces processing complexity while maintaining accurate attribute extraction. Each segment handles a specific aspect of the transformation from unstructured text to structured attributes.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11615453B2Systems and methods for intelligent extraction of attributes from product titles
Publication Date: 2023.03.28 COUPANG CORP
  • US11615453B2 patent drawing
  • US11615453B2 patent drawing
  • US11615453B2 patent drawing

AI summary

Some aspects of the present disclosure are directed to computerized methods for extracting attributes from product titles. The method may include: retrieving a title associated with a product listing and historical product title data; refining the title; determining at least one tag associated with an attribute; generating, based on the at least one extracted tag and the historical title data, a first combination of one or more attributes; determining whether the title includes at least one plurality of product options, and if so: determining, for each product option hi the plurality of product options, a second combination of one or more attributes by removing attributes associated with alternative product options from the first combination; and generating a product identifier based on the second combination; and if the title does not include at least one plurality of product options, generating, the product identifier based on the first combination.