Automated Attribute Extraction from Natural Language Documents

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current techniques fail to efficiently and automatically extract product attributes and values from natural language documents, limiting their application in retail data analysis for applications like demand forecasting and product comparison.

Innovation Solution

The use of classification algorithms, including supervised, semi-supervised, and unsupervised methods, to identify and associate product attributes and values from natural language documents, forming attribute-value pairs for improved data representation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If manual processes are used to extract product attributes and values from product descriptions, then extraction accuracy can be maintained through human inspection, but the process becomes inefficient and expensive

Engineering Contradiction:
Improveextraction efficiencyVSAvoidextraction cost
Core Design Contradiction:
ProductivityVSEase of manufacture

Solution Approach 1:

The system enables automated self-extraction of product attributes and values using classification algorithms that process product descriptions without human intervention. The algorithms automatically identify and classify attributes (e.g., brand, size, color) and their corresponding values, allowing the system to serve itself rather than requiring manual inspection for each product.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual inspection process with computational classification algorithms. These algorithms use supervised, semi-supervised, or unsupervised learning methods to automatically extract and classify product attributes from text descriptions, substituting human cognitive work with automated computational processes that are both faster and more cost-effective.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Device complexity

If products are treated as atomic entities with few attributes, then data storage and processing are simplified, but the effectiveness of applications like demand forecasting and product comparison is hindered

Engineering Contradiction:
Improvedata structure simplicityVSAvoidapplication effectiveness
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent segments each product into multiple discrete attribute-value pairs rather than treating it as a single atomic entity. For example, a product description is broken down into attributes such as brand, size, color, material, and their corresponding values. This segmentation enriches the data structure with detailed product characteristics while maintaining organized, queryable formats that improve application effectiveness.

Inventive Principle:
Principle #1Segmentation

3Productivity

If classification algorithms are used to automatically extract attributes and values, then processing efficiency is greatly improved, but the complexity of the extraction system increases

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidextraction system complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent employs different classification algorithm parameters and approaches (supervised, semi-supervised, unsupervised learning) depending on the specific extraction task and data availability. By adjusting algorithmic parameters and selecting appropriate classification strategies, the system achieves high processing efficiency while managing complexity through flexible, context-appropriate method selection rather than a single fixed complex system.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS7996440B2Extraction of attributes and values from natural language documents
Publication Date: 2011.08.09 ACCENTURE GLOBAL SERVICES LTD
  • US7996440B2 patent drawing
  • US7996440B2 patent drawing
  • US7996440B2 patent drawing

AI summary

One or more classification algorithms are applied to at least one natural language document in order to extract both attributes and values of a given product. Supervised classification algorithms, semi-supervised classification algorithms, unsupervised classification algorithms or combinations of such classification algorithms may be employed for this purpose. The at least one natural language document may be obtained via a public communication network. Two or more attributes (or two or more values) thus identified may be merged to form one or more attribute phrases or value phrases. Once attributes and values have been extracted in this manner, association or linking operations may be performed to establish attribute-value pairs that are descriptive of the product. In a presently preferred embodiment, an (unsupervised) algorithm is used to generate seed attributes and values which can then support a supervised or semi-supervised classification algorithm.