Attribute Extraction from Text Sources Using NLP Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional product description management systems face challenges in handling diverse, voluminous, and multilingual data due to manual processes, human subjectivity, and the complexity of integrating data from various sources.

Innovation Solution

An attribute analysis system that utilizes advanced natural language processing (NLP) and machine learning (ML) techniques, including question-answering models, distribution-based corrections, and harmonization processes, to automatically extract, standardize, and analyze product attributes across different languages and formats.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual tagging and attribute extraction are used, then human judgment and flexibility are applied, but labor intensity increases and human error susceptibility increases

Engineering Contradiction:
Improveattribute extraction accuracyVSAvoidprocessing efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system enables automated self-service through NLP and ML models that automatically extract attributes from product descriptions without human intervention. The models process multilingual data, perform normalization, and generate structured outputs autonomously, eliminating manual tagging while maintaining extraction accuracy through sophisticated language understanding capabilities.

Inventive Principle:
Principle #25Self-service

2Productivity

If automated NLP and ML systems are used, then processing speed and scalability improve, but handling multilingual data and real-world complexities becomes more difficult

Engineering Contradiction:
Improvedata processing speedVSAvoidmultilingual data handling capability
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The system implements multi-functional NLP and ML models capable of processing multiple languages and various data formats simultaneously. The unified architecture handles product descriptions, specifications, and attributes across different languages and sources, providing both high processing speed and broad adaptability through a single versatile platform.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system dynamically adjusts processing parameters based on input data characteristics, including language detection, format identification, and complexity assessment. This enables the automated system to adapt to multilingual and diverse real-world data while maintaining efficient processing through parameter optimization.

Inventive Principle:
Principle #35Parameter changes

3Quantity of substance

If data from multiple sources is integrated, then data comprehensiveness improves, but normalization complexity and processing overhead increase

Engineering Contradiction:
Improvedata volume and comprehensivenessVSAvoidnormalization process complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The system segments the normalization process into distinct modular components: language detection, format identification, attribute extraction, and standardization. Each module handles specific aspects of data from different sources independently, reducing overall complexity while comprehensively processing multilingual and multi-format data inputs.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12292908B1Attribute extraction from text sources
Publication Date: 2025.05.06 TREDENCE INC
  • US12292908B1 patent drawing
  • US12292908B1 patent drawing
  • US12292908B1 patent drawing

AI summary

Embodiments include a method. The method includes receiving, by the computer system, a first input. The method includes generating, by at least using a language detection model, a set of predicted languages. The method includes generating, using different translation models, a different language translation of the first input for each language in the set of predicted languages. The method includes generating, using a first question-answer model, a set of answers. The method includes performing a distribution-based correction, by at least using a distribution source and at least one answer included in the set of answers, to determine the attribute represented by the arbitrary values included in the first input. The method includes causing a device to present a recommendation based at least in part on the attribute represented by the arbitrary values included in the first input.