Header-token text segmentation via unsupervised lexical analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Supervised text segmentation techniques face challenges such as high costs due to manual tagging of data and limited scalability across different product types, making it difficult to efficiently segment product descriptions on e-commerce websites.

Innovation Solution

The method employs automatic text segmentation using header tokens as hints of relevance, estimating probabilities of token relevance and irrelevance to identify the most relevant segment in a description without requiring expensive manual tagging, utilizing unsupervised learning and lexical associations to adapt to various item categories.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If supervised text segmentation techniques are used, then text can be segmented into meaningful units, but manual tagging costs are high and scalability is limited

Engineering Contradiction:
Improvetext segmentation accuracyVSAvoidmanual tagging cost
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

The system performs self-service by automatically segmenting product descriptions using unsupervised learning algorithms without requiring manual human tagging. The algorithm autonomously identifies relevant segments by analyzing lexical associations and token probabilities, eliminating the need for expensive human annotators while maintaining segmentation quality.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical process of manual human tagging with an automated computational system. Instead of relying on human annotators to manually mark relevant portions of text, the system uses unsupervised learning algorithms, lexical association analysis, and probability calculations to automatically perform the segmentation task that previously required human mechanical effort.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If supervised text segmentation techniques are used, then text segmentation can be performed, but scalability across different product types is limited

Engineering Contradiction:
Improvetext segmentation accuracyVSAvoidscalability across product types
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system achieves universality by designing a product-description-specific unsupervised learning algorithm that can handle multiple product types and categories without requiring retraining or adaptation to specific domains. The algorithm processes any product description by analyzing lexical associations and token probabilities, making it universally applicable across diverse e-commerce product types while maintaining consistent segmentation performance.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If a large number of tagged training cases are used, then segmentation rules can be learned accurately, but the process becomes expensive and time-consuming

Engineering Contradiction:
Improvesegmentation rule accuracyVSAvoidtraining data preparation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system extracts and utilizes lexical associations between tokens as the basis for segmentation, eliminating the need for extensive tagged training data. By focusing on the inherent lexical relationships and probability distributions in the text itself, the algorithm derives segmentation rules directly from the data structure rather than requiring time-consuming preparation of large annotated training sets.

Inventive Principle:
Principle #2Taking out (Extraction)

4Measurement precision

If manual tagging is performed to create training data, then supervised segmentation models can be trained, but annotation costs increase significantly

Engineering Contradiction:
Improvesegmentation model performanceVSAvoidannotation cost
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

The system performs self-service by automatically segmenting product descriptions using unsupervised learning algorithms without requiring manual human tagging. The algorithm autonomously identifies relevant segments by analyzing lexical associations and token probabilities, eliminating the need for expensive human annotators while maintaining segmentation quality.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces expensive, time-consuming manual annotation with a computationally efficient unsupervised learning approach that processes text directly without requiring costly training data preparation. The system uses affordable computational resources to perform lexical association analysis and probability calculations, substituting expensive human labor with cheaper automated processing.

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

Data Source

PatentUS9529862B2Header-token driven automatic text segmentation
Publication Date: 2016.12.27 PAYPAL INC
  • US9529862B2 patent drawing
  • US9529862B2 patent drawing
  • US9529862B2 patent drawing

AI summary

A method and a system to automatically segment text based on header tokens is described. A relevance value and an irrelevance value are determined for each token in a description, assuming no tokens are left out of computations. The irrelevance value is based on occurrences of a token in a sample set of descriptions. The relevance value is an estimated probability of relevance based on the header of the description being segmented.