Unstructured Text Token Extraction for Enterprise Labeling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies are ineffective in labeling unstructured text documents within an enterprise context, as they focus on text analysis rather than enterprise-specific terms and require longer texts for statistical models, making it difficult to leverage unstructured data for analytics and machine learning applications.

Innovation Solution

A computer-implemented method that extracts unrecognized tokens from unstructured text documents and relates them to structured data elements from predefined data sources, using natural language and non-natural language elements to identify relevant labels, thereby enabling the labeling of short text snippets and bridging the gap between structured and unstructured data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If statistical models are used to extract terms from unstructured text, then term classification accuracy is improved, but text length requirements increase making it unusable for short documents

Engineering Contradiction:
Improveterm classification accuracyVSAvoidtext length
Core Design Contradiction:
Measurement precisionVSLength of moving object

Solution Approach 1:

The patent introduces an intermediary approach by using token extraction and matching against predefined data sources rather than directly applying statistical models to the entire text. This mediator layer (token extraction and matching) enables accurate classification without requiring long texts, as it focuses on identifying key tokens and their relationships to structured data elements.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent extracts specific tokens from unstructured text documents and separates them for analysis. By taking out key tokens and matching them against structured data sources, the system achieves accurate classification without needing to analyze the entire text, thereby eliminating the requirement for long documents while maintaining precision.

Inventive Principle:
Principle #2Taking out (Extraction)

2Difficulty of detecting and measuring

If existing text-focused classification techniques are applied, then text analysis capability is improved, but enterprise-specific context labeling capability deteriorates

Engineering Contradiction:
Improvetext analysis capabilityVSAvoidenterprise-specific context labeling
Core Design Contradiction:
Difficulty of detecting and measuringVSAdaptability or versatility

Solution Approach 1:

The patent merges text analysis capabilities with enterprise-specific context by combining token extraction from unstructured text with matching against structured data sources that contain enterprise terminology. This integration allows the system to maintain strong text analysis while simultaneously achieving enterprise-specific context labeling through the combination of both capabilities.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent creates a multi-functional system that can handle both general text analysis and enterprise-specific context labeling through a unified approach. The same token extraction and matching mechanism serves both purposes, making the system versatile across different labeling needs without requiring separate specialized techniques.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If manual term assignment is used for unstructured documents, then labeling accuracy is improved, but processing productivity deteriorates

Engineering Contradiction:
Improvelabeling accuracyVSAvoidprocessing speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent enables self-service automated labeling by having the system automatically extract tokens from unstructured documents and match them against structured data sources without human intervention. This self-service capability maintains high labeling accuracy through precise token matching while dramatically improving processing productivity by eliminating manual term assignment requirements.

Inventive Principle:
Principle #25Self-service

4Productivity

If automated classification is implemented without enterprise context, then processing efficiency is improved, but data utility for analytics deteriorates

Engineering Contradiction:
Improveclassification efficiencyVSAvoidenterprise context information
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent performs preliminary action by pre-defining structured data sources containing enterprise-specific terms and relationships before the classification process. This preliminary preparation enables the automated classification system to efficiently match tokens against pre-organized enterprise context, maintaining high processing efficiency while preserving enterprise context information through the pre-established structured data framework.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20230186023A1Automatically assign term to text documents
Publication Date: 2023.06.15 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US20230186023A1 patent drawing
  • US20230186023A1 patent drawing
  • US20230186023A1 patent drawing

AI summary

In an approach, a processor receives an unstructured text document. A processor extracts at least one unrecognized token from the unstructured text document. A processor identifies at least one structured data element in a predefined set of data sources, where the at least one structured data element is related to the at least one extracted unrecognized token from the unstructured text document. A processor relates a label associated with the identified at least one structured data element to the unstructured text document.