Entity Relationship Map Creation via Token Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing automated data processing systems face challenges in efficiently extracting and processing tokens, particularly numeric values and special characters, which are not effectively handled by prior token mechanisms, limiting their ability to create accurate entity relationship maps.

Innovation Solution

A method and system that identify unique and recurring lexical tokens, eliminate outliers based on standard deviation, and create an entity relationship map by associating unique tokens with their neighbors, using a cluster computer network and text analytics system to process and store relevant lexical matter.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional token mechanisms are used to process all lexical and non-lexical content, then complete data coverage is achieved, but processing efficiency and accuracy deteriorate due to unnecessary inclusion of numeric values and special characters

Engineering Contradiction:
Improvetoken processing accuracyVSAvoiddata processing efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent extracts and separates lexical tokens from non-lexical content (numeric values, special characters) in the document stream. The collector module identifies and isolates meaningful lexical matter, excluding irrelevant non-lexical elements from further processing, thereby improving both accuracy and efficiency

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies different processing qualities to different types of content: lexical tokens receive full analytical processing while non-lexical content is excluded or handled differently. This localized quality approach ensures resources are focused on meaningful data without the burden of processing all content uniformly

Inventive Principle:
Principle #3Local quality

2Loss of information

If all tokens including numeric values and special characters are processed, then comprehensive data reclamation is achieved, but system complexity and processing burden increase

Engineering Contradiction:
Improvedata reclamation completenessVSAvoidprocessing system complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The system extracts only the necessary lexical information from the document stream, separating meaningful tokens from irrelevant numeric values and special characters. This extraction approach maintains data reclamation effectiveness while reducing processing complexity

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

Instead of including all content and filtering out irrelevant parts, the patent inverts the approach by defaulting to exclusion of non-lexical content and only including meaningful lexical tokens. This inversion simplifies the processing system while maintaining comprehensive data reclamation of relevant information

Inventive Principle:
Principle #13The other way round (Inversion)

3Measurement precision

If outlier elimination based on standard deviation is implemented, then entity relationship map accuracy is improved, but additional processing time is required

Engineering Contradiction:
Improveentity relationship accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary frequency analysis and outlier elimination on tokens before building the entity relationship map. By pre-processing and eliminating outliers early in the pipeline, the system reduces the computational burden on subsequent relationship mapping operations, balancing accuracy improvement with time efficiency

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11568142B2Extraction of tokens and relationship between tokens from documents to form an entity relationship map
Publication Date: 2023.01.31 INFOSYS LTD
  • US11568142B2 patent drawing
  • US11568142B2 patent drawing
  • US11568142B2 patent drawing

AI summary

A system and method of creating an entity relationship map includes receiving a stream of lexical matter associated with one or more categories (302) and identifying one or more tokens from the received lexical matter based on the one or more categories (304). A frequency of one or more of unique lexical token and recurring lexical token are determined (306) and one or more outliers based on a standard deviation range associated with the at least one category is eliminated (308). Sentences with the one or more recurring lexical tokens are selected (310) to find one or more lexical neighbors and the entity relationship map is created based on an association between the unique lexical tokens and the at least one lexical neighbor (312).