Entity Relationship Map Creation via Token Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing automated data processing systems face challenges in efficiently extracting and processing tokens, particularly numeric values and special characters, which are not effectively handled by prior token mechanisms, limiting their ability to create accurate entity relationship maps.
Innovation Solution
A method and system that identify unique and recurring lexical tokens, eliminate outliers based on standard deviation, and create an entity relationship map by associating unique tokens with their neighbors, using a cluster computer network and text analytics system to process and store relevant lexical matter.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional token mechanisms are used to process all lexical and non-lexical content, then complete data coverage is achieved, but processing efficiency and accuracy deteriorate due to unnecessary inclusion of numeric values and special characters
Solution Approach 1:
The patent extracts and separates lexical tokens from non-lexical content (numeric values, special characters) in the document stream. The collector module identifies and isolates meaningful lexical matter, excluding irrelevant non-lexical elements from further processing, thereby improving both accuracy and efficiency
Solution Approach 2:
The patent applies different processing qualities to different types of content: lexical tokens receive full analytical processing while non-lexical content is excluded or handled differently. This localized quality approach ensures resources are focused on meaningful data without the burden of processing all content uniformly
2Loss of information
If all tokens including numeric values and special characters are processed, then comprehensive data reclamation is achieved, but system complexity and processing burden increase
Solution Approach 1:
The system extracts only the necessary lexical information from the document stream, separating meaningful tokens from irrelevant numeric values and special characters. This extraction approach maintains data reclamation effectiveness while reducing processing complexity
Solution Approach 2:
Instead of including all content and filtering out irrelevant parts, the patent inverts the approach by defaulting to exclusion of non-lexical content and only including meaningful lexical tokens. This inversion simplifies the processing system while maintaining comprehensive data reclamation of relevant information
3Measurement precision
If outlier elimination based on standard deviation is implemented, then entity relationship map accuracy is improved, but additional processing time is required
Solution Approach 1:
The patent performs preliminary frequency analysis and outlier elimination on tokens before building the entity relationship map. By pre-processing and eliminating outliers early in the pipeline, the system reduces the computational burden on subsequent relationship mapping operations, balancing accuracy improvement with time efficiency
Data Source
AI summary
A system and method of creating an entity relationship map includes receiving a stream of lexical matter associated with one or more categories (302) and identifying one or more tokens from the received lexical matter based on the one or more categories (304). A frequency of one or more of unique lexical token and recurring lexical token are determined (306) and one or more outliers based on a standard deviation range associated with the at least one category is eliminated (308). Sentences with the one or more recurring lexical tokens are selected (310) to find one or more lexical neighbors and the entity relationship map is created based on an association between the unique lexical tokens and the at least one lexical neighbor (312).


