Unsupervised Competition-Based Encoding for Entity Identification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Identifying similar entities from voluminous word-based data is challenging due to the complexity of extracting meaningful patterns and relationships.
Innovation Solution
A method that generates phrase vectors from frequency data, calculates similarity metrics, and uses machine learning to create embedded vectors for clustering and categorization, enabling the identification of similar entities and generating recommendations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional word-based data analysis methods are used to identify similar entities, then the process is simple to implement, but the ability to accurately identify similar entities from voluminous data deteriorates
Solution Approach 1:
The patent introduces phrase vectors as an intermediary representation between raw word-based data and similarity analysis. These vectors serve as a mediator that transforms unstructured text into a structured format suitable for computational comparison, enabling accurate identification of similar entities while managing system complexity through standardized transformation processes
Solution Approach 2:
The patent transforms word-based data into phrase vectors by changing the parameter representation from discrete words to continuous vector values. This parameter transformation enables mathematical operations and similarity calculations that are not feasible with raw text, thereby improving measurement precision in identifying similar entities
2Measurement precision
If phrase vectors and machine learning models are used to improve entity identification accuracy, then the precision of identifying similar entities improves, but the computational complexity and processing time worsen
Solution Approach 1:
The patent performs preliminary actions by pre-processing word-based data into phrase vectors and pre-training machine learning models before actual similarity analysis. This preliminary transformation of data into vector representations enables faster comparison and querying operations, reducing processing time during actual entity identification tasks
Solution Approach 2:
The patent creates vector copies of textual data that preserve the semantic meaning while enabling efficient computational operations. These vector copies serve as surrogates for the original text, allowing rapid similarity calculations without repeatedly processing the full text data, thereby reducing processing time
3Quantity of substance
If voluminous word-based data is processed to extract meaningful patterns, then the quantity of analyzed data increases, but the difficulty of detecting and measuring relationships worsens
Solution Approach 1:
The patent replaces manual or simple mechanical text analysis methods with machine learning-based vector processing systems. This substitution enables automated detection of patterns and relationships in voluminous data through learned representations, significantly reducing the difficulty of extracting meaningful insights from large datasets
Data Source
AI summary
A method collects word-based data corresponding to a first identifier. A first phrase vector is generated for the first identifier by extracting frequency data from the word-based data. A similarity metric is generated corresponding to the first identifier and a second identifier by comparing the first phrase vector of the first identifier to a second phrase vector of the second identifier. A tuple is generated that includes the first identifier and the second identifier using the similarity metric. A machine learning model is trained with the tuple to generate an embedded vector corresponding to the first identifier.


