Entity Resolution Indexing for Complex Activity Data Matching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing computer systems face inefficiencies in evaluating entity activities due to unorganized and unintelligible data collection, requiring significant resource costs and difficult-to-understand results.
Innovation Solution
A computer-based method utilizing a processor to receive, process, and match data items from multiple databases using heuristic searches and machine learning models to generate entity activity indexes, enabling efficient data organization and visualization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If existing systems collect and evaluate entity activities without organization, then comprehensive data collection is achieved, but resource cost increases and results become difficult to understand
Solution Approach 1:
The patent segments entity activities by matching them to specific entity records in a database. Each activity is associated with a unique entity identifier, dividing the unorganized data into structured segments that can be individually tracked and evaluated. This segmentation transforms the chaotic data collection problem into an organized framework where each activity has a clear ownership and context.
Solution Approach 2:
The patent introduces an intermediary matching process that connects activities to entity records through heuristic search and machine learning models. This intermediary layer acts as a mediator between raw activity data and structured entity information, automatically resolving relationships and organizing data without requiring manual intervention. The intermediary process simplifies the evaluation complexity while maintaining comprehensive data coverage.
2Measurement precision
If heuristic search is performed to match candidate entities, then matching precision is improved, but processing time increases
Solution Approach 1:
The patent performs preliminary actions by pre-processing activity data and entity records before the actual matching process. Features are extracted and standardized in advance, and entity records are pre-organized with relevant attributes. This preliminary preparation reduces the computational burden during the heuristic search phase, allowing for more precise matching without proportional increases in processing time.
Solution Approach 2:
The patent changes parameters dynamically during the matching process by adjusting search criteria and evaluation thresholds based on the complexity of the data and the specific matching requirements. The system can shift between different matching strategies (e.g., exact match vs. fuzzy match) and adjust the depth of heuristic search based on available computational resources, optimizing the balance between precision and time efficiency.
3Measurement precision
If machine learning model is used to predict entity matches, then matching accuracy is improved, but computational resource cost increases
Solution Approach 1:
The patent applies partial action by using machine learning models selectively rather than universally. The ML model is deployed only for activities that require sophisticated matching, while simpler activities can be processed using faster, less resource-intensive rules-based approaches. This partial application of ML reduces overall computational energy consumption while maintaining high accuracy where needed most.
Solution Approach 2:
The patent creates simplified representations or copies of complex entity relationships and activity patterns that can be processed more efficiently. Instead of working with the full complexity of original data throughout the entire process, the system creates distilled feature representations and summary statistics that capture essential patterns at lower computational cost, enabling accurate matching with reduced energy expenditure.
Data Source
AI summary
In order to facilitate the entity resolution and entity activity tracking and indexing, systems and methods include receiving first source records from a first database and second source records from a record database. A candidate set of second source records is determined by a heuristic search in the set of second source records. A candidate pair feature vector associated with each candidate pair of first and second source records is generated. An entity matching machine learning model predicts matching first source records for each candidate second source record based on the respective candidate pair feature vector. An aggregate quantity associated with the matching first source records is aggregated from a quantity associated with each first source record, and a quantity index for each candidate second source record is determined based the aggregate quantities. Each quantity index is displayed to a user.


