Peer-Based ML Estimation for Incomplete Entity Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Estimating the value of a metric for an entity is challenging due to noisy and incomplete information, leading to inaccurate manual generation of estimates.
Innovation Solution
A machine learning architecture that utilizes clusters of reference entities, calculates distances to select peer entities, and applies a trained model to estimate the metric value for a target entity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual estimation methods are used, then simplicity is maintained, but accuracy of metric estimation deteriorates due to noisy and incomplete information
Solution Approach 1:
The patent introduces an automated estimation system that acts as an intermediary between the noisy input data and the required metric estimates. This system uses machine learning models trained on historical data to process and interpret noisy, incomplete information, thereby improving estimation accuracy while managing complexity through automation rather than manual methods
Solution Approach 2:
The patent transforms the estimation problem by changing parameters from manual judgment to automated machine learning predictions. By using trained models that have learned optimal parameter relationships from historical data, the system can accurately estimate metrics even when input information is noisy or incomplete, resolving the contradiction between accuracy and complexity
2Loss of information
If more information about the entity is collected, then completeness of data improves, but noise and processing complexity increase
Solution Approach 1:
The patent applies preliminary action by pre-training machine learning models on comprehensive historical data before actual estimation tasks. This pre-processing phase captures relationships and patterns in advance, allowing the model to handle noisy, incomplete data during deployment without requiring complex real-time processing, thus improving information utilization while managing complexity
Solution Approach 2:
The patent uses copying by creating trained machine learning model copies that encapsulate knowledge from historical data. These model copies can process new entity information efficiently without requiring the full complexity of the original training process, enabling complete information utilization with reduced processing complexity
Data Source
AI summary
A method may include obtaining a cluster. The cluster may include a subset of reference entities. The method may further include calculating distances between features of a target entity and features of the subset of reference entities, selecting, based on the distances, peer entities from the subset, and generating an estimated value of a metric. The generating may include applying, to the features of the target entity, a machine learning model trained using training data including values of the features for the peer entities labeled with a value of the metric. The method may further include presenting the estimated value of the metric.


