Entity Selection Metrics for Predictive Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for optimizing predictive machine learning models fail to effectively assess and compare the suitability of initial entities inputted, leading to difficulties in evaluating their impact on model performance.
Innovation Solution
A system and method for generating comparison metrics based on data from knowledge graphs and predictive models, allowing users to select and evaluate entities by aggregating predictions, extracting metadata, and computing metrics such as overlap, correlations, and protein-protein interactions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If researchers pre-select entities to investigate in a knowledge graph, then the focus of predictive models can be directed to specific areas, but the number of similar or related entities becomes too large making quality assessment difficult
Solution Approach 1:
The patent introduces an intermediary evaluation system that mediates between the selected entities and the final results. This system computes multiple metrics (data quality, model performance, result utility) to differentiate entity quality, transforming the undifferentiated large set of entities into a ranked hierarchy that guides researchers to the most promising candidates.
Solution Approach 2:
The patent changes the parameters used to assess entities from simple counts to multi-dimensional metrics including data quality scores, model performance indicators, and result utility measures. This parameter transformation enables differentiation of entity quality even when entities are similar, allowing researchers to identify the most valuable entities efficiently.
2Reliability
If multiple predictive models are used to generate predictions, then the coverage and depth of analysis increase, but the complexity of evaluating and comparing model suitability increases
Solution Approach 1:
The patent merges the evaluation of multiple predictive models into a unified framework that assesses all models simultaneously across multiple dimensions. By combining data quality assessment, model performance evaluation, and result utility measurement into a single integrated system, the patent reduces the complexity of comparing multiple models while maintaining comprehensive reliability assessment.
Solution Approach 2:
The patent creates a universal evaluation framework that can assess any predictive model regardless of its specific type or configuration. This multi-functional system handles diverse models through standardized metrics, enabling straightforward comparison and selection without requiring model-specific evaluation procedures.
3Measurement precision
If comprehensive metrics are computed for entity evaluation, then the quality assessment becomes more accurate, but the computational resources and time required increase
Solution Approach 1:
The patent performs preliminary actions by pre-computing data quality metrics and entity characteristics before running predictive models. This advance preparation stores reusable information that accelerates the evaluation process, allowing comprehensive metrics to be computed efficiently when entities need to be assessed without repeating all computational steps.
Data Source
AI summary
Embodiments of present disclosure provide a system, apparatus and method(s) for generating a set of metrics for evaluating entities used with a predictive machine learning model, the method comprising: selecting one or more sets of entities from a data sources for generating a plurality of predictions aggregated from said one or more sets of entities using one or more pre-trained predictive models; selecting a subset of predictions from the plurality of predictions based on said one or more sets of entities in relation to the data source; extracting metadata from the data source associated with the subset of predictions, where the metadata comprises entity metadata and predicted metadata; generating the set of metrics based on the metadata extracted and the subset of predictions; and outputting the set of metrics for evaluation.


