Model Selection Using Feature Health Scores
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing model selection systems face challenges in choosing the top K performing subset of models from a model ensemble for prediction in domains with faulty and unreliable sensors, as they often rely on ensemble performance rather than identifying superior individual models.
Innovation Solution
The proposed solution involves a multi-stage approach that clusters health score vectors from sensors, compares model score distributions within clusters to the ensemble distribution, and selects the top K performing models for each cluster, thereby leveraging superior individual models instead of the ensemble.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If model ensemble is used for prediction, then system robustness is improved, but individual model performance is diluted
Solution Approach 1:
The patent segments the model ensemble into individual models and evaluates each model's performance independently using health score vectors. By clustering models based on their health scores and feature importance, the system identifies top-performing individual models rather than relying on aggregated ensemble predictions, thus resolving the contradiction between robustness and precision.
Solution Approach 2:
The patent changes the selection parameter from ensemble-wide metrics to individual model health score vectors. By transforming the selection criterion to focus on models with higher health scores and better feature importance alignments, the system achieves both robustness through health-based filtering and precision through top-model selection.
2Device complexity
If all models in ensemble are deployed, then system complexity is reduced, but computational resources are wasted on poor performing models
Solution Approach 1:
The patent extracts and removes poor-performing models from the deployment set by evaluating each model's health score vector and feature importance. Only the top K performing models are selected for deployment, extracting the unnecessary computational overhead while maintaining deployment simplicity through automated selection.
Solution Approach 2:
Instead of deploying all models in the ensemble (excessive action), the patent applies partial action by selecting only the top K models based on health scores. This partial deployment reduces computational resource consumption while maintaining sufficient prediction accuracy, resolving the contradiction between simplicity and resource efficiency.
3Measurement precision
If feature health scores are collected continuously, then model selection accuracy is improved, but data transmission overhead increases
Solution Approach 1:
The patent performs preliminary action by collecting and processing health score vectors locally at edge nodes before transmission to the central node. By accumulating and pre-processing this data at the edge, the system reduces the need for continuous transmission, thereby improving model selection accuracy while reducing transmission energy consumption.
Solution Approach 2:
The patent introduces edge nodes as intermediaries between sensor nodes and the central node. These intermediaries aggregate health score vectors and perform local processing, reducing the volume and frequency of transmissions to the central node. This intermediary layer maintains selection accuracy while significantly reducing transmission energy overhead.
Data Source
AI summary
Techniques are disclosed for model selection using feature health scores with unreliable sensors. One example method includes clustering health score vectors received from nodes operating in an environment, the health score vectors including feature health scores for sensors used by machine learning models; comparing a model score distribution for an ensemble of the models with model score distributions per cluster, to obtain a set of top K performing models for each cluster, upon receiving new data for prediction, identifying an associated health score vector for the data and using the top K performing models corresponding to the cluster for the associated health score vector to select the top K performing models; and deploying the clusters and model ensembles to the nodes.


