Data Selection for ML Models via Subset Metric Profiling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data profiling and selection techniques for machine learning models are limited by their focus on numeric and categorical variables, leading to suboptimal performance due to environmental and security constraints, requiring excessive computing resources and memory, especially in secure environments.
Innovation Solution
A computer-implemented method for data selection that generates models based on dataset metric values, identifies data subsets, and optimizes performance by calculating subset metric values, allowing for efficient data profiling and selection within secure environments without exhaustive searches, reducing resource requirements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If current data profiling techniques are used that focus on numeric and categorical variables, then data selection can be performed, but model performance is suboptimal and excessive computing resources and memory are required
Solution Approach 1:
The patent changes the parameters used for data profiling from traditional numeric and categorical variables to include unstructured data characteristics and security environment constraints. This enables more effective data selection that improves model performance while reducing the computing resources needed by focusing on relevant data attributes.
Solution Approach 2:
The patent segments the data selection process into multiple stages: initial data profiling, subset identification based on performance thresholds, and iterative refinement. This segmentation allows the system to avoid exhaustive search of all possible data subsets, thereby reducing memory and computing resource requirements while still achieving optimal model performance.
2Measurement precision
If exhaustive data profiling of multiple datasets is performed to optimize model performance, then model accuracy improves, but excessive computing resources and memory are required
Solution Approach 1:
The patent applies partial action by performing data profiling on a selective basis rather than exhaustively profiling all possible data subsets. The system identifies data subsets that meet performance thresholds and focuses profiling efforts on those candidates, achieving high model accuracy without requiring memory resources for exhaustive analysis of all potential subsets.
3Reliability
If data is transmitted within secure environments for profiling, then model optimization can occur, but security limitations and computing resource allocation are constrained
Solution Approach 1:
The patent introduces an intermediary data profiling system that operates within secure environments to evaluate data subsets before model training. This intermediary layer enables security-conscious data selection by assessing data characteristics and generating performance predictions without requiring transmission of confidential information, thereby maintaining security while improving profiling efficiency.
Data Source
AI summary
A method, computer system, and a computer program for data selection is provided. The present invention may include generating a first model associated with a dataset. The present invention may further include determining a first model performance level associated with the first model based on a plurality of dataset metric values of the dataset. The present invention may further include a plurality of data subsets of a dataset based on the first model performance level failing to exceed a performance threshold and calculating a plurality of subset metric values associated with the plurality of data subsets. The present invention may further include generating a second model associated with at least one data subset based on the plurality of subset metric values and determining an optimization associated with the first model based on a second model performance level associated with the second model exceeding the performance threshold.


