Dataset Merit Evaluation for Machine Learning Model Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The challenge lies in efficiently selecting and evaluating candidate datasets for machine-learning models, as existing datasets can be costly and time-consuming to obtain, with varying effectiveness in improving model performance, and there is a need to identify datasets that add significant value to existing data.
Innovation Solution
A method using machine-learning techniques to assess and select candidate datasets by identifying unique entities in a multi-dimensional space, determining a merit attribute based on weights and the number of unique entities, and selecting datasets with a merit attribute greater than a threshold value for use in software applications.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If expensive and time-consuming datasets are obtained to improve machine-learning model performance, then model accuracy and sensitivity are improved, but cost and training time increase significantly
Solution Approach 1:
The system performs preliminary evaluation of candidate datasets before actual training by computing merit attributes that predict dataset quality and compatibility with target applications. This preliminary assessment identifies high-value datasets without requiring full training cycles, saving time while ensuring model performance improvement.
Solution Approach 2:
The patent replaces the traditional mechanical trial-and-error approach of dataset selection with an automated machine-learning-based evaluation system. The system uses algorithms to compute merit attributes and predict dataset effectiveness, substituting manual evaluation and random selection with intelligent, data-driven decision-making.
2Quantity of substance
If more datasets are collected and accumulated to improve model training quality, then dataset quantity and diversity increase, but acquisition cost and resource expenditure increase
Solution Approach 1:
The system extracts and evaluates only the most valuable subsets of candidate datasets by computing merit attributes for individual datasets and their combinations. Instead of collecting and processing all available data, the system identifies and selects only those datasets with highest predicted value for the specific target application.
Solution Approach 2:
The patent changes the evaluation parameter from simple dataset size or source to a computed merit attribute that reflects dataset quality, compatibility, and expected contribution to model performance. This parameter transformation enables identification of high-value datasets regardless of their size or acquisition cost.
3Quantity of substance
If datasets with similar numbers of entries are used for training, then data volume is standardized, but datasets may add different values or improvements to existing datasets and the trained model
Solution Approach 1:
The system evaluates each dataset's local quality and specific contribution to the target application rather than treating all datasets uniformly. The merit attribute computation considers dataset-specific characteristics, features, and compatibility with the target model, enabling identification of high-value datasets even with smaller entry counts.
4Measurement precision
If extensive dataset evaluation and selection processes are performed to identify valuable datasets, then dataset selection accuracy is improved, but computational complexity and processing time increase
Solution Approach 1:
The evaluation process is segmented into distinct computational stages: computing merit attributes for individual datasets, evaluating dataset combinations, and selecting optimal subsets for training. This segmentation allows the system to manage complexity by breaking down the overall evaluation task into smaller, more tractable sub-problems.
Data Source
AI summary
Techniques disclosed herein relate generally to evaluating and selecting candidate datasets for use by software applications, such as selecting candidate datasets for training machine-learning models used in software applications. Various machine-learning and other data science techniques are used to identify unique entities in a candidate dataset that are likely to be part of target entities for a software application. A merit attribute is then determined for the candidate dataset based on the number of unique entities that are likely to be part of the target entities, and weights associated with these unique entities. The merit attribute is used to identify the most efficient or most cost-effective candidate dataset for the software application.


