Dataset Merit Evaluation for Machine Learning Model Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The challenge lies in efficiently selecting and evaluating candidate datasets for machine-learning models, as existing datasets can be costly and time-consuming to obtain, with varying effectiveness in improving model performance, and there is a need to identify datasets that add significant value to existing data.

Innovation Solution

A method using machine-learning techniques to assess and select candidate datasets by identifying unique entities in a multi-dimensional space, determining a merit attribute based on weights and the number of unique entities, and selecting datasets with a merit attribute greater than a threshold value for use in software applications.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If expensive and time-consuming datasets are obtained to improve machine-learning model performance, then model accuracy and sensitivity are improved, but cost and training time increase significantly

Engineering Contradiction:
Improvemodel performanceVSAvoidtraining time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary evaluation of candidate datasets before actual training by computing merit attributes that predict dataset quality and compatibility with target applications. This preliminary assessment identifies high-value datasets without requiring full training cycles, saving time while ensuring model performance improvement.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces the traditional mechanical trial-and-error approach of dataset selection with an automated machine-learning-based evaluation system. The system uses algorithms to compute merit attributes and predict dataset effectiveness, substituting manual evaluation and random selection with intelligent, data-driven decision-making.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Quantity of substance

If more datasets are collected and accumulated to improve model training quality, then dataset quantity and diversity increase, but acquisition cost and resource expenditure increase

Engineering Contradiction:
Improvedataset quantityVSAvoidresource expenditure
Core Design Contradiction:
Quantity of substanceVSLoss of energy

Solution Approach 1:

The system extracts and evaluates only the most valuable subsets of candidate datasets by computing merit attributes for individual datasets and their combinations. Instead of collecting and processing all available data, the system identifies and selects only those datasets with highest predicted value for the specific target application.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent changes the evaluation parameter from simple dataset size or source to a computed merit attribute that reflects dataset quality, compatibility, and expected contribution to model performance. This parameter transformation enables identification of high-value datasets regardless of their size or acquisition cost.

Inventive Principle:
Principle #35Parameter changes

3Quantity of substance

If datasets with similar numbers of entries are used for training, then data volume is standardized, but datasets may add different values or improvements to existing datasets and the trained model

Engineering Contradiction:
Improvedata volumeVSAvoiddataset value
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The system evaluates each dataset's local quality and specific contribution to the target application rather than treating all datasets uniformly. The merit attribute computation considers dataset-specific characteristics, features, and compatibility with the target model, enabling identification of high-value datasets even with smaller entry counts.

Inventive Principle:
Principle #3Local quality

4Measurement precision

If extensive dataset evaluation and selection processes are performed to identify valuable datasets, then dataset selection accuracy is improved, but computational complexity and processing time increase

Engineering Contradiction:
Improvedataset evaluation accuracyVSAvoidevaluation process complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The evaluation process is segmented into distinct computational stages: computing merit attributes for individual datasets, evaluating dataset combinations, and selecting optimal subsets for training. This segmentation allows the system to manage complexity by breaking down the overall evaluation task into smaller, more tractable sub-problems.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11704598B2Machine-learning techniques for evaluating suitability of candidate datasets for target applications
Publication Date: 2023.07.18 ADOBE INC
  • US11704598B2 patent drawing
  • US11704598B2 patent drawing
  • US11704598B2 patent drawing

AI summary

Techniques disclosed herein relate generally to evaluating and selecting candidate datasets for use by software applications, such as selecting candidate datasets for training machine-learning models used in software applications. Various machine-learning and other data science techniques are used to identify unique entities in a candidate dataset that are likely to be part of target entities for a software application. A merit attribute is then determined for the candidate dataset based on the number of unique entities that are likely to be part of the target entities, and weights associated with these unique entities. The merit attribute is used to identify the most efficient or most cost-effective candidate dataset for the software application.