Representative Informative Dataset for ML Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Machine learning models face increased computational complexity and reduced prediction accuracy due to the presence of redundant and irrelevant data in training datasets, making it difficult to identify predictive data.
Innovation Solution
A method to create a representative and informative (RAI) dataset by selecting informative attributes from an initial training dataset using model descriptions, removing duplicates, and applying clustering algorithms to generate representative data records, thereby reducing the dataset size and focusing the training process.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If the training data contains redundant and irrelevant data, then the training data size is large, but the computational complexity increases and prediction accuracy decreases
Solution Approach 1:
The patent extracts only the predictive features from the training data by identifying and selecting attributes that have been proven to improve prediction accuracy. This extraction process removes redundant and irrelevant data while retaining only the essential information needed for training, thereby reducing computational complexity without sacrificing data quantity needed for accurate models.
Solution Approach 2:
The patent introduces an intermediary feature selection process that acts as a mediator between the raw training data and the model training process. This intermediary layer identifies and filters out non-predictive attributes, creating a streamlined dataset that maintains sufficient information for accurate predictions while significantly reducing the computational burden during training.
2Quantity of substance
If the training data contains redundant and irrelevant data, then the training data size is large, but the prediction accuracy decreases due to interference
Solution Approach 1:
The patent extracts and retains only the predictive features that have been validated to improve prediction accuracy. By removing non-predictive attributes from the training data, the system maintains sufficient data quantity for accurate models while eliminating the interference from redundant information that would otherwise reduce prediction reliability.
Solution Approach 2:
The patent changes the parameters of the training data by transforming it into a reduced feature space that contains only the most predictive attributes. This parameter transformation maintains the essential information needed for accurate predictions while eliminating noise that would interfere with model performance.
3Device complexity
If the training data is reduced to remove redundant data, then computational complexity decreases, but data representativeness may be compromised
Solution Approach 1:
The patent employs feedback mechanisms to continuously monitor and validate the quality of the reduced training data. By feedback-driven feature selection and validation processes, the system ensures that the reduced dataset maintains sufficient representativeness and predictive power, preventing compromise of data quality while achieving lower computational complexity.
Solution Approach 2:
The patent performs preliminary feature selection and validation actions before model training begins. This preliminary action identifies and retains only the most predictive features in advance, ensuring that the reduced training data maintains representativeness and predictive accuracy while significantly reducing computational complexity during the actual training process.
Data Source
AI summary
In some aspects, techniques for creating representative and informative training datasets for the training of machine-learning models are provided. For example, a risk assessment system can receive a risk assessment query for a target entity. The risk assessment system can compute an output risk indicator for the target entity by applying a machine learning model to values of informative attributes associated with the target entity. The machine learning model may be trained using training samples selected from a representative and informative (RAI) dataset. The RAI dataset can be created by determining the informative attributes based on attributes used by a set of models and further extracting representative data records from an initial training dataset based on the determined informative attributes. The risk assessment system can transmit a responsive message including the output risk indicator for use in controlling access of the target entity to an interactive computing environment.


