Rare-Event Risk Model Training with Variable-Based Data Augmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Machine learning algorithms for predicting rare events are less efficient due to unbalanced training data sets, where data for common events outnumber data for rare events, leading to poor performance in predicting rare occurrences.
Innovation Solution
A method and device for generating balanced training data by determining a subset of variables associated with a risk threshold, generating pseudo-random variations, and applying unsupervised learning to create three groups of data, with optional reclassification using multiple classification models to improve data balance and prediction accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If machine learning algorithms are trained on unbalanced training data sets with many common events and few rare events, then the algorithm can be trained efficiently with abundant data, but the prediction performance for rare events deteriorates
Solution Approach 1:
The patent applies synthetic data generation techniques to create artificial copies of rare event data. By generating synthetic training samples that replicate the characteristics of rare events, the method increases the effective quantity of rare event data without requiring additional real-world observations, thereby improving prediction performance while maintaining training efficiency
Solution Approach 2:
The patent transforms the training data by applying parameter changes to balance the distribution between common and rare events. This involves modifying data sampling strategies, applying weighting schemes, or transforming features to ensure that rare events are adequately represented in the training set, thus resolving the contradiction between data quantity and prediction reliability
2Ease of manufacture
If the training data set contains predominantly common events, then the data is abundant and easy to collect, but the algorithm becomes biased toward common events and fails to detect rare events
Solution Approach 1:
The patent applies preliminary actions by pre-processing the training data to balance event representation before model training. This includes techniques such as oversampling rare events, undersampling common events, or applying stratified sampling to ensure that rare events are adequately represented from the outset, preventing algorithmic bias while maintaining ease of data collection
Solution Approach 2:
The patent introduces intermediary processing steps between data collection and model training. These intermediaries include data balancing algorithms, synthetic data generation modules, or reweighting mechanisms that mediate between the abundance of common event data and the scarcity of rare event data, ensuring balanced representation without complicating the original data collection process
3Reliability
If more rare event data is collected to improve prediction accuracy, then prediction performance improves, but the complexity and cost of data collection increases
Solution Approach 1:
The patent eliminates the need for complex data collection systems by generating synthetic copies of rare event data through computational methods. This approach achieves improved prediction accuracy for rare events without requiring additional sensors, monitoring systems, or data collection infrastructure, thereby avoiding increased device complexity
Solution Approach 2:
The patent replaces mechanical or physical data collection systems with computational data generation methods. Instead of using complex monitoring equipment to collect more rare event data, the system uses algorithms to synthesize rare event samples, substituting physical data collection infrastructure with software-based solutions that reduce overall system complexity
Data Source
Figure 1
Figure 2
Figure 3
AI summary
The invention relates to a method and device for generating training data for machine learning of a model for predicting the risk of the occurrence of a rare event, said training data being formed of vectors of variables representing a state of at least one system monitored before the occurrence of said rare event, from input training data (4) comprising a first initial set (6) of data associated with the absence of the occurrence of said rare event, and a second initial set (8) of data associated with the presence of the occurrence of said rare event.The device implements determination modules (20), by a machine learning method, of at least a subset of variables associated with a risk of occurrence of said rare event greater than a risk threshold, of generation (22) of a third set of data by pseudo-random variations on said subset of data, allowing to obtain an augmented training database.