Rare-Event Risk Model Training with Variable-Based Data Augmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Machine learning algorithms for predicting rare events are less efficient due to unbalanced training data sets, where data for common events outnumber data for rare events, leading to poor performance in predicting rare occurrences.

Innovation Solution

A method and device for generating balanced training data by determining a subset of variables associated with a risk threshold, generating pseudo-random variations, and applying unsupervised learning to create three groups of data, with optional reclassification using multiple classification models to improve data balance and prediction accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If machine learning algorithms are trained on unbalanced training data sets with many common events and few rare events, then the algorithm can be trained efficiently with abundant data, but the prediction performance for rare events deteriorates

Engineering Contradiction:
Improvequantity of training dataVSAvoidprediction performance for rare events
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent applies synthetic data generation techniques to create artificial copies of rare event data. By generating synthetic training samples that replicate the characteristics of rare events, the method increases the effective quantity of rare event data without requiring additional real-world observations, thereby improving prediction performance while maintaining training efficiency

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent transforms the training data by applying parameter changes to balance the distribution between common and rare events. This involves modifying data sampling strategies, applying weighting schemes, or transforming features to ensure that rare events are adequately represented in the training set, thus resolving the contradiction between data quantity and prediction reliability

Inventive Principle:
Principle #35Parameter changes

2Ease of manufacture

If the training data set contains predominantly common events, then the data is abundant and easy to collect, but the algorithm becomes biased toward common events and fails to detect rare events

Engineering Contradiction:
Improveease of data collectionVSAvoiddetection accuracy of rare events
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent applies preliminary actions by pre-processing the training data to balance event representation before model training. This includes techniques such as oversampling rare events, undersampling common events, or applying stratified sampling to ensure that rare events are adequately represented from the outset, preventing algorithmic bias while maintaining ease of data collection

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces intermediary processing steps between data collection and model training. These intermediaries include data balancing algorithms, synthetic data generation modules, or reweighting mechanisms that mediate between the abundance of common event data and the scarcity of rare event data, ensuring balanced representation without complicating the original data collection process

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If more rare event data is collected to improve prediction accuracy, then prediction performance improves, but the complexity and cost of data collection increases

Engineering Contradiction:
Improveprediction accuracy for rare eventsVSAvoidcomplexity of data collection system
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent eliminates the need for complex data collection systems by generating synthetic copies of rare event data through computational methods. This approach achieves improved prediction accuracy for rare events without requiring additional sensors, monitoring systems, or data collection infrastructure, thereby avoiding increased device complexity

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces mechanical or physical data collection systems with computational data generation methods. Instead of using complex monitoring equipment to collect more rare event data, the system uses algorithms to synthesize rare event samples, substituting physical data collection infrastructure with software-based solutions that reduce overall system complexity

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentEP4488893A1Generation of training data for machine learning of a model for predicting a risk of occurrence of a rare event
Publication Date: 2025.01.08 COMMISSARIAT A LENERGIE ATOMIQUE ET AUX ENERGIES ALTERNATIVES
  • EP4488893A1 patent drawingFigure 1
  • EP4488893A1 patent drawingFigure 2
  • EP4488893A1 patent drawingFigure 3

AI summary

The invention relates to a method and device for generating training data for machine learning of a model for predicting the risk of the occurrence of a rare event, said training data being formed of vectors of variables representing a state of at least one system monitored before the occurrence of said rare event, from input training data (4) comprising a first initial set (6) of data associated with the absence of the occurrence of said rare event, and a second initial set (8) of data associated with the presence of the occurrence of said rare event.The device implements determination modules (20), by a machine learning method, of at least a subset of variables associated with a risk of occurrence of said rare event greater than a risk threshold, of generation (22) of a third set of data by pseudo-random variations on said subset of data, allowing to obtain an augmented training database.