Synthetic Training Data for Neural Networks Predicting High-Pollution Events

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing neural networks struggle to accurately predict rare high-pollution events due to insufficient training data, leading to poorer predictions for these critical situations.

Innovation Solution

Generate synthetic pollutant concentration measurement series by modifying measured variables using a transmission model, incorporating chemical conversion processes and pollutant emission data to expand the training dataset for neural networks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If historical pollutant concentration data is used to train the neural network, then the average pollutant concentration can be predicted with sufficient accuracy, but the prediction for rare high-stress events deteriorates due to insufficient training data

Engineering Contradiction:
Improveprediction accuracy for average pollutant concentrationVSAvoidprediction accuracy for rare high-stress events
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent creates synthetic copies of rare high-pollution events by generating artificial measurement series that replicate the characteristics of actual high-stress events. These synthetic data copies are then added to the training dataset, allowing the neural network to learn from multiple examples of rare events without requiring additional real-world occurrence data.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent modifies parameters of the training data by adjusting the frequency and weighting of high-stress event data. Specifically, the method increases the representation of rare events in the training dataset through synthetic data generation and resampling techniques, changing the data distribution parameters to prioritize learning from rare but critical events.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If data or measurement series are weighted differently to improve rare event prediction, then prediction for high-pollution events improves, but prediction for average pollution deteriorates

Engineering Contradiction:
Improveprediction accuracy for high-pollution eventsVSAvoidprediction accuracy for average pollutant concentration
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

Instead of simply weighting existing data, the patent generates synthetic copies of high-pollution events that are indistinguishable from real events. This approach allows the model to treat all training data uniformly while still achieving improved rare event prediction, avoiding the need for explicit weighting that would bias average predictions.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent performs preliminary data preparation by generating synthetic high-stress event data before training the neural network. This preliminary action ensures that the training dataset already contains sufficient representations of rare events, eliminating the need for post-training weighting adjustments that would compromise average prediction accuracy.

Inventive Principle:
Principle #10Preliminary action

3Quantity of substance

If a fully model-based approach is used to calculate pollutant emissions and concentrations, then no training data is needed, but the computational effort is significant and known methods provide values that are too low

Engineering Contradiction:
Improvetraining data requirementsVSAvoidcomputational effort and model complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary approach by combining neural networks with simplified emission models. The neural network handles the complex concentration prediction based on emission data, while a simplified transmission model provides the emission calculations. This intermediary structure reduces the computational burden compared to fully model-based approaches while avoiding the need for extensive training data.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentEP4107522B1Computer-assisted method for generating training data for a neural network for predicting a concentration of pollutants at a measuring station
Publication Date: 2025.07.02 SIEMENS AG
  • EP4107522B1 patent drawingFigure 1
  • EP4107522B1 patent drawing

AI summary

The invention relates to a computer-assisted method for generating training data for a neural network, wherein the neural network is configured to detect a concentration of pollutants at at least one measuring station from at least one emission of pollutants. For this purpose in particular, a synthetic measurement series is used as training data by changing ∆C of a value C 0 of a provided measured measurement series of the pollutant concentrations, wherein the change ∆C takes place using the relative change ∆I/I 0 of values of pollutant emissions calculated by means of a transmission model I 0, I 1. The invention further relates to a computer-assisted method for training a neural network and to a method for detecting a concentration of pollutants using the neural network trained in this way.