Multimodal Time-Series Augmentation for Industrial Forecasting

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Complex neural networks used for predicting industrial time-dependent processes face challenges such as overfitting due to limited availability of large quantities of real-world data, which is expensive and difficult to collect, and are affected by nonstationarities in process dynamics, making it hard to implement effective training for long-term forecasting.

Innovation Solution

A data-driven generative model is used to generate synthetic samples that expand the training dataset by learning a joint time-dependent representation of condition parameters and key performance indicators (KPIs), allowing for data augmentation and improved generalization of predictive models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If more real-world training data is collected to train complex neural networks, then model accuracy improves, but data collection cost and time increase significantly

Engineering Contradiction:
Improvemodel accuracyVSAvoiddata collection time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent generates synthetic training data that copies the statistical properties and temporal dependencies of real-world industrial time series data. A generative model learns from limited real data and produces artificial samples that preserve autocorrelation, multivariate relationships, and process dynamics, enabling effective model training without extensive real data collection

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent transforms the training approach by changing from using only raw real-world data to using augmented training data that includes synthetically generated samples. This parameter change in data composition allows the system to achieve better generalization with smaller real data sets by combining real and synthetic samples

Inventive Principle:
Principle #35Parameter changes

2Reliability

If more real-world training data is collected to train complex neural networks, then model accuracy improves, but collection cost increases

Engineering Contradiction:
Improvemodel accuracyVSAvoiddata collection cost
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent creates synthetic copies of real-world industrial time series data that preserve essential statistical properties including autocorrelation, multivariate dependencies, and process dynamics. This copying approach eliminates the need for expensive real data collection while providing sufficient training material for complex neural networks

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system uses the limited real-world data available to train a generative model, which then serves itself to produce additional training samples. This self-service mechanism allows the system to bootstrap effective training from minimal real data without requiring external data collection resources

Inventive Principle:
Principle #25Self-service

3Ease of operation

If standard machine learning models are used for long-term forecasting, then implementation is simple, but performance degrades due to nonstationarities in process dynamics

Engineering Contradiction:
Improveimplementation simplicityVSAvoidforecasting performance
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent employs a generative model that learns and captures the dynamic, nonstationary behavior of industrial processes. The model adapts to changing process dynamics by learning from the temporal patterns in the data, enabling it to generate realistic synthetic samples even when process characteristics evolve over time

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20230045548A1Augmentation of multimodal time series data for training machine-learning models
Publication Date: 2023.02.09 BASF SE
  • US20230045548A1 patent drawing
  • US20230045548A1 patent drawing
  • US20230045548A1 patent drawing

AI summary

The present invention relates to training predictive data-driven model for predicting an industrial time dependent process. A data driven generative model is introduced for modelling and generating complex sequential data comprising multiple modalities, by learning a joint time-dependent representation of the different modalities. The model may be configured to handle any combination of missing modalities, which enables conditional generation based on known modalities, providing a high degree of control over the properties of the generated sequences.