Variational Autoencoder Surrogate Modeling for High-Dimensional Uncertainty

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

High-dimensional data processing under uncertainty is computationally intensive, particularly in machine learning applications, due to factors like noise, data mismatch, and incomplete data, where existing uncertainty quantification techniques become prohibitively expensive.

Innovation Solution

A method using a variational autoencoder (VAE) for dimensionality reduction and polynomial chaos expansion (PCE) to map high-dimensional input data to a low-dimensional latent space, enabling uncertainty estimation without prior statistical assumptions, by sampling new data samples and learning their distributions to perform estimation under uncertainty such as missing values.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If uncertainty quantification techniques are applied to high-dimensional data, then data representation learning under uncertainty is improved, but computational cost becomes prohibitively expensive

Engineering Contradiction:
Improveuncertainty quantificationVSAvoidcomputational cost
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The patent applies dimensionality reduction by transforming high-dimensional input data into a lower-dimensional latent space using a variational autoencoder (VAE). This allows uncertainty quantification to be performed in the reduced-dimensional space, significantly lowering computational costs while preserving essential data characteristics and uncertainty information.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent creates a surrogate model that copies the essential statistical properties and uncertainty characteristics of the original high-dimensional data into a simplified low-dimensional representation. This surrogate model enables uncertainty quantification without requiring computationally intensive analysis of the full high-dimensional data.

Inventive Principle:
Principle #26Copying

2Measurement precision

If high-dimensional data is processed under uncertainty, then data representation learning is improved, but computational intensity increases

Engineering Contradiction:
Improvedata representationVSAvoidcomputational intensity
Core Design Contradiction:
Measurement precisionVSPower

Solution Approach 1:

The patent reduces the dimensionality of high-dimensional data through a VAE-based latent space representation. This transformation maintains the essential information and uncertainty structures while operating in a lower-dimensional space, thereby reducing computational intensity requirements for data representation learning.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent changes the parameter space from high-dimensional to low-dimensional by learning a compressed representation. This parameter transformation preserves the statistical properties needed for accurate data representation while significantly reducing the computational power required for processing.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If polynomial chaos expansion is used to map latent space to output distributions, then uncertainty estimation is improved, but model complexity increases

Engineering Contradiction:
Improveuncertainty estimationVSAvoidmodel complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the uncertainty estimation process into distinct components: the VAE for dimensionality reduction and the PCE for uncertainty quantification. This segmentation allows each component to be optimized independently, managing overall model complexity while improving uncertainty estimation reliability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a latent space representation as an intermediary between the high-dimensional input data and the uncertainty estimation model. This intermediary layer simplifies the relationship between inputs and outputs, enabling more reliable uncertainty estimation with manageable model complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20240135185A1High dimensional surrogate modeling for learning uncertainty
Publication Date: 2024.04.25 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US20240135185A1 patent drawing
  • US20240135185A1 patent drawing
  • US20240135185A1 patent drawing

AI summary

A method to determine data uncertainty is provided. The method receives a high dimensional data input and a corresponding data output. The method trains a variational autoencoder (VAE) with the high dimensional data input to learn a low dimensional latent space representation of the high dimensional data input. An encoder part of the VAE outputs a set of distributions of the high dimensional dataset in a latent space. The method samples new data samples in the latent space using the set of distributions outputs from the encoder part of the VAE. The method learns a polynomial chaos expansion to map the new data samples in the latent space to the corresponding data output to learn the set of distributions and their relation to perform estimation with high-dimensional dataset under uncertainty such as missing values by estimating the values using the set of distributions.