Conditional Generative Model for Heterogeneous Data Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current machine learning techniques face challenges in modeling complex probability distributions with diverse and heterogeneous data sets, particularly in fields like health informatics where labeled data is scarce, making it difficult to predict outcomes in clinical trials and other applications.

Innovation Solution

The development of conditional generative models, specifically combining probabilistic models like Conditional Restricted Boltzmann Machines (CRBMs) with point prediction models, to generate samples and refine time-series data, enabling the use of heterogeneous and unlabeled data for training stochastic unsupervised machine-learning models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If traditional machine learning techniques are used to model probability distributions with heterogeneous data, then the model structure remains simple, but the ability to handle diverse and unlabeled data is insufficient

Engineering Contradiction:
Improveability to handle heterogeneous and unlabeled dataVSAvoidmodel structure complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent combines probabilistic models (CRBMs) with point prediction models into a unified conditional generative model framework. This merging allows the system to simultaneously handle heterogeneous data types and unlabeled data while maintaining a coherent model structure that can capture complex probability distributions.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The conditional generative model uses a composite architecture that integrates different model components (probabilistic CRBM layer and point prediction layer) with distinct functionalities. This composite structure enables the system to process diverse data types effectively while managing complexity through modular design.

Inventive Principle:
Principle #40Composite materials

2Measurement precision

If more labeled data is collected to improve prediction accuracy, then the predictive capability improves, but the cost and time required for data collection and labeling increases

Engineering Contradiction:
Improveprediction accuracyVSAvoiddata collection and labeling time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The conditional generative model performs self-service by generating synthetic training data that automatically supplements the available labeled data. The model uses its learned probability distributions to create realistic sample data, eliminating the need for manual data collection and labeling while improving prediction accuracy through enhanced training datasets.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs preliminary data generation before the actual prediction task. By pre-generating synthetic training data and pre-training the model on this augmented dataset, the system prepares in advance to improve prediction accuracy without incurring time costs during the actual prediction phase.

Inventive Principle:
Principle #10Preliminary action

3Quantity of substance

If synthetic data is generated to supplement training data, then the amount of training data increases, but the quality and realism of the generated data must be maintained

Engineering Contradiction:
Improveamount of training dataVSAvoidquality and realism of generated data
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The training process incorporates feedback mechanisms where the generated synthetic data is evaluated and used to refine the model parameters. The CRBM model learns from the feedback signal provided by the reconstruction error and probability distribution matching, continuously improving the quality and realism of generated data while increasing the training dataset size.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system adjusts model parameters (such as temperature parameters in sampling, regularization strengths, and network architecture parameters) to optimize the balance between quantity and quality of generated data. By carefully tuning these parameters, the model generates realistic synthetic data that maintains high quality while substantially increasing the available training data volume.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20240169188A1Systems and Methods for Training Conditional Generative Models
Publication Date: 2024.05.23 UNLEARN AI INC
  • US20240169188A1 patent drawing
  • US20240169188A1 patent drawing
  • US20240169188A1 patent drawing

AI summary

Systems and techniques for adjusting experiment parameters are illustrated. One embodiment includes a method that defines a joint distribution, wherein the joint distribution corresponds to a combination of a probabilistic model and a point prediction model, and wherein the point prediction model is configured to obtain a measurement of regression accuracy. The method derives an energy function for the joint distribution. The method obtains, from the energy function for the joint distribution, an approximation for a conditional distribution, wherein an output of the point prediction model is a parameter of the approximation. The method determines, from a loss function, at least one training parameter. The method trains the probabilistic based on the at least one parameter to operate as a conditional generative model, wherein the trained probabilistic model follows the conditional distribution. The method applies the trained probabilistic model to a dataset corresponding to a randomized trial.