Automated Stochastic DNN Architecture Search with Irregular Beliefs

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing automated machine learning (AutoML) frameworks for deep neural networks (DNNs) face challenges in efficiently identifying the best probabilistic model and hyperparameters for stochastic DNNs, particularly when underlying data statistics and uncertainty are unspecified, leading to suboptimal performance and increased exploration time.

Innovation Solution

The system employs an automated variational Bayesian inference framework that explores irregular combinations of posterior, prior, and likelihood beliefs, along with mismatched discrepancy measures, using hypergradient methods to optimize stochastic DNNs for unspecified datasets, allowing for heterogenous and mismatched pairings of distributions and connectivity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If homogeneous normal distribution is used for latent representations, then computational convenience is improved, but model accuracy deteriorates when data statistics are unspecified

Engineering Contradiction:
Improvecomputational convenienceVSAvoidmodel accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent changes the distributional parameters of latent representations from homogeneous normal distribution to heterogeneous distributions (normal, Laplace, Cauchy, logistic, Gumbel, student-t, uniform, exponential, hyper-exponential) selected based on data characteristics. This allows the model to adapt to unspecified data statistics while maintaining computational tractability through automated selection mechanisms.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces dynamic selection of distribution types for latent representations based on automated analysis of data statistics. The system dynamically determines which distribution family (normal, Laplace, Cauchy, etc.) best fits the underlying data characteristics, enabling adaptability rather than static homogeneous assumptions.

Inventive Principle:
Principle #15Dynamics

2Adaptability or versatility

If automated exploration of different DNN architectures is performed, then adaptability is improved, but exploration time increases

Engineering Contradiction:
ImproveadaptabilityVSAvoidexploration time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent applies local quality by allowing different latent representations to have different distribution types (normal, Laplace, Cauchy, logistic, Gumbel, student-t, uniform, exponential, or hyper-exponential) based on their specific data characteristics. This localized adaptation reduces the need for exhaustive global architecture search while maintaining optimality for each specific representation.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system performs self-service by automatically analyzing data statistics and selecting appropriate distribution types for latent representations without requiring manual architecture design or extensive hyperparameter tuning. The automated selection mechanism serves the model configuration needs internally, reducing external exploration requirements.

Inventive Principle:
Principle #25Self-service

Data Source

PatentEP4396730B1Automated variational inference using stochastic models with irregular beliefs
Publication Date: 2025.09.24 MITSUBISHI ELECTRIC CORP
  • EP4396730B1 patent drawingFigure 1A~1C
  • EP4396730B1 patent drawingFigure 2
  • EP4396730B1 patent drawingFigure 3

AI summary

A system and method for automated construction of a stochastic deep neural network (DNN) architecture is provided. The framework of invention automatically searches for most relevant stochastic modes underlaying datasets for variational Bayesian inference. The invention provides a way to use heterogenous, irregular, and mismatched beliefs in stochastic sampling for intermediate representation in DNNs with a capability of an automatically tuning mechanism of posterior, prior, and likelihood models to enable accurate generative models and uncertainty models for machine learning tasks. The system further allows adjustable discrepancy measure to regularize intermediate representation by variants of divergence metrics including Renyi's alpha, beta, and gamma divergences. The invention enables diverse mixture combinations of stochastic models for misspecified and unspecified probabilistic relations in an automatic fashion. Accordingly, the representation capability of variational autoencoders, variational information bottlenecks, denoising diffusion probabilistic models and other stochastic DNNs are improved.