Deep Neural Network Training with Conditional Mutual Information

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Deep neural networks (DNNs) face challenges in achieving high accuracy and robustness while maintaining interpretability, particularly due to their complex architecture and reliance on error rates as the primary performance metric.

Innovation Solution

The proposed solution involves training DNNs using a method that optimizes both the error function and a network mapping function, which represents predicted label distribution geometry properties such as intra-class concentration and inter-class separation. This approach enhances the model's robustness and interpretability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If deep neural networks are trained using traditional error function optimization only, then training simplicity is maintained, but model robustness and interpretability deteriorate

Engineering Contradiction:
Improvemodel robustnessVSAvoidtraining complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent combines the error function optimization with network mapping function optimization into a unified training framework. By merging these two objectives, the model simultaneously achieves high accuracy through error minimization and improved robustness/interpretability through geometry property optimization, resolving the contradiction between model reliability and training complexity.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces new parameters related to network mapping geometry properties (such as intra-class concentration and inter-class separation) alongside traditional error metrics. By changing the optimization parameters to include both error function values and mapping geometry properties, the training process achieves better robustness without excessive complexity increase.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If deep neural networks focus on error rate minimization, then prediction accuracy improves, but interpretability and geometric understanding of predictions worsen

Engineering Contradiction:
Improveprediction accuracyVSAvoidgeometric information loss
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent performs preliminary optimization of network mapping geometry properties during the training phase, before deployment. By pre-shaping the prediction distribution geometry to have desirable properties (such as proper concentration and separation), the model maintains both accuracy and interpretability, preventing geometric information loss that would otherwise occur.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements feedback mechanisms where the network mapping function's geometry properties are continuously evaluated during training and used to adjust the optimization process. This feedback loop ensures that both accuracy metrics and geometric properties are maintained, preventing the loss of geometric information while pursuing accuracy improvements.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20250086455A1Methods and systems for conditional mutual information constrained deep learning
Publication Date: 2025.03.13 MULTICOM TECH INC
  • US20250086455A1 patent drawing
  • US20250086455A1 patent drawing
  • US20250086455A1 patent drawing

AI summary

A system, method and computer program product for training a deep neural network. The deep neural network can be trained using a learning process that is defined to optimize both an error function of the deep neural network as well as a network mapping function of the deep neural network. The network mapping function can represent a predicted label distribution geometry property of the deep neural network. This learning process can improve the accuracy of the trained deep neural network model as well as its robustness again adversarial attacks. Optimizing the network mapping function can also provide increased insight into the operation of the trained deep neural network model, which may promote increased interpretability of the trained model and thus encourage uptake of the trained model.