Neural Network Training with Extremal Entropy Loss
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing machine learning systems that use entropy-based loss functions often overlook outliers in the training data, leading to models that do not generalize well due to overfitting typical examples.
Innovation Solution
A multi-layer node network is trained to minimize the worst-case error instead of average error, by identifying the input category with the maximum difference between expected and actual output probability distributions and adjusting the network parameters to minimize this maximum difference.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If entropy-based loss functions are used to train the network, then the average error is minimized, but the worst-case error (outliers) is overlooked and model generalization deteriorates
Solution Approach 1:
The patent changes the loss function parameter from entropy (average error) to extremal entropy (worst-case error). Specifically, it uses the maximum entropy among all output nodes as the loss function, which transforms the optimization target from minimizing average error to minimizing the worst-case error. This parameter change allows the model to focus on outlier cases that were previously overlooked.
Solution Approach 2:
Instead of minimizing the average error (sum of errors across all samples), the patent inverts the approach by maximizing the minimum error (focusing on the worst-case scenario). The loss function is defined as the maximum entropy among all output nodes, which inverts the conventional wisdom of averaging errors and instead amplifies the most problematic cases for optimization.
2Productivity
If typical examples dominate the training data, then the network learns common patterns well, but outlier examples are stranded and not learned properly
Solution Approach 1:
The patent changes the aggregation method for computing loss from summation (average) to maximization (extremal). By using the maximum entropy among output nodes as the loss function, it transforms the training objective to focus on the worst-performing sample rather than the average, thereby giving outliers sufficient gradient magnitude to be learned properly despite their small proportion in the dataset.
3Loss of time
If the optimization algorithm uses average error minimization, then training converges quickly on typical cases, but gradients for outliers become too small and training terminates prematurely
Solution Approach 1:
The patent changes the loss function from average error (sum of entropies) to extremal error (maximum entropy). This parameter change ensures that the gradient magnitude is determined by the worst-case sample rather than being diluted by averaging, preventing premature termination and ensuring adequate gradient signals for outlier learning.
Data Source
AI summary
Some embodiments provide a method for configuring a machine-trained (MT) network that includes multiple configurable weights to train. The method propagates a set of inputs through the MT network to generate a set of output probability distributions. Each input has a corresponding expected output probability distribution. The method calculates a value of a continuously-differentiable loss function that includes a term approximating an extremum function of the difference between the expected output probability distributions and generated set of output probability distributions. The method trains the weights by back-propagating the calculated value of the continuously-differentiable loss function.


