Neural Network Training with Extremal Entropy Loss

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine learning systems that use entropy-based loss functions often overlook outliers in the training data, leading to models that do not generalize well due to overfitting typical examples.

Innovation Solution

A multi-layer node network is trained to minimize the worst-case error instead of average error, by identifying the input category with the maximum difference between expected and actual output probability distributions and adjusting the network parameters to minimize this maximum difference.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If entropy-based loss functions are used to train the network, then the average error is minimized, but the worst-case error (outliers) is overlooked and model generalization deteriorates

Engineering Contradiction:
Improveaverage errorVSAvoidmodel generalization
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent changes the loss function parameter from entropy (average error) to extremal entropy (worst-case error). Specifically, it uses the maximum entropy among all output nodes as the loss function, which transforms the optimization target from minimizing average error to minimizing the worst-case error. This parameter change allows the model to focus on outlier cases that were previously overlooked.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

Instead of minimizing the average error (sum of errors across all samples), the patent inverts the approach by maximizing the minimum error (focusing on the worst-case scenario). The loss function is defined as the maximum entropy among all output nodes, which inverts the conventional wisdom of averaging errors and instead amplifies the most problematic cases for optimization.

Inventive Principle:
Principle #13The other way round (Inversion)

2Productivity

If typical examples dominate the training data, then the network learns common patterns well, but outlier examples are stranded and not learned properly

Engineering Contradiction:
Improvelearning efficiencyVSAvoidoutlier detection
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent changes the aggregation method for computing loss from summation (average) to maximization (extremal). By using the maximum entropy among output nodes as the loss function, it transforms the training objective to focus on the worst-performing sample rather than the average, thereby giving outliers sufficient gradient magnitude to be learned properly despite their small proportion in the dataset.

Inventive Principle:
Principle #35Parameter changes

3Loss of time

If the optimization algorithm uses average error minimization, then training converges quickly on typical cases, but gradients for outliers become too small and training terminates prematurely

Engineering Contradiction:
Improvetraining timeVSAvoidgradient magnitude for outliers
Core Design Contradiction:
Loss of timeVSMeasurement precision

Solution Approach 1:

The patent changes the loss function from average error (sum of entropies) to extremal error (maximum entropy). This parameter change ensures that the gradient magnitude is determined by the worst-case sample rather than being diluted by averaging, preventing premature termination and ensuring adequate gradient signals for outlier learning.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250068912A1Training network to minimize worst-case error
Publication Date: 2025.02.27 AMAZON COM SERVICES LLC
  • US20250068912A1 patent drawing
  • US20250068912A1 patent drawing
  • US20250068912A1 patent drawing

AI summary

Some embodiments provide a method for configuring a machine-trained (MT) network that includes multiple configurable weights to train. The method propagates a set of inputs through the MT network to generate a set of output probability distributions. Each input has a corresponding expected output probability distribution. The method calculates a value of a continuously-differentiable loss function that includes a term approximating an extremum function of the difference between the expected output probability distributions and generated set of output probability distributions. The method trains the weights by back-propagating the calculated value of the continuously-differentiable loss function.