Learning Device Using Fisher Information Matrix Eigenvectors
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods face difficulties in improving generalization performance against adversarial examples, particularly due to the use of random numbers as initial values when searching for optimal models using proxy losses in TRADES.
Innovation Solution
A learning device that acquires data with labels and learns a model representing the probability distribution of the label using an eigenvector corresponding to the maximum eigenvalue in the Fisher information matrix, allowing for robustness against adversarial examples.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If random numbers are used as initial values when searching for optimal models using proxy losses in TRADES, then differentiation can be performed, but generalization performance against adversarial examples deteriorates
Solution Approach 1:
The patent changes the parameter of initial values from random numbers to eigenvectors corresponding to maximum eigenvalues in the Fisher information matrix. This parameter change enables the model to achieve both differentiability and improved generalization performance against adversarial examples, resolving the contradiction between performing differentiation and achieving robust generalization.
2Reliability
If conventional TRADES with proxy losses are used, then model training is enabled, but robustness to adversarial examples deteriorates
Solution Approach 1:
The patent introduces the Fisher information matrix as an intermediary component that bridges the gap between conventional TRADES training and adversarial robustness. By using the eigenvectors of the Fisher information matrix as initial values, the method enables robustness without requiring complete redesign of the training framework, thus improving robustness while managing training complexity.
Data Source
AI summary
A learning device includes processing circuitry configured to acquire data with a label to be predicted, and learn a model that represents probability distribution of the label of the acquired data using an eigenvector corresponding to a maximum eigenvalue in a Fisher information matrix for the data in the model.


