Training device, training method, and program
By optimizing C*yclic networks using a learning device that minimizes the maximum eigenvalue and maximizes the eigenvalues of the Hessian matrix, the accuracy of anomaly detection and other applications is enhanced.
Patent Information
- Application Number
- PCT/JP2024/001905
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-23
- Publication Date
- 2025-07-31
AI Technical Summary
The theoretical analysis of C*yclic networks is insufficient, leading to decreased accuracy in anomaly detection techniques, as the performance of these networks has not been adequately improved.
A learning device that learns neural networks with parameters represented by linear combinations of continuous functions on a compact Hausdorff space, minimizing the maximum eigenvalue of the Hessian matrix and maximizing the entire eigenvalues of the Hessian matrix through a regularization term, to optimize the C*yclic networks.
This approach enhances the generalization performance of C*yclic networks, improving inference accuracy in anomaly detection and other applications such as regression, classification, and data generation.
Smart Images

Figure JP2024001905_31072025_PF_FP_ABST
Abstract
Description
Learning device, learning method, and program
[0001] The present disclosure relates to a learning device, a learning method, and a program.
[0002] Neural networks are applied in a wide range of fields, and research into them is actively being conducted. In recent years, research into the continuation of neural networks has also been conducted, and C is being proposed as a framework for generalizing existing neural networks. * A ring network has been proposed (Non-Patent Document 1). * The ring network is a type of continuation of neural networks, and was proposed with the aim of achieving efficient learning by simultaneously training an infinite number of neural networks.
[0003] C * The circular network is expected to be applied to a wide range of fields, for example, in anomaly detection technology for automating the operation of ICT systems.
[0004] Yuka Hashimoto, Zhao Wang, and Tomoko Matsui. C*-algebra net: A new approach generalizing neural network parameters to C*-algebra. In Proceedings of the 39th International Conference on Machine Learning (ICML), 2022.
[0005] However, C * The theoretical analysis of ring networks is still insufficient, and it is unclear what kind of learning is necessary to * There are many unknowns as to whether the performance of the ring network will improve. * If the performance of the ring network cannot be improved, then when it is applied to anomaly detection technology, for example, the accuracy of anomaly detection will decrease.
[0006] The present disclosure has been made in view of the above points, * The purpose is to improve the performance of the ring network.
[0007] A learning device according to one aspect of the present disclosure is a learning device that trains a neural network having parameters whose elements are functions expressed as a linear combination of a finite number of continuous functions in a compact Hausdorff space, and includes: an input unit that inputs a training dataset for training the neural network; a Hessian matrix for the finite number of continuous functions, which is a function expressed as the square of the difference between the neural network with fixed values in the compact Hausdorff space and a function obtained by integrating the neural network using a predetermined measure in the compact Hausdorff space; and a model learning unit that uses the training dataset to learn the parameters so as to minimize the largest eigenvalue of the Hessian matrix at values obtained by integrating the finite number of continuous functions using the measure, and to minimize a loss function that includes a regularization term for maximizing all the eigenvalues.
[0008] C * The performance of the ring network can be improved.
[0009] Fig. 1 is a diagram illustrating an example of a hardware configuration of a learning device according to an embodiment; Fig. 2 is a diagram illustrating an example of a functional configuration of a learning device according to an embodiment; Fig. 3 is a flowchart illustrating an example of a model learning process according to an embodiment; Fig. 4 is a diagram illustrating an example of an accuracy rate for test data.
[0010] An embodiment of the present invention will be described in detail below with reference to the drawings. * A learning device 10 that can improve the performance of a ring network will be described. * The performance of the ring network is * It refers to the generalization performance acquired when optimizing a ring network through learning. Generalization performance is the ability or capability to make correct inferences about unknown data that is not included in the learning dataset.
[0011] C learned by the learning device 10 according to this embodiment * By applying the ring network to various fields, it is expected that the inference accuracy in those fields will be improved. For example, *When the ring network is applied to anomaly detection technology, it is expected that the accuracy of anomaly detection will improve. * An example of an area where the ring network can be applied is C * Ring networks can be applied to a variety of fields that can be reduced to problems such as regression problems, classification problems, data generation, and data distribution estimation.
[0012] <Theoretical structure> C * A theoretical framework for improving the performance of ring networks is described.
[0013] <Problem setting> Each element is C * Consider a K-layer neural network with ring-valued weight matrices and nonlinear activation functions.
[0014] Below, C * Let the ring be A. C * As the ring A, we use the space C(Z) of all continuous functions on the compact Hausdorff space Z.
[0015] For j = 1, ..., K, let d be a natural number that represents the width of the jth layer of the neural network (i.e., the dimension of the jth layer). j Also, each element is C * d takes values in the ring A j ×d j-1 The matrix is used as a weight matrix, and W j Let each element be C * d takes values in the ring A j Let the dimensional vector be the bias, and j In addition, d 0 Note that is the dimension of the input of the neural network.
[0016] A d_j From A d_j The activation function is a nonlinear mapping to σ j However, "d_j" is "d j ". Also, d j-1 Let x be a dimensional vector, and f (j) (x) = σ j (W j x + b j) The vector x when j = 1 is called the input vector. The input vector x is d 0 It is a real-valued vector of dimension, but we will regard it as a constant function that takes an input vector x as input and outputs the input vector x.
[0017] In this case, consider a neural network f(x) shown in the following equation (1).
[0018] The neural network f(x) shown in the above formula (1) is * "Circular Neural Network" and "C * In the following, we will use the neural network f(x) shown in the above formula (1) as a C * This will be called a ring network.
[0019] C * The output of the ring network f(x) is d K Since each element of a dimensional vector is an element of C(Z), f can be considered as a mapping that maps (z, x) to f(x)(z). Note that z∈Z. Therefore, if we fix z in f and consider it as a function of only x, we can call it f z I will write it as follows.
[0020] C * A ring network f is a network of infinite z such that f z This is equivalent to considering a normal neural network whose output is a real-valued vector. If we want a function (neural network) that outputs a real-valued vector, we can define f by a measure P on the compact Hausdorff space Z. z with respect to z (in other words, C by a measure P on the compact Hausdorff space Z) * By integrating the ring network f with respect to z, we obtain a single function. This is done by integrating multiple functions f, each of which outputs a real-valued vector. z This is equivalent to adding the
[0021] However, C * Since the ring A is an infinite-dimensional space, the weight matrix W j and bias b jIt is not possible to handle each element of C on a computer. * Consider a finite-dimensional subspace of the ring A, and define the weight matrix W j and bias b j Each element of v takes a value in this finite-dimensional subspace. 1 , ..., v m m continuous functions are prepared, and the weight matrix W j and bias b j Each element of C * Among the elements of the ring A, there are m continuous functions v 1 , ..., v m Only functions expressed as linear combinations of m are taken, where m is a predetermined natural number.
[0022] <Learning method for improving generalization performance> The loss function is a function L of two variables that always takes a positive value. The loss function L is continuous and convex with respect to the first variable. In this case, C is calculated so as to minimize the loss function L. * Learn the ring network f.
[0023] By a measure P on a compact Hausdorff space Z, f z The integral of with respect to z is defined as μ(x). That is, μ(x) is defined by the following equation (2).
[0024] Also, the derivative of the loss function L with respect to the first variable is ∂ 1 Let L be g(x, [v 1 (z), ..., v m (z)]) = (f z (x)-μ(x) 2 Furthermore, by a measure P on the compact Hausdorff space Z, [v 1 (z), ..., v m (z)] is integrated with respect to z to be v. That is, v is defined by the following equation (3).
[0025] In addition, the Hessian matrix for the second variable of g is H 2 Let g.
[0026] At this time, for the training data (x, y) represented by a pair of an input vector x and its training data y, L(f z Consider evaluating the difference between L(μ(x),y) and L(μ(x),y).
[0027] For any element z in the compact Hausdorff space Z, d 0 For any input vector x, which is a real-valued vector of dimension, |f z If (x)-μ(x)| is sufficiently small, the following equation (4) holds.
[0028] Here, O represents big O in order notation, and the following symbols represent tensor products:
[0029] Also, v(z) = [v 1 (z), ..., v m (z)]. Furthermore, T is the symbol representing transposition.
[0030] Note that due to Jensen's inequality, the left hand side of the above equation (4) is always positive.
[0031] Therefore, (v(z)-ν)H 2 g(x, ν)(v(z)-ν) T If we train to make L(f z The difference between L(μ(x), y) and L(μ(x), y) becomes large, and the measure P z μ, which is the integral of with respect to z, is z This means that a small loss function value can be achieved for
[0032] Here, H 2 When the eigenvalue of g(x,ν) is large, (v(z)-ν)H 2 g(x, ν)(v(z)-ν) T On the other hand, if the value of the loss function is multiplied by a scalar, H 2 g(x,ν) is also multiplied by a scalar. 2 While suppressing the norm of g(x,ν) (the magnitude of the largest eigenvalue), 2It can be seen that regularization should be performed so that the overall eigenvalue of g(x,ν) becomes as large as possible. For this reason, as the regularization term for the loss function L, for example, the regularization term shown in the following formula (5) is used.
[0033] where I is the identity matrix. 1 , λ 2 , η are regularization parameters that take predetermined positive real numbers.
[0034] From the above, for example, the loss function L shown in the following equation (6) reg (x, y) to minimize C * Learning the ring network f (i.e., the weight matrix W of μ(x) j and bias b j (Optimize the weight matrix W j and bias b j may be referred to as, for example, a "learning target parameter," a "parameter to be optimized," or a "parameter."
[0035] This allows us to develop a C with high generalization performance. * A function f is a function of a ring network, or more precisely, a function P on a compact Hausdorff space Z. z We can expect to obtain a trained neural network μ(x) expressed as a function obtained by integrating with respect to z.
[0036] <Example of Hardware Configuration of Learning Device 10> An example of the hardware configuration of the learning device 10 according to this embodiment will be described with reference to Fig. 1. Fig. 1 is a diagram showing an example of the hardware configuration of the learning device 10 according to this embodiment.
[0037] 1, a learning device 10 according to this embodiment includes an input device 101, a display device 102, an external I / F 103, a communication I / F 104, a random access memory (RAM) 105, a read-only memory (ROM) 106, an auxiliary storage device 107, and a processor 108. Each of these pieces of hardware is connected to each other via a bus 109 so as to be able to communicate with each other.
[0038] The input device 101 is, for example, a keyboard, a mouse, a touch panel, physical buttons, etc. The display device 102 is, for example, a display, a display panel, etc. Note that the learning device 10 does not necessarily have to have at least one of the input device 101 and the display device 102, for example.
[0039] The external I / F 103 is an interface with an external device such as a recording medium 103a. Examples of the recording medium 103a include a CD (Compact Disc), a DVD (Digital Versatile Disk), an SD memory card (Secure Digital memory card), and a USB (Universal Serial Bus) memory card.
[0040] The communication I / F 104 is an interface for connecting to a communication network. The RAM 105 is a volatile semiconductor memory (storage device) that temporarily stores programs and data. The ROM 106 is a non-volatile semiconductor memory (storage device) that can store programs and data even when the power is turned off. The auxiliary storage device 107 is a non-volatile storage device such as a hard disk drive (HDD), a solid state drive (SSD), or a flash memory. The processor 108 is a variety of arithmetic devices such as a central processing unit (CPU) or a graphic processing unit (GPU).
[0041] 1 is an example, and the hardware configuration of the learning device 10 is not limited to this. For example, the learning device 10 may have multiple auxiliary storage devices 107 or multiple processors 108, may not have some of the hardware shown in the figure, or may have various hardware other than the hardware shown in the figure.
[0042] <Example of Functional Configuration of Learning Device 10> An example of the functional configuration of the learning device 10 according to this embodiment will be described with reference to Fig. 2. Fig. 2 is a diagram showing an example of the functional configuration of the learning device 10 according to this embodiment.
[0043] As shown in FIG. 2 , the learning device 10 according to this embodiment includes an input unit 201 and a model learning unit 202. These units are realized, for example, by a process in which one or more programs installed in the learning device 10 are executed by the processor 108 or the like. The learning device 10 according to this embodiment also includes a training dataset storage unit 203 and a model storage unit 204. These units are realized, for example, by a storage area of the auxiliary storage device 107 or the like. However, for example, at least one of the training dataset storage unit 203 and the model storage unit 204 may be realized by a storage area of a storage device or the like (e.g., a storage device provided in a database server) communicatively connected to the learning device 10.
[0044] The input unit 201 inputs a learning data set stored in the learning data set storage unit 203. The learning data set is a set of learning data (x, y) represented by a pair of an input vector x and its training data y. Hereinafter, the i-th learning data will be referred to as (x i , y i ), and the training data set is D = {(x i , y i ) |i=1, . . . , |D|}.
[0045] The model learning unit 202 uses the learning data set D input by the input unit 201 to calculate the loss function L shown in the above formula (6). reg C stored in the model storage unit 204 so as to minimize (x, y). * Learning the ring network f (i.e., the weight matrix W of μ(x) j and bias b j (Optimize the C * An existing optimization method (for example, Adam, etc.) may be used to train the ring network f.
[0046] The learning data set storage unit 203 stores the given learning data set D. The model storage unit 204 stores C * Store the ring network f.
[0047] <Model Learning Process> An example of the model learning process according to this embodiment will be described with reference to Fig. 3. Fig. 3 is a flowchart showing an example of the model learning process according to this embodiment.
[0048] The input unit 201 inputs the learning data set D stored in the learning data set storage unit 203 (step S101).
[0049] The model learning unit 202 uses the learning data set D input in step S101 to calculate the C stored in the model storage unit 204. * The model learning unit 202 learns the ring network f (step S102). That is, the model learning unit 202 uses the learning data set D input in step S101 to calculate the loss function L reg (x, y) is minimized by the existing optimization method. j and bias b j Note that μ(x) is a measure P on the compact Hausdorff space Z that optimizes f z This is a neural network expressed as a function obtained by integrating with respect to z.
[0050] This results in a trained neural network μ(x) with high generalization performance.
[0051] <Experiment> An experiment for evaluating the learning device 10 according to this embodiment will be described below.
[0052] In this experiment, we used a dataset called KMNIST (Reference 1) to classify image data of handwritten hiragana characters. * As for the ring network f, the width of each layer is d 0 =28×28=784,d 1 = 1024, d 2 = 2048, d 3 = 1024, d 4 = 10, and the activation functions were ReLU for j = 1, 2, sigmoid for j = 3, and softmax for j = 4.
[0053] Also, Z is a subset of all real numbers. Furthermore, m=2, and v shown in the following equation (7) is i (z) (i=1, 2) was used.
[0054] However, z 1 = 1, z 2 =3.
[0055] Furthermore, as the measure P, N=15 and the measure P shown in the following equation (8) was used.
[0056] However, for k=0,...,4, s=1,2,3, for i=3k+s, ζ i are points independently sampled from a normal distribution with mean k and standard deviation 0.01. In this experiment, the coefficient β i This was also optimized through learning.
[0057] At this time, a training data set consisting of 1000 samples of training data is prepared, and three patterns (λ 1 , λ 2 ) = (0.1, 1), (0, 1), (0, 0), respectively, and the loss function L shown in the above formula (6) is reg (x, y) by Adam. * The learning rate for the ring network f was 10 -3 It was decided.
[0058] Also, C * When training the ring network f, the accuracy rate was evaluated for each epoch using 1000 test data samples. As a result of this evaluation, the accuracy rate (mean and standard deviation) for the 1000 test data samples at each epoch is shown in Figure 4. As shown in Figure 4, λ 1 = λ 2 = 0 (i.e., when no regularization term is used), the accuracy rate is lowest. 1 = 0, λ 2 = 1 (i.e., H 2 When only the norm of g(x,ν) (the magnitude of the largest eigenvalue) is suppressed, λ 1 = λ 2The accuracy rate is higher than when λ = 0. 1 = 0.1, λ 2 = 1 (i.e., H 2 While suppressing the norm of g(x, ν), H 2 It can be seen that the highest accuracy rate is achieved when all eigenvalues of g(x, ν) are maximized.
[0059] This shows that by using the regularization term shown in equation (5) above, a trained neural network μ(x) with high generalization performance can be obtained.
[0060] <Summary> As described above, the learning device 10 according to this embodiment can calculate g(x, [v 1 (z), ..., v m (z)]) = (f z (x)-μ(x) 2 The second variable of (i.e., [v 1 (z), ..., v m (z)]) 2 g and a measure P on the compact Hausdorff space Z such that [v 1 (z), ..., v m (z)] with respect to z and the value ν, 2 While suppressing the maximum eigenvalue of g(x, ν), H 2 The total eigenvalue of g(x, ν) is regularized to be as large as possible, and C * This allows us to learn a ring network f, which has high generalization ability. * A function f is a function of a ring network, or more precisely, a function P on a compact Hausdorff space Z. z We can expect to obtain a trained neural network μ(x) expressed as a function obtained by integrating with respect to z.
[0061] The learning device 10 according to the present embodiment may include, for example, an inference unit that performs inference using the trained neural network μ(x). A device having such an inference unit may be called, for example, an “inference device.”
[0062] The present invention is not limited to the above-described specifically disclosed embodiments, and various modifications, changes, and combinations with known technologies are possible without departing from the scope of the claims.
[0063] [References] Reference 1: Tarin Clanuwat, Mikel Bober-Irizar, Asanobu Kitamoto, Alex Lamb, Kazuaki Yamamoto, David Ha. Deep Learning for Classical Japanese Literature. arXiv:1812.01718.
[0064] 10 Learning device 101 Input device 102 Display device 103 External I / F 103a Recording medium 104 Communication I / F 105 RAM 106 ROM 107 Auxiliary storage device 108 Processor 109 Bus 201 Input unit 202 Model learning unit 203 Learning dataset storage unit 204 Model storage unit
Claims
1. A learning device that learns a neural network having parameters with functions represented by linear combinations of a finite number of continuous functions on a compact Hausdorff space, the learning device comprising: an input unit that inputs a learning data set for learning the neural network; and a model learning unit that uses a Hessian matrix with respect to the finite number of continuous functions of a function represented by the square of the difference between the neural network with fixed values in the compact Hausdorff space and the function obtained by integrating the neural network with a predetermined measure on the compact Hausdorff space, and the learning data set, and learns the parameters so as to minimize a loss function including a regularization term for minimizing the maximum eigenvalue of the Hessian matrix at the value obtained by integrating the finite number of continuous functions with the measure and maximizing the entire set of eigenvalues.
2. The learning device according to claim 1, wherein the regularization term includes a first term for minimizing the maximum eigenvalue of the Hessian matrix at the value obtained by integrating the finite number of continuous functions with the measure, and a second term for maximizing the entire set of eigenvalues of the Hessian matrix at the value obtained by integrating the finite number of continuous functions with the measure.
3. The learning device according to claim 2, wherein the regularization term includes a first regularization parameter multiplied by the first term and a second regularization parameter multiplied by the second term.
4. Let the input data included in the training data constituting the training data set be \(x\), the value obtained by integrating the finite number of continuous functions by the measurement be \(\nu\), and the Hessian matrix at the value \(\nu\) be \(H\). 2 As \(g(x, \nu)\), the first term is \(\|(\eta I + H\). 2 g(x, \nu))\| (where \(\eta\) is a predetermined positive real number and \(I\) is the identity matrix), and the second term is \(\|H\). -1 g(x, \nu)\|\). The learning device according to claim 2 or 3. 2 5. The learning data includes teacher data y for the input data x, and the loss function includes a loss between the value when the input data x is input to the function obtained by integrating the neural network with a predetermined measure on the compact Hausdorff space and the teacher data y, the first term, and the second term. The learning device according to claim 4.
6. A learning method for training a neural network with parameters whose elements are functions represented by linear combinations of a finite number of continuous functions on a compact Hausdorff space, the method comprising: an input procedure for inputting a training dataset for training the neural network; and a model learning procedure for training the parameters so as to minimize the maximum eigenvalue of the Hessian matrix with respect to the finite number of continuous functions of a function represented by the square of the difference between the neural network with fixed values in the compact Hausdorff space and the function obtained by integrating the neural network with a predetermined measure on the compact Hausdorff space, and to minimize a loss function including a regularization term for maximizing the entire set of eigenvalues, using the Hessian matrix and the training dataset. The learning method is executed by a computer.
7. A program for causing a computer that trains a neural network with parameters whose elements are functions represented by linear combinations of a finite number of continuous functions on a compact Hausdorff space to execute: an input procedure for inputting a training dataset for training the neural network; and a model learning procedure for training the parameters so as to minimize the maximum eigenvalue of the Hessian matrix with respect to the finite number of continuous functions of a function represented by the square of the difference between the neural network with fixed values in the compact Hausdorff space and the function obtained by integrating the neural network with a predetermined measure on the compact Hausdorff space, and to minimize a loss function including a regularization term for maximizing the entire set of eigenvalues, using the Hessian matrix and the training dataset.
Citation Information
Patent Citations
Analysis device and program
WO2023175681A1