Information processing device, information processing method, and program

LCNG optimizes ONN parameters by combining zeroth-order and natural gradient methods, addressing training challenges in ONNs with manufacturing variations, resulting in efficient and accurate parameter updates.

WO2025141804A1PCT designated stage expired Publication Date: 2025-07-03NT T INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2023/047052
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-27
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Conventional techniques struggle to effectively train neural networks, particularly optical neural networks (ONNs), due to manufacturing variations and repetitive structures like the Mach-Zehnder interferometer array, which cannot be accurately modeled, leading to difficulties in gradient calculation and unstable convergence.

Method used

The implementation of the Linear Combination of Natural Gradients (LCNG) optimization method, which combines zeroth-order optimization with the natural gradient method, to address the challenges of training ONNs by optimizing parameters through a linear combination of direction vectors, considering both parameter space discrepancy and function space discrepancy.

Benefits of technology

LCNG enables accurate training of ONNs with reduced computational cost and sensitivity to manufacturing variations, achieving faster convergence and efficient parameter updates compared to existing methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2023047052_03072025_PF_FP_ABST
    Figure JP2023047052_03072025_PF_FP_ABST
Patent Text Reader

Abstract

According to the present invention, an information processing device that is for optimizing the parameters of a neural network comprises: a calculation unit that calculates update amounts for the parameters on the basis of the amount of change in output from the neural network when the parameters have been changed by exactly some amount of change; and an updating unit that updates the parameters on the basis of the update amounts.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device, information processing method, and program

[0001] The present invention relates to techniques for training neural networks.

[0002] In recent years, optical neural networks (ONNs) using programmable integrated circuits have been attracting attention as neural networks capable of high-speed processing with low energy consumption.

[0003] The ONN is a hardware circuit, and there are manufacturing variations that cannot be observed. The ONN also has a repeating structure such as a Mach-Zehnder interferometer array (MZI array).

[0004] P. Zhao, P.-Y. Chen, S. Wang, and X. Lin. 2020. Towards query-efficient black-box adversary with zeroth-order natural gradient descent. In Proc. AAAI Conf.Artificial Intelligence, Vol. 34. 6909-6916.

[0005] Due to the above-mentioned characteristics of ONNs, conventional techniques such as backpropagation have not been able to properly train ONNs. This problem is not limited to neural networks of hardware circuits such as ONNs, but can occur in all neural networks.

[0006] The present invention has been made in view of the above points, and has an object to provide a technique for appropriately training a neural network.

[0007] According to the disclosed technology, there is provided an information processing device for optimizing parameters in a neural network, comprising: a calculation unit that calculates an update amount for the parameter based on an amount of change in output from the neural network when the parameter is changed by a certain amount; and an update unit that updates the parameter based on the update amount.

[0008] The disclosed technology provides a technology for properly training a neural network.

[0009] FIG. 1 is a diagram illustrating an example of the structure of an ONN. FIG. 2 is a diagram illustrating notation. FIG. 3 is a diagram illustrating an example of the configuration of an optimization device 100. FIG. 4 is a diagram illustrating an example of the configuration of an information processing device 200. FIG. 5 is a flowchart for explaining the operation of the optimization device 100. FIG. 6 is a diagram illustrating a feedforward ONN. FIG. 7 is a diagram illustrating training loss. FIG. 8 is a diagram illustrating convergence behavior. FIG. 9 is a diagram illustrating elapsed time and memory footprint of training. FIG. 10 is a diagram illustrating an example of the hardware configuration of the device.

[0010] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. The embodiment described below is merely an example, and the embodiment to which the present invention is applied is not limited to the following embodiment.

[0011] In this embodiment, an ONN is taken as the subject of training (which may also be called "learning"), but this is merely one example. For example, the subject of training may be an analog circuit other than an ONN, in which internal manufacturing variations cannot be observed, or a circuit other than an ONN, which has a repetitive structure such as an MZI array. Furthermore, the subject of training is not limited to a hardware circuit such as an ONN, but may also be a general neural network that operates as software on a computer.

[0012] (Problems and Outline of Technology According to the Present Embodiment) First, problems related to the technology according to the present embodiment will be described in detail, and then an outline of the technology according to the present embodiment will be described.

[0013] ONN, a programmable integrated circuit, is attracting attention as a promising system because it is capable of high-speed processing with low energy consumption.

[0014] Figure 1 shows an example of the structure of an ONN consisting of linear and nonlinear units. As shown in Figure 1, the linear units in the ONN are an array of Mach-Zehnder interferometers (MZIs) connected to a waveguide. The MZIs have a structure in which two "PS-DC" pairs, each consisting of a programmable phase shifter (PS) and a directional coupler (DC), are connected sequentially.

[0015] The ONN process is affected by manufacturing variations in the components that make it up, such as the split angle error γ and the attenuation phase error ξ, which includes phase variations in the waveguide and additional attenuation.

[0016] Training an ONN is more difficult than training a general neural network (NN) because errors caused by manufacturing variations cannot be accurately modeled and observed. That is, the gradient with respect to the parameters of an ONN cannot be calculated by backpropagation, which requires complete internal information. A conventional technique for calculating the gradient with respect to the parameters of an ONN is zeroth-order optimization, which directly evaluates an approximate gradient on the ONN.

[0017] There are also challenges arising from the structure of the MZI array. The structure of the MZI array differs from the matrices used in ordinary neural networks in terms of parameterization, and its parameters θ n_p are related to each other by a repeating structure.

[0018] The natural gradient method is known to be robust against local reparameterization of the model. This property is useful for training ONNs, since the MZI array can be regarded as a reparameterization of the matrix used in conventional neural networks. Therefore, a prior art method that directly combines zero-order optimization and the natural gradient method has been proposed (Non-Patent Document 1). However, the technique disclosed in Non-Patent Document 1 requires enormous computational costs.

[0019] In this embodiment, to solve the above-mentioned problems, an optimization device 100 (described later) performs training to optimize the parameters of an ONN using a linear combination natural gradient (LCNG), which is an extended method of zero-order optimization. LCNG is a new optimization method derived from the concept of the natural gradient method and is suitable for zero-order optimization.

[0020] In order to facilitate understanding of LCNG, we will first explain ONN, normal NN training, natural gradient methods, and zero-order optimization, and then explain LCNG in detail.

[0021] Figure 2 shows a summary of the meanings of the letters used in the following explanation. In the text of this specification, for convenience of description, bold letters are not used for vectors or matrices, but it is clear from the context that they indicate vectors or matrices. In addition, normal typeface letters are used for letters representing sets. Furthermore, "_" is used as a subscript for subscripts. "θ n_p " is an example. Also, the symbol T is used to represent transposition.

[0022] (ONN) As shown in Fig. 1, ONN is composed of linear units and nonlinear units. In this embodiment, a unitary matrix is ​​realized by using a Clements mesh as the MZI array of the linear units. In addition, modReLU, which is expressed by the function shown in Fig. 1, is used as the nonlinear unit.

[0023] Each MZI consists of two PS-DC pairs. The PS-DC pair functions as a transfer matrix as shown in Figure 1. In the transfer matrix shown in Figure 1, j is the imaginary unit. PS and modReLU have phase and bias parameters, θ n_p and θ n_b In this embodiment, in order to model the manufacturing variations, ζ is used as the component error in the MZI. 1 , ζ 2 , and γ. 1 , ζ 2 ∈C(|ζ 1 |,|ζ 2 |≦1) is the attenuation-phase error, and γ∈R is the splitting angle error.

[0024] In this embodiment, the phase parameter and the bias parameter are collectively referred to as θ n ∈R. θ=[θ1 , ..., θ N ]∈R N is a parameter vector. An ONN implements a function f with N parameters θ and has an input vector x i For output vector z i = f(x i , θ)∈R M Outside the ONN, for training the ONN, we generate a target vector t i loss function l(z i , t i ) is defined.

[0025] (About NN training and natural gradient method) A training dataset of size D, D = {x i , t i} i=1 D Given the above, the objective of training the NN is to optimize the parameter θ so as to minimize the total loss shown in equation (1) below.

[0026] A standard method for optimizing the parameter θ is a mini-batch processing method using mini-batch stochastic gradient descent (SGD), which iteratively updates the parameter using the following equations (2) and (3).

[0027]

[0028] In the above formula, B = {x i , t i} i=1 B ⊂D is a mini-batch drawn from the training dataset, and ∇ θ is the gradient with respect to θ, usually computed by backpropagation, and η>0 is the learning rate.

[0029] For NNs with potentially correlated parameters, such as ONNs with MZI arrays, gradient descent methods tend to converge slowly and become unstable, whereas natural gradient methods are effective in such situations because they are less sensitive to parameterization.

[0030] Many existing optimization methods, including gradient descent and natural gradient methods, can be derived from the proximal point method (PPM) approach, as shown below.

[0031] Here, the first term in the curly brackets in equation (4) corresponds to the loss function, and the weight λ P The second term with ≥ 0 indicates the parameter space discrepancy (PSD), and the weight λ F The third term with ≧0 indicates function space discrepancy (FSD).

[0032] The PSD functions to reduce the amount of parameter update, and the FSD functions to reduce the amount of change in the estimation result by the function f.

[0033] The gradient descent update formula (2) is a linear approximation to the loss function term in formula (4) and λ P ,=1,λ F = 0, which means that the FSD term is ignored. In the case of the natural gradient method, the probability distribution p θ and consider the FSD term expressed by the following equation (5).

[0034] In formula (5), d KL is the Kullback-Leibler divergence between two probability distributions. The Kullback-Leibler divergence is a measure of how similar two probability distributions are. By applying a quadratic approximation to the FSD term above, we can obtain the update formula for the natural gradient method shown in Equation (6) below.

[0035] In equation (6), I is the identity matrix.

[0036] Also, F in formula (6) θ is as shown in the above equation (7), and indicates the Fisher Information Matrix (FIM) for the parameter θ. Here, the target vector t i is not used, and pθ (t | z i ) and extract a sample t from

[0037] (Zeroth-order optimization) Zeroth-order optimization is a method in which the gradient of the loss function l cannot be calculated by backpropagation, but the function query l(z i , t i ), z i = f(x i , θ) are allowed. Zero-order optimization approximates the gradient using a zero-order gradient estimate as shown in equation (8) below.

[0038] In the above formula (8), δθ q (q=1,...,Q) is usually a direction vector randomly drawn from a multivariate normal distribution N(0,I / N).

[0039] δl q is a quotient difference with smoothing hyperparameter μ>0, as shown in equation (9) above. For a mini-batch of size B and Q direction vectors, B×(Q+1) queries are required. The parameter is the zeroth-order gradient estimate ^∇ defined in equation (8) as follows: θ l B (θ) is used to update in the same manner as in equation (2).

[0040] (Regarding Linear Combination Natural Gradient Method (LCNG)) In this embodiment, the optimization device 100 optimizes the parameter θ of the ONN using LCNG, which is a method that extends zero-order optimization. First, the contents of LCNG will be described, and then the configuration and operation of the optimization device 100 that executes LCNG will be described.

[0041] LCNG follows the principles of natural gradient methods, but differs from conventional natural gradient methods. The commonalities and differences between LCNG and natural gradient methods are listed below.

[0042] a) Both are derived from equation (4) taking into account the FSD term.

[0043] b) Unlike the natural gradient method, LCNG performs gradient estimation using a linear combination of directional vectors, similar to zero-order optimization (Equation (8)). As a result, the gradient estimation is optimized within a subspace spanned by Q directional vectors. As shown in Equation (12) below, the linear combination of directional vectors in LCNG differs from the linear combination of directional vectors in zero-order optimization (Equation (8)).

[0044] c) The natural gradient method uses an inverse matrix (λ) whose size N × N depends on the number of parameters N, as shown in equation (6). P I+λ F F θ ) -1 It is necessary to calculate the following. Note that since it is not practical to calculate the inverse matrix of N×N for a large N, an approximation method is used in actual calculations.

[0045] On the other hand, LCNG calculates the inverse matrix of a Q×Q matrix as shown in equation (13) described later. In practice, the number Q of directional vectors is at most several hundred, so the Q×Q matrix is ​​small and can be calculated.

[0046] LCNG will be described in detail below.

[0047] The direction vector δθ introduced in the explanation of zero-order optimization q and the difference quotient δl q are expressed in matrix and vector form as follows:

[0048] δΘ=[δθ 1 , .... , δθ Q ]∈R N×Q δl = [δl 1 , .... , δl Q ] T ∈R Q where T represents the transposition operator. i = [δz i1 , .... , δz iQ ]∈R M×Q is defined by the following equation (11).

[0049] The value for each q in equation (11) is the direction vector δθ qThis shows how the NN output is changed by changing the parameters by [mathematical formula - see original document]. Then, in LCNG, the parameters are updated by the following equation (12). ηδΘa indicates the amount of update.

[0050] The coefficient a is given by the following equation (13).

[0051] F in formula (13) z_i is as shown in the following equation (14).

[0052] Equation (14) is the output z i The derivation of equation (13) and the calculation procedure of equation (14) will be described later.

[0053] Gradient estimation in ordinary zero-order optimization (Equation (8)) corresponds to a special case of LCNG, where δΘ T δΘ=I,λ P = 1, and λ F = 0. For ordinary zero-order optimization and gradient descent, they are similar in that they both ignore the FSD terms, i.e., the third term in equation (4) and the second term in equation (13).

[0054] In contrast, in LCNG, the FSD term is weighted by λ F Since the second term in the curly brackets in equation (4), the PSD term, which is directly affected by the parameterization, can be considered using , the importance can be relatively reduced. As a result, the optimization using LCNG can be less sensitive to the parameterization of the MZI array in the ONN than the case of using the conventional zero-order optimization.

[0055] (Derivation of Equation (13)) The derivation of Equation (13), which is the coefficient a used in Equation (12), will be described. In LCNG, the coefficient a=[a 1 , .... , a Q ] T The purpose is to optimize the

[0056] Here, u is set as in the following equation (15), and it will be explained how the three terms in the curly brackets in equation (4) are expressed in LCNG.

[0057] The first term is approximated by linearization as shown in the following equation (16).

[0058] The second term is given by the following equation (17).

[0059] The third term is approximated by a second order Taylor series with a as follows:

[0060] Therefore, the objective function to be minimized with respect to a is given by the following equation (19).

[0061] By solving ∂ / ∂aC(a)=0 (partial differential of C(a) with respect to a=0), the following equation (20) is obtained. C(a) is minimized with the following a.

[0062] Here, we introduce the following Jacobian matrix:

[0063] This results in F θ can be expressed as follows:

[0064] Jacobian matrix J z_iθ can generally be calculated by backpropagation. However, in this embodiment, which uses a method based on zero-order optimization, we need to rely on another method. i -J z_iθ δΘ || 2 2 We use the least squares and minimum norm solution of , ie the Moore-Penrose inverse function.

[0065] Substituting equation (23) into equation (22) and equation (22) into equation (20) gives equation (13).

[0066] (Probability distribution and output FIM (regarding equation (14)) Since there are two typical probability distributions that can be defined for the output of a NN, the output zi There are two options for calculating the FIM (14) for . The first probability distribution is the multivariate normal distribution:

[0067] In this case, FIM is z_i =I.

[0068] The second probability distribution is the multinomial distribution, which can be expressed in one of two forms:

[0069]

[0070] Here, z i = [z 1_i , ..., z M_i ] T and t is a one-hot vector.

[0071] The above equation (27) is the softmax function. The derivative of the log-likelihood is obtained from (25) as follows:

[0072] T (uppercase τ) = {[1, 0, ..., 0] T , ..., [0, 0, ..., 1] T} are all possible outcomes, the calculation of equation (14) can be performed using (26) and (28) as the following equation (29):

[0073] (Device Configuration Example) Fig. 3 shows a configuration example of the optimization device 100 according to this embodiment. As shown in Fig. 3, the optimization device 100 includes an input unit 101, a mini-batch generation unit 102, a loss / output calculation unit 103, an output FIM calculation unit 104, a random direction vector generation unit 105, a difference quotient calculation unit 106, an output change amount calculation unit 107, an LCNG coefficient calculation unit 108, a parameter update unit 109, and an output unit 110.

[0074] Although it is assumed here that the target of parameter optimization is an ONN, the optimization device 100 is not limited to ONN and can be applied to any neural network.

[0075] The optimization device 100 does not include an ONN (or other neural network) that is the target of parameter optimization. However, this is not limited to this, and the optimization device 100 may include an ONN (or other neural network) that is the target of parameter optimization. In this case, the loss / output calculation unit 103 includes an ONN (or other neural network). In FIG. 3 , for convenience, it is assumed that the loss / output calculation unit 103 includes an ONN (or other neural network), and that parameter updating is performed on the ONN (or other neural network) included in the loss / output calculation unit 103. The operation of each unit is as follows.

[0076] The input unit 101 acquires training data from an external source. The mini-batch generation unit 102 extracts a part of the training data set as a mini-batch. Specifically, the mini-batch generation unit 102 extracts a part of the training data set as a mini-batch. i , t i} i=1 B is taken from the training dataset D.

[0077] The loss / output calculation unit 103 calculates the output and loss of the ONN for the mini-batch data. The loss is calculated using equation (3). The output is calculated by input z i Output z from ONN for i = f(x i , θ) is obtained.

[0078] The output FIM calculation unit 104 calculates the output z i As mentioned above, there are two calculation methods depending on the probability distribution (multivariate normal distribution or multinomial distribution) defined by the output of the ONN.

[0079] The random direction vector generator 105 generates a random direction vector δθ q (q=1, . . . , Q) is generated. Although the generation method is not limited to a specific method, in this embodiment, as described above, a random direction vector is generated using random numbers from a multivariate normal distribution.

[0080] The difference quotient calculation unit 106 calculates the change in loss when the parameter is changed by a small amount determined by the smoothing hyperparameter μ in the direction of the random direction vector, i.e., the difference quotient. Specifically, the difference quotient is calculated using equation (9).

[0081] The output change amount calculation unit 107 calculates the amount of change in the output when the parameter is changed by a small amount determined by the smoothing hyperparameter μ in the direction of the random direction vector. Specifically, the amount of change in the output is calculated using equation (11).

[0082] The LCNG coefficient calculation unit 108 calculates the coefficient a according to equation (13). The parameter update unit 109 updates the parameters according to equation (12). Note that the LCNG coefficient calculation unit 108 may calculate an update amount (ηδΘa) and pass the update amount to the parameter update unit 109, causing the parameter update unit 109 to update the parameters. The output unit 110 outputs the parameters after the training has converged.

[0083] Note that the device configuration for performing training using LCNG in this embodiment is not limited to the configuration in Fig. 3. For example, training using LCNG may be performed by an information processing device 200 shown in Fig. 4. The optimization device 100 may be considered to be an example of the information processing device 200.

[0084] 4 is an information processing device for optimizing parameters in a neural network. This "neural network" includes both a neural network using a circuit such as an ONN, and a neural network on a general computer.

[0085] The information processing device 200 has a calculation unit 201 and an update unit 202. The calculation unit 201 calculates an update amount for a parameter based on the amount of change in output from a neural network when the parameter is changed by a certain amount. The update unit 202 updates the parameter based on the update amount.

[0086] (Operation of the Optimization Device 100) The operation of the optimization device 100 will be described with reference to the flowchart shown in Fig. 5. As a premise, the optimization device 100 performs optimization on a training data set D={x i , t i} i=1 D Assume that we hold

[0087] In S1 (step 1), the mini-batch generation unit 102 generates a mini-batch B={x i , t i} i=1 B The subsequent training processes from S2 to S8 are performed using this mini-batch.

[0088] In S2, the loss / output calculation unit 103 calculates the loss for the mini-batch data using equation (3). The loss / output calculation unit 103 also calculates the input x i Output z from ONN for i = f(x i , θ).

[0089] In S3, the output FIM calculation unit 104 calculates the output z i Calculate the FIM (Fisher Information Matrix) of

[0090] In S4, the random direction vector generation unit 105 generates a random direction vector δθ q (q=1,...,Q).

[0091] In S5, the difference quotient calculation unit 106 calculates the difference quotient using equation (9), and the output change amount calculation unit 107 calculates the amount of change in the output using equation (11).

[0092] In S6, the LCNG coefficient calculation unit 108 calculates the coefficient a according to equation (13). In S7, the parameter update unit 109 updates the parameters according to equation (12).

[0093] If the training has not converged in S8, the process returns to S1. If the training has converged in S8, the process proceeds to S9. In S9, the output unit 110 outputs the parameter θ.

[0094] (Experimental Settings) In order to demonstrate the effect of the LCNG according to this embodiment, experiments were carried out using four optimization methods.

[0095] The four optimization methods are ZO, ZO-NG, LCNG, and LCNG-I. ZO is an optimization method based on ZO gradient estimation in equation (8) and parameter update in equation (10).

[0096] ZO-NG is an optimization method based on a combination of ZO gradient estimation in equation (8) and updating by the natural gradient method in equation (6) (method of Non-Patent Document 1).

[0097] Both LCNG and LCNG-I are optimization methods of LCNG using equations (12) and (13) according to this embodiment. LCNG and LCNG-I differ in the way they calculate the output FIM (equation (14)). LCNG uses equation (29), while LCNG-I uses F Z_i =I is used.

[0098] In the experiment, the hyperparameter μ = 10 ‐3 was applied to Equations (9) and (11). All optimization methods were implemented using PyTorch®, and some modules were accelerated by customized CUDA® kernels.

[0099] In the experiments, we performed a classification task using the MNIST handwritten digit database using cross-entropy loss. We used a feedforward ONN with three Clements meshes as the MZI array and two modReLU layers as the nonlinear units. Figure 6 shows the configuration of the ONN.

[0100] Let H be the dimensionality of the unitary matrix implemented by each Clements mesh. A discrete Fourier transform was applied to an MNIST image of 784 = 28 × 28 pixels, and the lowest H frequency bins were used as input. In the experiments, four dimensions H = 16, 32, 64, and 128 were investigated. The number N of parameters θ to be trained can be calculated as H = 3 × H (H − 1) + 2 × H. All parameters θ were initialized to 0.

[0101] The number of direction vectors, Q, was set as Q = H / 2, compromising between gradient estimation accuracy and query efficiency. The mini-batch size was set as B = 100. The number of training epochs was 100. Instead of the usual SGD, the Adam optimizer was used to mitigate the effect of variations in the magnitude of the estimated gradient. The best learning rate, η, was searched for ZO in each dimension, H, and the same learning rate was used for the other three methods.

[0102] In the experiment, the ONN was simulated using a conventional computer for each dimension H (i.e., each MZI array size) described above. For the component errors of the ONN, the attenuation phase error ζ and the split error γ were randomly set using the following equation (30).

[0103] In formula (30), r u1 and u2 is a random number sampled from the uniform distribution [0, 1), and r n is a random number sampled from the standard normal distribution N(0, 1). The error settings are as follows:

[0104] In equation (31), the amount of error was controlled by β, with the setting β=1 corresponding to the estimate of the Clements mesh actually calibrated on silicon photonics.

[0105] (Experimental Results) Figure 7 shows the training loss for all combinations of the four methods and the dimension H. The error setting is β = 1. To check the statistical significance, the experiment was performed five times. Each entry value in Figure 7 indicates the average of the five results, and the ± values ​​indicate the standard deviation. For the three methods related to natural gradient methods, ZO-NG, LCNG, and LCNG-I, λ is used as the weight of the FSD term. F =100 was set.

[0106] As shown in Figure 7, all three methods significantly outperform ZO. Among them, LCNG-I performs best except for the smallest dimension, H = 16. Figure 8 shows the convergence behavior for H = 64, sampled from five runs.

[0107] Fig. 9 shows the elapsed training time and memory footprint on the computer used in this experiment. As shown in Fig. 9, for ZO-NG, i.e., the existing combination of ZO optimization and natural gradient methods, these values ​​increase rapidly as the dimension H increases.

[0108] When H = 128, ZO-NG ran out of memory (48 GB). On the other hand, LCNG and LCNG-I can perform calculations with much less overhead than ZO.

[0109] (Hardware Configuration Example) Any of the devices described in this embodiment (the optimization device 100 and the information processing device 200) can be realized, for example, by causing a computer to execute a program. This computer may be a physical computer or a virtual machine on the cloud.

[0110] That is, the device can be realized by executing a program corresponding to the processing performed by the device using hardware resources such as a CPU and memory built into a computer. The program can be recorded on a computer-readable recording medium (such as a portable memory) and stored or distributed. The program can also be provided via a network such as the Internet or email.

[0111] Fig. 10 is a diagram showing an example of the hardware configuration of the computer. The computer in Fig. 10 includes a drive device 1000, an auxiliary storage device 1002, a memory device 1003, a CPU 1004, an interface device 1005, a display device 1006, an input device 1007, an output device 1008, and the like, all of which are interconnected via a bus B. The computer may further include a GPU.

[0112] The program that realizes the processing on the computer is provided by a recording medium 1001, such as a CD-ROM or a memory card. When the recording medium 1001 storing the program is set in the drive device 1000, the program is installed from the recording medium 1001 to the auxiliary storage device 1002 via the drive device 1000. However, the program does not necessarily have to be installed from the recording medium 1001, but may be downloaded from another computer via a network. The auxiliary storage device 1002 stores the installed program as well as necessary files, data, etc.

[0113] The memory device 1003 reads and stores a program from the auxiliary storage device 1002 when an instruction to start the program is received. The CPU 1004 realizes functions related to the device in accordance with the program stored in the memory device 1003. The interface device 1005 is used as an interface for connecting to a network, etc. The display device 1006 displays a GUI (Graphical User Interface) or the like according to the program. The input device 1007 is composed of a keyboard, mouse, buttons, a touch panel, etc., and is used to input various operation instructions. The output device 1008 outputs the results of calculations.

[0114] (Effects of the embodiment, etc.) As described above, the technology described in this embodiment provides a technology for appropriately training neural networks consisting of circuits such as ONN, and general neural networks on a computer.

[0115] In particular, it is possible to perform highly accurate training at a lower computational cost than conventional techniques (e.g., Non-Patent Document 1) for neural networks such as ONNs that have unobservable internal manufacturing variations or repetitive structures such as MZI arrays.

[0116] The following additional notes are provided regarding the above-described embodiments.

[0117] <Additional Notes> (Additional Item 1) An information processing device for optimizing parameters in a neural network, comprising: a memory; and at least one processor connected to the memory, wherein the processor calculates an update amount for the parameter based on an amount of change in output from the neural network when the parameter is changed by a certain amount, and updates the parameter based on the update amount. (Additional Item 2) The information processing device according to Additional Item 1, wherein the processor calculates an amount of change for the parameter used to determine the amount of change in the output, using a directional vector indicating a direction in which the parameter is changed and a value indicating a magnitude of the change in the parameter. (Additional Item 3) The information processing device according to Additional Item 2, wherein the processor calculates the update amount using a linear combination of a coefficient calculated using a Fisher information matrix of the output and the amount of change in the output, and the directional vector. (Additional Item 4) The information processing device according to Additional Item 3, wherein the coefficient includes a function space deviance. (Supplementary Item 5) The information processing device according to Supplementary Item 3, wherein the processor calculates the Fisher information matrix based on a probability distribution assumed in the output. (Supplementary Item 6) An information processing method executed by an information processing device for optimizing parameters in a neural network, the information processing method comprising: a calculation step of calculating an update amount for the parameter based on an amount of change in output from the neural network when the parameter is changed by a certain amount; and an update step of updating the parameter based on the update amount. (Supplementary Item 7) A non-transitory storage medium storing a program for causing a computer to function as each unit in the information processing device according to any one of Supplementary Item 1 to 5.

[0118] Although the present embodiment has been described above, the present invention is not limited to such a specific embodiment, and various modifications and changes are possible within the scope of the gist of the present invention described in the claims.

[0119] REFERENCE SIGNS LIST 100 Optimization device 101 Input unit 102 Mini-batch generation unit 103 Loss / output calculation unit 104 Output FIM calculation unit 105 Random direction vector generation unit 106 Divided difference calculation unit 107 Output change amount calculation unit 108 LCNG coefficient calculation unit 109 Parameter update unit 110 Output unit 200 Information processing device 201 Calculation unit 202 Update unit 1000 Drive device 1001 Recording medium 1002 Auxiliary storage device 1003 Memory device 1004 CPU 1005 Interface device 1006 Display device 1007 Input device 1008 Output device

Claims

1. An information processing apparatus for optimizing parameters in a neural network, comprising: a calculation unit that calculates an update amount of the parameters based on a change amount of an output from the neural network when the parameters are changed by a certain change amount; and an update unit that updates the parameters based on the update amount.

2. The information processing apparatus according to claim 1, wherein the calculation unit calculates the change amount for changing the parameters, which is used to obtain the change amount of the output, using a direction vector indicating a direction in which the parameters are changed and a value indicating a magnitude of the change in the parameters.

3. The information processing apparatus according to claim 2, wherein the calculation unit calculates the update amount using a linear combination of a coefficient calculated using a Fisher information matrix of the output and the change amount of the output, and the direction vector.

4. The information processing apparatus according to claim 3, wherein the coefficient includes a function space divergence.

5. The information processing apparatus according to claim 3, wherein the calculation unit calculates the Fisher information matrix based on a probability distribution assumed in the output.

6. An information processing method executed by an information processing apparatus for optimizing parameters in a neural network, comprising: a calculation step of calculating an update amount of the parameters based on a change amount of an output from the neural network when the parameters are changed by a certain change amount; and an update step of updating the parameters based on the update amount.

7. A program for causing a computer to function as each unit in the information processing apparatus according to any one of claims 1 to 5.