Training Method, Device, Medium, and Program Product of Optical Neural Network
By iteratively performing input training optical signals, parameter offset processing, gradient calculation and optimization operations in the optical neural network, the problem of low speed and accuracy of parameter gradient calculation of optical neural networks in the prior art is solved, and a more efficient training process is achieved.
Patent Information
- Application Number
- CN202510160011.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-02-13
AI Technical Summary
The prior art has low speed and accuracy in parameter gradient calculation of optical neural networks, which affects training efficiency.
The training efficiency of the optical neural network is improved by iteratively performing a series of operations, including inputting training optical signals, performing parameter offset processing, calculating the current parameter gradient, and optimizing phase parameters.
It realizes efficient and accurate calculation of the parameter gradient of optical neural networks, and improves training efficiency.
Smart Images

Figure CN119623577B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a training method, device, medium, and program product for an optical neural network. Background Art
[0002] Optical neural networks (ONNs) are a new type of computing architecture that uses photons as information carriers to implement neural network functions. It uses the wave nature and interference effects of photons to perform computing tasks, and has advantages such as high speed, low energy consumption, parallel processing, and three-dimensional integration. It is applied in fields such as image recognition, speech recognition, big data processing, complex system simulation, optical communication, and biomedicine.
[0003] In the related art, the finite difference method and the adjoint variable method are generally used to determine the parameter gradient during the training process of the optical neural network. However, the related art has low calculation speed and accuracy for the parameter gradient, which affects the training efficiency of the optical neural network. Summary of the Invention
[0004] In view of the problems, the present invention provides a training method, device, device, medium, and program product for an optical neural network.
[0005] According to a first aspect of the present invention, there is provided a training method for an optical neural network, the optical neural network including at least one optical interference unit, and the optical interference unit having at least one phase parameter; the training method includes: iteratively performing the following operations until an iteration stop condition is reached:
[0006] Inputting a training optical signal corresponding to a training sample into the optical neural network, and outputting first output data for the training sample;
[0007] Performing a parameter offset process on the current phase parameter of at least one optical interference unit using a parameter offset amount to obtain an optical neural network after parameter offset;
[0008] Inputting a training optical signal corresponding to a training sample into the optical neural network after parameter offset, and outputting second output data for the training sample;
[0009] Determining a current parameter gradient for the current phase parameter according to the first output data, the second output data, and a preset loss function;
[0010] Optimizing the current phase parameter of the optical neural network using the current parameter gradient to obtain an optimized optical neural network.
[0011] According to an embodiment of the present invention, the optical neural network after parameter offset includes an optical neural network after positive offset and an optical neural network after negative offset;
[0012] Performing parameter offset processing on the current phase parameter of at least one optical interference unit using a parameter offset amount, the obtained optical neural network after parameter offset includes:
[0013] Adding the current phase parameter and the parameter offset amount to obtain the optical neural network after positive offset;
[0014] Subtracting the parameter offset amount from the current phase parameter to obtain the optical neural network after negative offset.
[0015] According to an embodiment of the present invention, the second output data includes positive offset output data and negative offset output data;
[0016] Inputting the training optical signal corresponding to the training sample into the optical neural network after parameter offset, the output of the second output data for the training sample includes:
[0017] Inputting the training optical signal corresponding to the training sample into the optical neural network after positive offset, and outputting positive offset output data;
[0018] Inputting the training optical signal corresponding to the training sample into the optical neural network after negative offset, and outputting negative offset output data.
[0019] According to an embodiment of the present invention, the output of the optical neural network is represented by an objective function, and the objective function is a function of the phase parameter; determining the current parameter gradient for the current phase parameter according to the first output data, the second output data, and a preset loss function includes:
[0020] Processing the second output data using a preset parameter offset formula to obtain the first gradient of the objective function with respect to the current phase parameter;
[0021] Determining the second gradient of the loss function with respect to the first output data;
[0022] Determining the current parameter gradient for the current phase parameter according to the first gradient and the second gradient.
[0023] According to an embodiment of the present invention, the second output data includes positive offset output data and negative offset output data;
[0024] Processing the second output data using a preset parameter offset formula to obtain the first gradient of the objective function with respect to the current phase parameter includes:
[0025] Multiplying the difference between the positive offset output data and the negative offset output data by a preset value to obtain the first gradient.
[0026] According to an embodiment of the present invention, the loss function is a function of the objective function;
[0027] Determining the second gradient of the loss function with respect to the first output data includes:
[0028] Determining the derivative of the loss function with respect to the objective function;
[0029] Processing the first output data using the derivative to obtain the second gradient.
[0030] According to an embodiment of the present invention, determining the current parameter gradient for the current phase parameter based on the first gradient and the second gradient includes:
[0031] Multiplying the first gradient and the second gradient to obtain the current parameter gradient.
[0032] According to an embodiment of the present invention, the optical interference unit includes at least one phase shifter, and the phase shifter is represented by a transfer matrix; the parameter offset is determined by the following steps:
[0033] Processing the transfer matrix using the offset calculation formula to determine the parameter offset.
[0034] According to an embodiment of the present invention, the offset calculation formula includes matrix transformation parameters, and the matrix transformation parameters are represented by a matrix transformation formula;
[0035] Processing the transfer matrix using the offset calculation formula includes:
[0036] Converting the transfer matrix into an exponential form of a preset format, where the exponential form includes a Hermitian matrix;
[0037] Processing the Hermitian matrix using the matrix transformation formula to determine the parameter values of the matrix transformation parameters;
[0038] Determining the parameter offset according to the parameter values and the offset calculation formula.
[0039] According to an embodiment of the present invention, processing the Hermitian matrix using the matrix transformation formula to determine the parameter values of the matrix transformation parameters includes:
[0040] Determining two eigenvalues of the Hermitian matrix;
[0041] Processing the two eigenvalues using the matrix transformation formula to obtain the parameter values of the matrix transformation parameters.
[0042] According to an embodiment of the present invention, optimizing the current phase parameter of the optical neural network using the current parameter gradient to obtain the optimized optical neural network includes:
[0043] Processing the current parameter gradient and the current phase parameter using a preset gradient descent formula to obtain the optimized phase parameter;
[0044] Adjusting the current phase parameter of at least one optical interference unit according to the optimized phase parameter to obtain the optimized optical neural network.
[0045] According to an embodiment of the present invention, adjusting the current phase parameter of at least one optical interference unit according to the optimized phase parameter includes:
[0046] For each optical interference unit, according to the optimized phase parameter for the optical interference unit, by adjusting the voltage and / or temperature of the optical interference unit, the current phase parameter of the optical interference unit is adjusted.
[0047] A second aspect of the present invention provides a training device for an optical neural network. The optical neural network has at least one optical interference unit, and the optical interference unit includes at least one phase parameter. The device includes:
[0048] A first input / output module, configured to input a training optical signal corresponding to a training sample into the optical neural network and output first output data for the training sample;
[0049] A parameter offset module, configured to perform a parameter offset process on the current phase parameter of at least one optical interference unit by using a parameter offset amount to obtain an optical neural network with offset parameters;
[0050] A second input / output module, configured to second input a training optical signal corresponding to a training sample into the optical neural network with offset parameters and output second output data for the training sample;
[0051] A determination module, configured to determine a current parameter gradient for the current phase parameter according to the first output data, the second output data, and a preset loss function;
[0052] An optimization module, configured to optimize the current phase parameter of the optical neural network by using the current parameter gradient to obtain an optimized optical neural network.
[0053] A third aspect of the present invention provides an electronic device, including: one or more processors; a memory, configured to store one or more computer programs, wherein the above one or more processors execute the above one or more computer programs to implement the steps of the above method.
[0054] A fourth aspect of the present invention further provides a computer-readable storage medium, on which a computer program or instruction is stored, and when the computer program or instruction is executed by a processor, the steps of the above method are implemented.
[0055] A fifth aspect of the present invention further provides a computer program product, including a computer program or instruction, and when the computer program or instruction is executed by a processor, the steps of the above method are implemented. Description of the Drawings
[0056] Through the following description of the embodiments of the present invention with reference to the accompanying drawings, the above content and other objects, features, and advantages of the present invention will become clearer. In the drawings:
[0057] Figure 1 is an application scenario diagram of the training method of the optical neural network according to an embodiment of the present invention.
[0058] Figure 2 is a flowchart of the training method of the optical neural network according to an embodiment of the present invention.
[0059] Figure 3 is a schematic diagram of the training method of the optical neural network according to an embodiment of the present invention.
[0060] Figure 4 is a schematic diagram of the method for determining the parameter offset according to an embodiment of the present invention.
[0061] Figure 5 is a schematic diagram of determining the current parameter gradient according to an embodiment of the present invention.
[0062] Figure 6 is a structural block diagram of the training device of the optical neural network according to an embodiment of the present invention.
[0063] Figure 7 is a block diagram of an electronic device suitable for implementing the above method according to an embodiment of the present invention. Detailed Embodiments
[0064] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. In the following detailed description, for the sake of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present invention. However, obviously, one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present invention.
[0065] The terms used herein are merely for describing specific embodiments and are not intended to limit the present invention. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0066] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0067] In the case of using expressions such as "at least one of A, B, and C", generally, it should be interpreted according to the meaning that those skilled in the art usually understand this expression (for example, "a system having at least one of A, B, and C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).
[0068] With the development of big data resources, advanced algorithms, and high-performance computing hardware, artificial intelligence has played a significant role in various industries. However, neural network computing has gradually reached a physical bottleneck, and the demand for new computing methods is increasing day by day. As a new computing mode, optical computing has developed very rapidly and can provide more efficient computing power and development potential in dealing with some computing processes. For example, silicon-based optoelectronics provides an attractive platform for optical computing due to its high-speed signal processing ability and compatibility with existing microelectronic manufacturing processes. Integrating optical components on a silicon chip allows the creation of complex optoelectronic circuits that can perform neural network operations at the speed of light.
[0069] An important process in optical neural network applications is the training of the model. Commonly used training methods are gradient descent methods such as the relatively basic finite difference method and the adjoint variable method. The finite difference method is to calculate, for each target parameter for which the gradient needs to be calculated , by adding a small perturbation to the parameter , the change in the loss function L to calculate . For the adjoint variable method, by converting the gradient calculation formula into an expression of physical quantities (such as the dielectric constant in a waveguide, etc.) in the optical computing process, an analytical solution for calculating the gradient value is obtained. However, the related technology has low speed and accuracy in calculating the parameter gradient, which affects the training efficiency of the optical neural network.
[0070] Embodiments of the present invention provide a training method for an optical neural network. The optical neural network has at least one optical interference unit, and the optical interference unit includes at least one phase parameter. The method includes: iteratively performing the following operations until an iteration stop condition is reached: inputting a training optical signal corresponding to a training sample into the optical neural network to output first output data for the training sample; performing a parameter offset process on the current phase parameters of at least one optical interference unit using a parameter offset amount to obtain an optical neural network with offset parameters; inputting the training optical signal corresponding to the training sample into the optical neural network with offset parameters to output second output data for the training sample; determining a current parameter gradient for the current phase parameters according to the first output data, the second output data, and a preset loss function; and optimizing the current phase parameters of the optical neural network using the current parameter gradient to obtain an optimized optical neural network.
[0071] Figure 1 It is an application scenario diagram of the training method of the optical neural network according to an embodiment of the present invention.
[0072] As Figure 1 shown, the application scenario 100 according to this embodiment may include an optical neural network 101 composed of at least one optical interference unit 110, a first terminal device 102, a second terminal device 103, a server 104, and a network 105. The network 105 is used as a medium to provide communication links between the optical neural network 101 and the first terminal device 102, the second terminal device 103, and the server 104. The network 105 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc. It should be noted that Figure 1 in the network connection relationship where at least one optical interference unit 110 is rectangular, the principle of other connection relationships such as triangular network connection is the same.
[0073] The user can input a training optical signal into the optical neural network 101 to obtain the output data of the optical neural network 101; then send the output data of the optical neural network 101 to the first terminal device 102, the second terminal device 103, or the server 104 through the network 105. The first terminal device 102, the second terminal device 103, or the server 104 can determine the current parameter gradient based on the output data and a preset loss function, so as to adjust the current phase parameters of the optical neural network 101 according to the current parameter gradient.
[0074] Various communication client applications may be installed on the first terminal device 102 and the second terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only for example).
[0075] The first terminal device 102 and the second terminal device 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablets, laptop computers, and desktop computers, etc.
[0076] The server 104 may be a server providing various services. For example, the server can perform data analysis and other processing on the output data of the optical neural network and feedback the processing result to the optical neural network 101, so as to adjust the phase parameters of the optical neural network and realize the training of the optical neural network.
[0077] It should be noted that Figure 1 in the network connection relationship where at least one optical interference unit 110 is rectangular, the principle of other connection relationships such as triangular network connection is the same.
[0078] It should be understood that Figure 1The numbers of the optical neural networks, terminal devices, networks, and servers in it are merely illustrative. According to the implementation requirements, there can be any number of optical neural networks, terminal devices, networks, and servers.
[0079] Based on the Figure 1 described scenario, the training method of the optical neural network of the disclosed embodiments will be described in detail through Figures 2 to 5 the following.
[0080] Figure 2 is a flowchart of the training method of the optical neural network according to an embodiment of the present invention.
[0081] According to an embodiment of the present invention, the optical neural network has at least one optical interference unit, and the optical interference unit includes at least one phase parameter. For example, the optical interference unit can be a Mach-Zehnder interferometer (MZI), and an MZI generally consists of two beam splitters and two phase shifters. Among them, the adjustable phase parameter is in the phase shifter. The transfer matrix of the phase shifter can be expressed as where is the phase parameter, which is also the parameter to be trained in the optical neural network.
[0082] As Figure 2 shown, the training method 200 of the optical neural network in this embodiment can iteratively execute operation S210 to operation S250 until the iteration stop condition is reached, and the training method of this optical neural network can be executed by the server.
[0083] The iteration stop condition can preset the number of training rounds, and the training is completed when the number of iterations reaches the set number of training rounds. For example, the number of training rounds is 5000, and the training is completed when the network iteration number reaches 5000 times.
[0084] The iteration stop condition can also be that the value of the loss function is less than a certain value. For example, the training is completed when the value of the loss function is less than the preset value.
[0085] It should be noted that the present invention does not limit the iteration stop condition, and it can be flexibly configured according to actual application needs.
[0086] In operation S210, the training optical signal corresponding to the training sample is input into the optical neural network, and the first output data for the training sample is output.
[0087] Before performing operation S210, it can also include obtaining a training set including multiple training samples and the sample labels of each training sample.
[0088] In operation S220, parameter offset processing is performed on the current phase parameter of at least one optical interference unit by using a parameter offset amount to obtain an optical neural network with parameters offset.
[0089] In operation S230, the training optical signal corresponding to the training sample is input into the optical neural network with the parameter offset of the optical signal, and the second output data for the training sample is output.
[0090] In operation S240, according to the first output data, the second output data, and the preset loss function, the current parameter gradient for the current phase parameter is determined.
[0091] In operation S250, the current phase parameter of the optical neural network is optimized by using the current parameter gradient to obtain the optimized optical neural network.
[0092] The training sample can be an image sample, a voice sample, etc. The training optical signal corresponding to the training sample can be formed by encoding the training sample into the optical signal. One training sample corresponds to a set of training optical signals, and a set of training optical signals can include multiple training optical signals.
[0093] One method for obtaining the above training set can be obtained by the user himself collecting various historical image samples or historical voice samples at different times and performing corresponding manual annotations. Another method can be to obtain the actually required image samples with sample labels by selecting data samples from a public image sample database provided by a third party, or to obtain the actually required voice samples with sample labels by selecting data samples from a public voice sample database provided by a third party.
[0094] After obtaining the training set by the above method, the training sample is encoded into a training optical signal for training the optical neural network.
[0095] It can be understood that the above optical neural network is the optical neural network that has not been optimized in this embodiment.
[0096] The output of the optical neural network can be represented by an objective function related to the phase parameter In one embodiment, the objective function can be represented by the following formula (1).
[0097] (1);
[0098] Where, is the current phase parameter, , is the vector composed of the input training optical signals, is the transfer matrix of the optical neural network, is transpose.
[0099] In some other embodiments, the objective function can be expressed as , where , is a vector composed of the input training optical signals, is the transfer matrix of the optical neural network, is transpose of, O is the output selection matrix, and the elements on the diagonal of the output selection matrix are 1 or 0. An element of 1 indicates that the light in the waveguide is selected as the output, and 0 indicates that it is not selected.
[0100] For example, if the light in the first waveguide among two waveguides is selected as the output, then the output selection , for example, if the light in the second waveguide among two waveguides is selected as the output, then the matrix , for example, if the light in all waveguides among two waveguides is selected as the output, then the matrix . It should be noted that the case of multiple waveguides can be deduced by analogy.
[0101] The parameter offset can be determined according to a preset method and pre-configured. For example, the parameter offset can be .
[0102] The parameter offset processing can include adding the current phase parameter to the parameter offset and subtracting the parameter offset from the current phase parameter.
[0103] For example, using the parameter offset to perform parameter offset processing on the current phase parameter of at least one optical interference unit, the optical neural network after parameter offset obtained can include: adding the current phase parameter to the parameter offset can obtain the optical neural network after positive offset; subtracting the parameter offset from the current phase parameter to obtain the optical neural network after negative offset.
[0104] The optical neural network after positive offset can be expressed by the following formula (2).
[0105] (2);
[0106] where is the parameter offset.
[0107] The optical neural network after negative offset can be expressed by the following formula (3).
[0108] (3);
[0109] where is the parameter offset.
[0110] Specifically, for example, if an optical neural network includes an optical interference unit A and an optical interference unit B, then the current phase parameters of the optical interference unit A and the current phase parameters of the optical interference unit B are respectively added to the parameter offset to obtain a positively offset optical neural network; the current phase parameters of the optical interference unit A and the current phase parameters of the optical interference unit B are respectively subtracted from the parameter offset to obtain a negatively offset optical neural network.
[0111] The second output data may include positively offset output data output by the positively offset optical neural network and negatively offset output data output by the negatively offset optical neural network.
[0112] For example, inputting a training optical signal corresponding to a training sample into the parameter-offset optical neural network may include inputting the training optical signal corresponding to the training sample into the positively offset optical neural network to output positively offset output data; inputting the training optical signal corresponding to the training sample into the negatively offset optical neural network to output negatively offset output data.
[0113] By performing parameter offset processing on the optical neural network using the parameter offset, a positively offset optical neural network and a negatively offset optical neural network are obtained, so as to obtain positively offset output data and negatively offset output data, thereby enabling the determination of the current parameter formula using the parameter offset formula.
[0114] The preset loss function may be a mean squared error loss function, a logarithmic loss function, a cross-entropy loss function, etc.
[0115] Optimizing the current phase parameters of the optical neural network using the current parameter gradient may include processing the current parameter gradient and the current phase parameters using a preset gradient descent formula to obtain optimized phase parameters; adjusting the current phase parameters of at least one optical interference unit according to the optimized phase parameters.
[0116] The gradient descent formula may adopt the following formula (4).
[0117] (4);
[0118] where is the current phase parameter, is the current parameter gradient, is the optimized phase parameter, is the learning rate, and the learning rate can be pre-configured.
[0119] Optimizing the current phase parameters of the optical neural network using the current parameter gradient can also use an optimizer method to optimize the current phase parameters. For example, the Adam optimization algorithm can be used to optimize the current phase parameters.
[0120] According to an embodiment of the present invention, adjusting the current phase parameters of at least one optical interference unit according to the optimized phase parameters includes: for each optical interference unit, adjusting the current phase parameters of the optical interference unit by adjusting the voltage and / or temperature of the optical interference unit according to the optimized phase parameters for the optical interference unit.
[0121] According to an embodiment of the present invention, by inputting a training optical signal corresponding to a training sample into the optical neural network, first output data processed by the optical neural network is output, then parameter offset processing is performed on the current phase parameters of the optical neural network using a parameter offset amount, and the training optical signal corresponding to the training sample is input into the optical neural network after parameter offset, and second output data is output; then, using the first output data, the second output data, and a preset loss function to calculate the current parameter gradient, thereby realizing an efficient and accurate calculation of the parameter gradient, which in turn helps to improve the training efficiency of the optical neural network.
[0122] Figure 3 is a schematic diagram of a training method of an optical neural network according to an embodiment of the present invention.
[0123] As Figure 3 shown, the optical neural network 320 of this embodiment includes current phase parameters and current phase parameters ; subtracting the current phase parameters and current phase parameters from the parameter offset amount to obtain a negatively offset optical neural network 330, where the negatively offset optical neural network 330 includes negatively offset phase parameters and negatively offset phase parameters ; adding the current phase parameters and current phase parameters to the parameter offset amount to obtain a positively offset optical neural network 340, where the positively offset optical neural network 340 includes positively offset phase parameters and positively offset phase parameters ; Then, input the training optical signal 310 into the optical neural network 320, the negatively offset neural network 330, and the positively offset neural network 340 respectively, and output the first output data 321, the negatively offset output data 331, and the positively offset output data 341; then determine the current parameter gradient 350 according to the first output data 321, the negatively offset output data 331, the positively offset output data 341, and a preset loss function; afterwards, use the current parameter gradient 350 to optimize the current phase parameters of the optical neural network 320, realizing one iteration of the optical neural network, and then loop to execute the above operations until the iteration stop condition is reached.
[0124] According to an embodiment of the present invention, the optical interference unit includes at least one phase shifter, and the phase shifter is represented by a transfer matrix; the parameter offset is determined by the following steps: processing the transfer matrix with an offset calculation formula to determine the parameter offset.
[0125] According to an embodiment of the present invention, the offset calculation formula includes matrix transformation parameters, and the matrix transformation parameters are represented by a matrix transformation formula.
[0126] The offset calculation formula can be represented by the following formula (5).
[0127] (5);
[0128] Wherein, represents the parameter offset, and r represents the matrix transformation parameter.
[0129] Processing the transfer matrix with the offset calculation formula to determine the parameter offset may include:
[0130] First, convert the transfer matrix into an exponential form of a preset format, wherein the exponential form includes a Hermitian matrix.
[0131] For example, convert the transfer matrix into the exponential form , such that , and then solve , obtaining a = -1 as a real number, being a Hermitian matrix.
[0132] Then, process the Hermitian matrix with the matrix transformation formula to determine the parameter value of the matrix transformation parameter.
[0133] The matrix transformation formula can be the following formula (6).
[0134] (6);
[0135] Where e 1 , e 0are two eigenvalues of the Hermitian matrix G, and a is a real constant of -1.
[0136] According to an embodiment of the present invention, processing the Hermitian matrix using a matrix transformation formula and determining the parameter value of the matrix transformation parameter includes: determining two eigenvalues of the Hermitian matrix; processing the two eigenvalues using the matrix transformation formula to obtain the parameter value of the matrix transformation parameter.
[0137] According to , it can be obtained that e 1 = 1, e 0 = 0, and thus the matrix transformation parameter is determined according to formula (6).
[0138] After that, the parameter offset is determined according to the parameter value and the offset calculation formula.
[0139] For example, substituting the parameter value of the matrix transformation parameter into the above formula (5) to obtain the parameter offset .
[0140] By transforming the transfer matrix into an exponential form, the Hermitian matrix is determined, and then the parameter offset is determined according to the Hermitian matrix and the offset calculation formula, so that the determination of the parameter offset is closely related to the phase shifter, and thus the parameter gradient can be calculated efficiently and accurately.
[0141] Figure 4 is the schematic diagram of the method for determining the parameter offset according to an embodiment of the present invention.
[0142] As Figure 4 shown, the transfer matrix 410 is converted into an exponential form 420 to obtain the equation , and then the Hermitian matrix G430 and the parameter a440 are solved, where a = -1 is a real number, ; then two eigenvalues of the Hermitian matrix G are determined to obtain the eigenvalue e 1 450 and the eigenvalue e 0 460; after that, the eigenvalue e 1 450, the eigenvalue e 0 460 and the parameter a440 are substituted into the matrix transformation formula to obtain the parameter value 470 of the matrix transformation parameter r; then the parameter value 470 of the matrix transformation parameter r is substituted into the offset calculation formula to determine the parameter offset 480.
[0143] According to an embodiment of the present invention, the output of the optical neural network is represented by an objective function, which is a function of phase parameters. Determining the current parameter gradient for the current phase parameter according to the first output data, the second output data, and a preset loss function includes:
[0144] Processing the second output data using a preset parameter offset formula to obtain a first gradient of the objective function with respect to the current phase parameter; determining a second gradient of the loss function with respect to the first output data; and determining the current parameter gradient for the current phase parameter according to the first gradient and the second gradient.
[0145] For the current phase parameter The current parameter gradient Can be represented by the following formula (7).
[0146] (7);
[0147] Wherein, Represents the second gradient, Represents the first gradient.
[0148] According to an embodiment of the present invention, the second output data may include positive offset output data and negative offset output data. Processing the second output data using a preset parameter offset formula to obtain a first gradient of the objective function with respect to the current phase parameter may include: multiplying the difference between the positive offset output data and the negative offset output data by a preset value to obtain the first gradient.
[0149] The parameter offset formula may be the following formula (8).
[0150] (8);
[0151] Wherein, Represents the first gradient, And Are the second output data, wherein, May represent the positive offset output data, May represent the negative offset output data; r is a matrix transformation parameter, , Is the parameter offset amount, and .
[0152] Inputting the positive offset output data and the negative offset output data into the above formula (8) can obtain the first gradient.
[0153] According to an embodiment of the present invention, the loss function is a function of the objective function. Determining the second gradient of the loss function with respect to the first output data includes: determining the derivative of the loss function with respect to the objective function; processing the first output data using the derivative to obtain the second gradient.
[0154] In one embodiment, the loss function can obtain the MSE loss function according to the following formula (9).
[0155] (9);
[0156] where represents the objective function and label represents the sample label.
[0157] At this time, the derivative of the loss function with respect to the objective function can be expressed by the following formula (10).
[0158] (10).
[0159] Processing the first output data using the derivative to obtain the second gradient can include inputting the first output data into formula (10) to obtain the second gradient.
[0160] According to an embodiment of the present invention, determining the current parameter gradient for the current phase parameter based on the first gradient and the second gradient includes: multiplying the first gradient and the second gradient to obtain the current parameter gradient.
[0161] Figure 5 is the schematic diagram of determining the current parameter gradient according to an embodiment of the present invention.
[0162] As Figure 5 shown, the method for determining the current parameter gradient in this embodiment includes determining the first gradient and determining the second gradient. Among them, determining the first gradient includes: obtaining the positive-offset output data 510 of the optical neural network output after positive offset and the negative-offset output data 520 of the optical neural network output after negative offset. Among them, the positive-offset output data 510 can be expressed as , and the negative-offset output data 520 can be expressed as ; then, inputting the positive-offset output data 510 , the negative-offset output data 520 into the parameter offset formula and using the formula to obtain which is the first gradient 530.
[0163] Determining the second gradient includes: determining the derivative 560 of the loss function 540 with respect to the objective function. Among them, the loss function 540 can be expressed as , Taking it as the objective function, the derivative 560 can be expressed as ; then input the first output data 550 into the derivative 560 to obtain the second gradient 570; then multiply the first gradient 530 and the second gradient 570 to obtain the current parameter gradient 580.
[0164] Based on the characteristics and parameter offset method of optical computing implemented by the MZI, the present invention provides a method for training using gradient descent of an optical neural network model. This method realizes the efficient and accurate calculation of parameter gradients by obtaining the outputs of the optical neural network model after adding or subtracting parameter offsets to the current phase parameters, such as the forward offset output data and the negative offset output data, and further calculating the outputs, thereby improving the model training efficiency.
[0165] Based on the above training method of the optical neural network, the present invention also provides a training device for the optical neural network. The following will be combined with Figure 6 to describe this device in detail.
[0166] Figure 6 is a structural block diagram of a training device for an optical neural network according to an embodiment of the present invention.
[0167] As Figure 6 shown, the optical neural network of this embodiment has at least one optical interference unit, and the optical interference unit includes at least one phase parameter. The training device 600 of the optical neural network includes a first input-output module 610, a parameter offset module 620, a second input-output module 630, a determination module 640, and an optimization module 650.
[0168] The first input-output module 610 is configured to input a training optical signal corresponding to a training sample into the optical neural network and output first output data for the training sample.
[0169] The parameter offset module 620 is configured to perform parameter offset processing on the current phase parameters of at least one optical interference unit by using a parameter offset amount to obtain an optical neural network with offset parameters.
[0170] The second input-output module 630 is configured to second input a training optical signal corresponding to a training sample into the optical neural network with offset parameters and output second output data for the training sample.
[0171] The determination module 640 is configured to determine the current parameter gradient for the current phase parameters according to the first output data, the second output data, and a preset loss function.
[0172] An optimization module 650 is configured to optimize the current phase parameters of the optical neural network by using the current parameter gradient, so as to obtain an optimized optical neural network.
[0173] According to an embodiment of the present invention, the optical neural network with parameter offset includes a forward-offset optical neural network and a backward-offset optical neural network.
[0174] According to an embodiment of the present invention, the parameter offset module includes: a first offset sub-module and a second offset sub-module.
[0175] The first offset sub-module is configured to add the current phase parameter and the parameter offset amount to obtain a forward-offset optical neural network.
[0176] The second offset sub-module is configured to subtract the parameter offset amount from the current phase parameter to obtain a backward-offset optical neural network.
[0177] According to an embodiment of the present invention, the second output data includes forward-offset output data and backward-offset output data.
[0178] According to an embodiment of the present invention, the second input / output module includes: a first input / output sub-module and a second input / output sub-module.
[0179] The first input / output sub-module is configured to input a training optical signal corresponding to a training sample into the forward-offset optical neural network and output forward-offset output data.
[0180] The second input / output sub-module is configured to input a training optical signal corresponding to a training sample into the backward-offset optical neural network and output backward-offset output data.
[0181] According to an embodiment of the present invention, the output of the optical neural network is represented by an objective function, and the objective function is a function of the phase parameter.
[0182] According to an embodiment of the present invention, the determination module includes: a first processing sub-module, a first determination sub-module, and a second determination sub-module.
[0183] The first processing sub-module is configured to process the second output data by using a preset parameter offset formula to obtain a first gradient of the objective function with respect to the current phase parameter.
[0184] The first determination sub-module is configured to determine a second gradient of the loss function with respect to the first output data.
[0185] The second determination sub-module is configured to determine a current parameter gradient for the current phase parameter according to the first gradient and the second gradient.
[0186] According to an embodiment of the present invention, the second output data includes positive offset output data and negative offset output data.
[0187] According to an embodiment of the present invention, the first processing sub-module includes: a difference multiplication unit.
[0188] The difference multiplication unit is configured to multiply the difference between the positive offset output data and the negative offset output data by a preset value to obtain a first gradient.
[0189] According to an embodiment of the present invention, the loss function is a function of the objective function.
[0190] According to an embodiment of the present invention, the first determination sub-module includes: a first determination unit and a first processing unit.
[0191] The first determination unit is configured to determine the derivative of the loss function with respect to the objective function.
[0192] The first processing unit is configured to process the first output data using the derivative to obtain a second gradient.
[0193] According to an embodiment of the present invention, the second determination sub-module includes:
[0194] A multiplication unit configured to multiply the first gradient and the second gradient to obtain a current parameter gradient.
[0195] According to an embodiment of the present invention, the optical interference unit includes at least one phase shifter, and the phase shifter is represented by a transfer matrix; the parameter offset is determined by the following module:
[0196] A processing module configured to process the transfer matrix using an offset calculation formula to determine the parameter offset.
[0197] According to an embodiment of the present invention, the offset calculation formula includes matrix transformation parameters, and the matrix transformation parameters are represented by a matrix transformation formula.
[0198] According to an embodiment of the present invention, the processing module includes: a conversion sub-module, a second processing sub-module, and a third determination sub-module.
[0199] The conversion sub-module is configured to convert the transfer matrix into an exponential form of a preset format, where the exponential form includes a Hermitian matrix.
[0200] The second processing sub-module is configured to process the Hermitian matrix using the matrix transformation formula to determine the parameter value of the matrix transformation parameter.
[0201] The third determination sub-module is configured to determine the parameter offset according to the parameter value and the offset calculation formula.
[0202] According to an embodiment of the present invention, the second processing sub-module includes: a second determination unit and a second processing unit.
[0203] A second determination unit, configured to determine two eigenvalues of a Hermitian matrix.
[0204] A second processing unit, configured to process the two eigenvalues by using a matrix transformation formula to obtain parameter values of matrix transformation parameters.
[0205] According to an embodiment of the present invention, the optimization module includes: a third processing sub-module and an adjustment sub-module.
[0206] The third processing sub-module is configured to process a current parameter gradient and a current phase parameter by using a preset gradient descent formula to obtain an optimized phase parameter.
[0207] The adjustment sub-module is configured to adjust a current phase parameter of at least one optical interference unit according to the optimized phase parameter to obtain an optimized optical neural network.
[0208] According to an embodiment of the present invention, the adjustment sub-module includes: an adjustment unit.
[0209] The adjustment unit is configured to, for each optical interference unit, adjust the current phase parameter of the optical interference unit by adjusting the voltage and / or temperature of the optical interference unit according to the optimized phase parameter for the optical interference unit.
[0210] According to an embodiment of the present invention, any plurality of modules among the first input / output module 610, the parameter offset module 620, the second input / output module 630, the determination module 640, and the optimization module 650 may be combined and implemented in one module, or any one of them may be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present invention, at least one of the first input / output module 610, the parameter offset module 620, the second input / output module 630, the determination module 640, and the optimization module 650 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or any other reasonable way of integrating or packaging circuits, etc., implemented by hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Alternatively, at least one of the first input / output module 610, the parameter offset module 620, the second input / output module 630, the determination module 640, and the optimization module 650 may be at least partially implemented as a computer program module, and when the computer program module is run, the corresponding functions may be executed.
[0211] Figure 7 It is a block diagram of an electronic device suitable for implementing the above method according to an embodiment of the present invention.
[0212] As Figure 7 shown, the electronic device 700 according to an embodiment of the present invention includes a processor 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage section 708 into a random access memory (RAM) 703. The processor 701 may include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), etc. The processor 701 may also include on-board memory for caching purposes. The processor 701 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.
[0213] In the RAM 703, various programs and data required for the operation of the electronic device 700 are stored. The processor 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. The processor 701 performs various operations of the method flow according to an embodiment of the present invention by executing the programs in the ROM 702 and / or the RAM 703. It should be noted that the program may also be stored in one or more memories other than the ROM 702 and the RAM 703. The processor 701 may also perform various operations of the method flow according to an embodiment of the present invention by executing the programs stored in the one or more memories.
[0214] According to an embodiment of the present invention, the electronic device 700 may further include an input / output (I / O) interface 705, and the input / output (I / O) interface 705 is also connected to the bus 704. The electronic device 700 may further include one or more of the following components connected to the input / output (I / O) interface 705: an input section 706 including a keyboard, a mouse, etc.; an output section 707 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, a modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the input / output (I / O) interface 705 as needed. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 710 as needed so that a computer program read from it can be installed into the storage section 708 as needed.
[0215] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist independently without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of the present invention is implemented.
[0216] According to an embodiment of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium, and may include, for example, but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present invention, the computer-readable storage medium may include the above-described ROM 702 and / or RAM 703 and / or one or more memories other than ROM 702 and RAM 703.
[0217] An embodiment of the present invention further includes a computer program product, which includes a computer program that contains program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to cause the computer system to implement the method provided by the embodiments of the present invention.
[0218] When the computer program is executed by the processor 701, the above functions defined in the system / apparatus of the embodiments of the present invention are executed. According to an embodiment of the present invention, the above-described systems, apparatuses, modules, units, etc. may be implemented by computer program modules.
[0219] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and be downloaded and installed through the communication part 709, and / or be installed from the removable medium 711. The program code included in the computer program may be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0220] In such an embodiment, the computer program can be downloaded and installed from a network through the communication part 709, and / or installed from the removable medium 711. When the computer program is executed by the processor 701, the above functions defined in the system of the embodiments of the present invention are executed. According to the embodiments of the present invention, the systems, devices, apparatuses, modules, units, etc. described above can be implemented by computer program modules.
[0221] According to the embodiments of the present invention, the program code for executing the computer program provided by the embodiments of the present invention can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedures and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include but are not limited to, such as Java, C++, python, the "C" language, or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, by connecting through an Internet service provider via the Internet).
[0222] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combinations of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0223] Those skilled in the art can understand that the features described in the various embodiments of the present invention can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features described in the various embodiments of the present invention can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.
[0224] The embodiments of the present invention have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Without departing from the scope of the present invention, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present invention.
Claims
1. A method for training an optical neural network, characterized in that: The optical neural network includes at least one optical interference unit, and the optical interference unit has at least one phase parameter; the training method includes: Iteratively perform the following operations until the iteration stop condition is reached: The optical neural network is used as the current optical neural network, and a parameter offset is used to perform parameter offset processing on the current phase parameter of the at least one optical interference unit included in the current optical neural network to obtain an optical neural network after parameter offset; Inputting the training light signal corresponding to the training sample into the optical neural network and the optical neural network after the parameter shift respectively, and outputting the first output data and the second output data for the training sample; Determining a current parameter gradient for the current phase parameter according to the first output data, the second output data and a preset loss function; and The current phase parameter of the optical neural network is optimized using the current parameter gradient to obtain an optimized optical neural network, and the optimized optical neural network is used as the current optical neural network.
2. The training method according to claim 1, characterized in that: The optical neural network after parameter shift includes an optical neural network after positive shift and an optical neural network after negative shift; The step of performing parameter offset processing on the current phase parameter of the at least one optical interference unit by using the parameter offset to obtain the optical neural network after the parameter offset comprises: Adding the current phase parameter to the parameter offset to obtain the optical neural network after the forward offset; The current phase parameter is subtracted from the parameter offset to obtain the optical neural network after the negative offset.
3. The training method according to claim 2, characterized in that: The second output data includes positive offset output data and negative offset output data; Inputting the training light signal corresponding to the training sample into the optical neural network after parameter shift, and outputting second output data for the training sample comprises: Inputting a training optical signal corresponding to the training sample into the optical neural network after forward shift, and outputting the forward shift output data; The training light signal corresponding to the training sample is input into the negatively offset optical neural network, and the negatively offset output data is output.
4. The training method according to claim 1, characterized in that: The output of the optical neural network is represented by an objective function, and the objective function is a function of the phase parameter; and determining a current parameter gradient for the current phase parameter according to the first output data, the second output data, and a preset loss function comprises: Processing the second output data using a preset parameter offset formula to obtain a first gradient of the objective function relative to the current phase parameter; determining a second gradient of the loss function with respect to the first output data; A current parameter gradient for the current phase parameter is determined according to the first gradient and the second gradient.
5. The training method according to claim 4, characterized in that: The second output data includes positive offset output data and negative offset output data; The step of processing the second output data by using a preset parameter offset formula to obtain a first gradient of the objective function relative to the current phase parameter comprises: The first gradient is obtained by multiplying a difference between the positive offset output data and the negative offset output data by a preset value.
6. The training method according to claim 4, characterized in that: The loss function is a function of the objective function; Determining a second gradient of the loss function relative to the first output data comprises: Determining a derivative of the loss function with respect to the objective function; The first output data is processed using the derivative to obtain the second gradient.
7. The training method according to any one of claims 4 to 6, characterized in that: Determining a current parameter gradient for the current phase parameter according to the first gradient and the second gradient includes: The first gradient and the second gradient are multiplied to obtain the current parameter gradient.
8. The training method according to claim 1, characterized in that: The optical interference unit includes at least one phase shifter, and the phase shifter is represented by a transfer matrix; the parameter offset is determined by the following steps: The transfer matrix is processed using an offset calculation formula to determine the parameter offset.
9. The training method according to claim 8, characterized in that: The offset calculation formula includes matrix transformation parameters, and the matrix transformation parameters are expressed by a matrix transformation formula; The processing of the transfer matrix by using the offset calculation formula comprises: Converting the transfer matrix into an exponential form of a preset format, wherein the exponential form includes a Hermitian matrix; Processing the Hermitian matrix using the matrix transformation formula to determine parameter values of the matrix transformation parameters; The parameter offset is determined according to the parameter value and the offset calculation formula.
10. The training method according to claim 9, characterized in that: The using the matrix transformation formula to process the Hermitian matrix and determine the parameter value of the matrix transformation parameter comprises: determining two eigenvalues for the Hermitian matrix; The two eigenvalues are processed using the matrix transformation formula to obtain parameter values of the matrix transformation parameters.
11. The training method according to claim 5, characterized in that: The step of optimizing the current phase parameter of the optical neural network by using the current parameter gradient to obtain an optimized optical neural network comprises: Processing the current parameter gradient and the current phase parameter using a preset gradient descent formula to obtain an optimized phase parameter; The current phase parameter of the at least one optical interference unit is adjusted according to the optimized phase parameter to obtain the optimized optical neural network.
12. The training method according to claim 11, characterized in that: The step of adjusting the current phase parameter of the at least one optical interference unit according to the optimized phase parameter comprises: For each of the optical interference units, according to the optimized phase parameter for the optical interference unit, the current phase parameter of the optical interference unit is adjusted by adjusting the voltage and / or temperature of the optical interference unit.
13. An electronic device, comprising: one or more processors; a memory for storing one or more computer programs, It is characterized in that the one or more processors execute the one or more computer programs to implement the steps of the training method according to any one of claims 1 to 12.
14. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instruction is executed by a processor, the steps of the training method according to any one of claims 1 to 12 are implemented.
15. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the steps of the training method according to any one of claims 1 to 12 are implemented.
Citation Information
Patent Citations
Model training method and device andimage recognition method and device
CN113657596A