Optical neural network training method and system

By constructing a local training closed loop for the optical neural network and utilizing coherent light processing to achieve gradient calculation and parameter updates, the problems of high training latency and sensitivity to errors in existing technologies are solved, thereby improving the training efficiency and stability of the optical neural network.

CN122491377APending Publication Date: 2026-07-31UNIV OF SHANGHAI FOR SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
UNIV OF SHANGHAI FOR SCI & TECH
Filing Date
2026-05-07
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing optical neural network training methods rely on strict reverse physical propagation conditions and high digital assistance during actual hardware deployment and in-situ learning, resulting in high training latency, sensitivity to minor manufacturing errors and environmental disturbances, and difficulty in achieving all-optical in-situ learning.

Method used

By introducing a coherent light source, a complex amplitude coding unit, a physical modulation layer, a local error physical calculation unit, and a parallel gradient interferometry acquisition unit, a local training closed loop is constructed to reduce the dependence on backpropagation of global errors. Gradient calculation and parameter updates are completed using coherent processing in the optical domain.

Benefits of technology

It improves the adaptability of in-situ training of optical neural networks and the stability of training control, reduces the dependence on strict reverse physical propagation conditions, and enhances training efficiency and the physical robustness of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122491377A_ABST
    Figure CN122491377A_ABST
Patent Text Reader

Abstract

This application provides an optical neural network training method and system, belonging to the field of optical computing. The system includes: a coherent light source unit for generating a coherent light beam; a complex amplitude encoding unit for generating a target light field based on a virtual target field; a physical modulation layer for forward modulation of the training input light field to output the actual output light field corresponding to the layer to be updated; a local error physical calculation unit for interferometry processing of the actual output light field and the target light field to obtain the local error field corresponding to the layer to be updated; a parallel gradient interferometry acquisition unit for extracting gradient information corresponding to the layer to be updated based on the interferogram corresponding to the local error field; and a central controller for determining error processing information, determining the virtual target field of the layer to be updated based on the actual output light field and the error processing information, and updating the forward modulation parameters of the layer to be updated based on the gradient information. This application can improve the adaptability and stability of in-situ training of optical neural networks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of optical computing technology, and in particular to an optical neural network training method and system. Background Technology

[0002] Optical neural networks have attracted attention in the field of artificial intelligence computing due to their parallel processing capabilities and low power consumption. Current training of optical neural networks typically relies on backpropagation algorithms, and some implementations also require the use of digital processing units to assist in the calculation of intermediate variables, loss information, or gradient information. For practical optical hardware and in-situ training scenarios, the backpropagation of error signals along the forward propagation path usually depends on relatively strict physical propagation conditions. In cases involving non-reciprocal devices, spatial filtering structures, or optical path alignment deviations, existing training methods may struggle to stably adapt to real-world physical systems. Furthermore, the construction of physical operators related to weight matrix transposition or backpropagation in optical physical entities is complex, easily leading to additional hardware configuration and computational overhead. On the other hand, existing training processes often require waiting for global forward propagation and loss calculation to complete before performing layer-by-layer updates, which can result in high training latency in deep architectures. Moreover, traditional training results are highly sensitive to minor manufacturing errors, environmental disturbances, or alignment deviations, and some solutions still heavily rely on electronic computing resources for gradient calculation and intermediate information processing. Therefore, how to reduce the dependence of the training process on strict reverse physical propagation conditions and high digital assistance in the in-situ training scenario of optical neural networks, while taking into account both training efficiency and physical implementation adaptability, has become one of the technical problems to be addressed in related technologies.

[0003] It should be noted that the above content is not necessarily prior art, nor is it intended to limit the scope of patent protection of this application. Summary of the Invention

[0004] This application provides an optical neural network training method and system to solve or alleviate one or more of the technical problems mentioned above.

[0005] One aspect of this application provides an optical neural network training system, including: A coherent light source unit is used to generate coherent light beams; Complex amplitude encoding unit, used to generate target light field based on virtual target field; A physical modulation layer is used to forward modulate the training input light field to output the actual output light field corresponding to the layer to be updated. The physical modulation layer includes at least a forward modulation component and an error convolution component. The local error physical calculation unit is used to perform interference processing on the actual output light field and the target light field to obtain the local error field corresponding to the layer to be updated. A parallel gradient interferometry acquisition unit is used to interfere the local error field with the modulation phase term that characterizes the modulation state of the layer to be updated, and to extract the gradient information corresponding to the layer to be updated based on the interferogram corresponding to the local error field. The central controller is used to determine error processing information based on the local error information of the next layer and the error convolution unit of the next layer, determine the virtual target field of the layer to be updated according to the actual output light field and the error processing information, and update the forward modulation parameters of the layer to be updated according to the gradient information.

[0006] Optionally, the physical modulation layer includes multiple cascaded diffraction modulation sub-layers, each of which is provided with the forward modulation component and the error convolution component.

[0007] Optionally, the complex amplitude encoding unit includes a spatial light modulator, and the central controller is used to control the spatial light modulator to encode the virtual target field in order to output the target light field.

[0008] Optionally, the complex amplitude encoding unit includes a dual-phase holographic encoding spatial light modulator and a low-pass filter component, wherein the low-pass filter component is used to reconstruct the complex amplitude of the encoded light field.

[0009] Optionally, the local error physical calculation unit includes a beam splitter, a phase shifter, and a detector. The phase shifter is used to introduce a π phase shift into one of the optical fields participating in the interference, so that the actual output optical field and the target optical field undergo destructive interference.

[0010] Optionally, the parallel gradient interferometry acquisition unit includes a phase shifter and a detector. The parallel gradient interferometry acquisition unit is used to cause the local error field to interfere with the modulation phase term under multiple phase shift conditions, and to record multiple frames of original interferometric patterns under multiple phase shift conditions. The central controller is used to extract the gradient information based on the multiple frames of original interferometric patterns.

[0011] Optionally, the central controller is further configured to perform complex field recovery processing on the multi-frame original interferometric patterns to obtain a complex field recovery result, the complex field recovery result including a complex gradient field.

[0012] Optionally, the central controller is further configured to perform smoothing processing on the phase distribution corresponding to the complex gradient field in the complex field recovery result, so as to suppress high-frequency noise and induce the formation of a smooth phase distribution.

[0013] Optionally, the central controller is further configured to simultaneously update the error convolution parameters of the layer to be updated while updating the forward modulation parameters of the layer to be updated.

[0014] Optionally, the central controller is also used to employ an asynchronous parallel pipeline scheduling method, so that the m-th layer and the (m+1)-th layer simultaneously perform local error calculation, gradient acquisition and parameter refresh, where m is an integer greater than or equal to 1.

[0015] Optionally, the physical modulation layer includes at least one of a programmable spatial light modulator and a three-dimensional diffraction chip formed by two-photon polymerization.

[0016] Another aspect of this application provides an optical neural network training method, applied to an optical neural network training system, the method comprising: The training input light field is forward modulated to output the actual output light field corresponding to the layer to be updated; Error processing information is determined based on the local error information of the next layer and the error convolution component of the next layer, and the virtual target field of the layer to be updated is determined based on the actual output light field and the error processing information. Generate a target light field based on the virtual target field; The actual output light field and the target light field are subjected to interference processing to obtain the local error field corresponding to the layer to be updated; Based on the interferogram corresponding to the local error field, extract the gradient information corresponding to the layer to be updated. The forward modulation parameters of the layer to be updated are updated based on the gradient information.

[0017] Optionally, determining the error processing information based on the local error information of the next layer and the error convolution unit of the next layer includes: Obtain the local error information of the next layer; A correction coefficient is applied to the error convolution component of the next layer, and a Fourier transform is performed on the error convolution component after the correction coefficient is applied; The transformed result is then processed with the local error information of the next layer to obtain error processing information.

[0018] Optionally, generating the target light field based on the virtual target field includes: The spatial light modulator is controlled to encode the virtual target field; Output the encoded target light field.

[0019] Optionally, generating the target light field based on the virtual target field includes: The virtual target field is decomposed into two phase components using dual-phase holographic coding technology; The decomposed light field is reconstructed by using a low-pass filter to obtain the target light field.

[0020] Optionally, the step of interferometric processing of the actual output light field and the target light field to obtain the local error field corresponding to the layer to be updated includes: The actual output light field and the target light field are respectively connected to the two input terminals of the local error physical calculation unit; The phase shifter is controlled to introduce a π phase shift into one of the optical fields involved in the interference; This causes the two coherent beams, after being phase-shifted, to undergo destructive interference. Obtain the local error field after interference.

[0021] Optionally, extracting the gradient information corresponding to the layer to be updated based on the interferogram corresponding to the local error field includes: The local error field is spatially interfered with the modulation phase term using phase-shifting interferometry. Record multiple frames of original interferograms under various phase shift conditions; The original interferometric patterns of the multiple frames are demodulated to extract the gradient information.

[0022] Optionally, the method further includes: The original interferograms of the multiple frames are subjected to complex field recovery processing to obtain a complex field recovery result, which includes a complex gradient field.

[0023] Optionally, the method further includes: The complex field recovery result is smoothed to suppress high-frequency noise and induce a smooth phase distribution.

[0024] Optionally, updating the forward modulation parameters of the layer to be updated based on the gradient information includes: Update the forward modulation parameters of the layer to be updated based on the gradient information; The error convolution parameters of the layer to be updated are updated synchronously.

[0025] Optionally, the method further includes: When performing gradient updates at the (m+1)th layer, gradient acquisition and parameter refresh are simultaneously triggered at the mth layer, where m is an integer greater than or equal to 1.

[0026] Optionally, the forward modulation of the training input light field to output the actual output light field corresponding to the layer to be updated includes: The training input light field is spatially phase modulated by multiple cascaded diffraction modulation sublayers. The actual output light field is obtained on the detection plane corresponding to the layer to be updated.

[0027] Optionally, before forward modulation of the training input light field, the method further includes: The physical modulation layer is initialized to a random phase distribution or a plane wave state; Initialize the forward modulation parameters to zero; The error convolution parameters are initialized to uniformly distributed random noise.

[0028] The embodiments of this application employing the above technical solution may include the following advantages: the central controller determines the virtual target field of the layer to be updated based on the local error information of the next layer and the error convolution component of the layer to be updated, and the complex amplitude encoding unit generates the target light field. Then, the local error physical calculation unit performs interference processing on the actual output light field and the target light field to obtain the local error field. The parallel gradient interference acquisition unit acquires gradient information based on the interference pattern corresponding to the local error field and updates the forward modulation parameters of the layer to be updated accordingly. This enables the formation of a local training closed loop around the layer to be updated, from target reference construction, local error acquisition, gradient information acquisition to parameter update. Based on this local training closed loop, the dependence of the training process on the backpropagation of global error along the complete forward propagation path can be reduced, and it is beneficial to combine the parameter update process with the actual optical training object, thereby improving the adaptability of in-situ training of optical neural networks and the stability of training control. Attached Figure Description

[0029] The accompanying drawings exemplify embodiments and form part of the specification, serving together with the textual description to explain exemplary implementations of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0030] Figure 1 A schematic diagram illustrating the architecture of an optical neural network training system according to Embodiment 1 of this application is shown. Figure 2 A flowchart illustrating an optical neural network training method according to Embodiment 2 of this application is shown schematically. Figure 3 Schematic illustration Figure 2 A detailed flowchart of the steps for determining error processing information based on the local error information of the next layer and the error convolution component of the next layer; Figure 4 Schematic illustration Figure 2 A detailed flowchart of the steps for generating a target light field based on the virtual target field is provided below. Figure 5 Schematic illustration Figure 2 A detailed flowchart of the steps for generating a target light field based on the virtual target field is provided below. Figure 6 Schematic illustration Figure 2The detailed flowchart of the steps for interfering the actual output light field and the target light field to obtain the local error field corresponding to the layer to be updated is as follows: Figure 7 Schematic illustration Figure 2 A detailed flowchart of the steps for extracting gradient information of the layer to be updated based on the interferogram corresponding to the local error field is provided. Figure 8 The flowchart illustrating the additional steps of the optical neural network training method according to Embodiment 2 of this application is shown in the schematic diagram. Figure 9 A schematic diagram of the phase distribution is shown. Detailed Implementation

[0031] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application. All other embodiments obtained by those skilled in the art based on the embodiments in this application without inventive effort are within the scope of protection of this application.

[0032] It should be noted that the descriptions involving "first," "second," etc., in the embodiments of this application are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined with "first" or "second" may explicitly or implicitly include at least one of that feature. Furthermore, the technical solutions of the various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. If the combination of technical solutions is contradictory or impossible to implement, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed in this application.

[0033] It should be noted that, in any stage of this application involving the collection, storage, use, transmission, and processing of data, each stage strictly adheres to the laws, regulations, industry standards, and regulatory requirements of the data source, usage location, and relevant countries and regions to ensure the legality and compliance of data activities. In the collection stage, the purpose, method, and scope of collection are clearly communicated to the data subject in a prominent manner. Collection is conducted only after obtaining the data subject's legal authorization, ensuring that the collection process follows the "minimum necessary" principle and does not exceed the scope of data collection. In the storage stage, storage periods are limited, and data is promptly deleted or anonymized / encrypted after the storage purpose is achieved. In the usage stage, a strict data security protection mechanism is implemented, using field-level desensitization technology and processing the original data according to preset desensitization rules. For different types of data, multiple desensitization strategies, such as data generalization, data anonymization, and data encryption, are employed to effectively mitigate the risk of sensitive information leakage and ensure that all data used is securely processed and desensitized, comprehensively protecting the rights and interests of data subjects and data security. In the transmission and processing stages, the confidentiality and security of data are ensured during transmission and processing.

[0034] In the description of this application, it should be understood that the numerical labels before the steps do not indicate the order of the steps, but are only used to facilitate the description of this application and to distinguish each step, and therefore should not be construed as a limitation of this application.

[0035] First, a definition of the terminology used in this application is provided: Optical Neural Networks (ONNs) are a class of physical systems that use photons instead of electrons to perform neural network computation tasks. They aim to overcome the physical limits of traditional electronic chips in terms of bandwidth, latency, and power consumption by mapping the mathematical operations of neural networks (such as matrix multiplication and nonlinear activation) to optical physical phenomena (such as interference, diffraction, and modulation).

[0036] Diffractive Deep Neural Networks (D²NNs) are all-optical machine learning architectures built on the principles of physical optical diffraction. They are physical neural networks composed of multiple stacked diffractive layers. Each layer contains thousands of programmable neurons, which are essentially transparent or opaque pixels (or phase modulation units). When light waves pass through these layers, physical diffraction occurs, forming a specific intensity distribution on the output plane, thus achieving a mapping from input to output (such as image classification).

[0037] Spatial Light Modulator (SLM): A photoelectric device that can dynamically control the spatial distribution (amplitude, phase, or polarization) of a light beam at the pixel level.

[0038] Virtual target field: In the local target propagation training of an optical neural network, the virtual target field refers to the ideal complex amplitude optical field data calculated by the central controller based on the actual output optical field of the layer to be updated, the local error field of the next layer (m+1 layer), and the error convolution unit, using local learning rules. This virtual target field is the locally ideal output optical field set for the layer to be updated.

[0039] Secondly, to facilitate understanding of the technical solutions provided in the embodiments of this application by those skilled in the art, the relevant technologies are described below: Optical neural networks (ONNs) have shown great potential in the field of artificial intelligence computing due to their high-speed parallel processing capabilities and low power consumption. Currently, mainstream physical architectures such as diffractive deep neural networks (D²NNs) primarily rely on digitally assisted backpropagation (BP) algorithms for training. However, in practical hardware deployment and in-situ training, existing technologies suffer from the following significant drawbacks: 1. Physical Path Reciprocity Constraint. Traditional backpropagation (BP) algorithms require the error signal to propagate back along the reverse path of forward propagation, which physically demands strict optical time-reversal symmetry. However, practical optical hardware commonly contains non-reciprocal components (such as isolators), spatial filters, and optical path alignment errors, causing the physically propagated error field to deviate significantly from the theoretical gradient direction, resulting in training failing to converge.

[0040] 2. Global Coupling Delay. Traditional architectures require waiting for the optical field to complete its forward propagation across the entire network and for electronic devices to calculate the global loss before initiating layer-by-layer gradient backpropagation. This strong coupling logic limits training throughput and makes it difficult to support parallel updates in deep architectures.

[0041] 3. Extreme sensitivity to physical defects. Because traditional gradient descent methods tend to learn high-frequency interference modes, the resulting phase masks often contain a large number of fine, sharp structures. This mode can cause the optical feature extraction capability to collapse rapidly when faced with micro / nano fabrication errors (such as edge roughness) or slight misalignments caused by environmental vibrations.

[0042] 4. High reliance on digitalization makes true all-optical in-situ learning difficult. While existing technologies attempt to perform some computations in the optical domain, they still heavily rely on high-bandwidth digital memory and electronic processors for gradient calculation and intermediate variable storage. This not only generates additional energy consumption but also fails to fully utilize the low-latency advantage of optical parallel computing, resulting in the overall system energy efficiency falling short of optimal levels.

[0043] To address this, this application provides an optical neural network training system. In this system, by introducing independent trainable error kernels, a local loss function is constructed for each layer, achieving physical localization of gradient calculation and completely eliminating the dependence on global error chain differentiation and physical reciprocity paths. Utilizing the superposition characteristics of coherent light fields, complex subtraction and gradient sampling are directly performed in the optical domain through physical coherence. Compared to existing technologies, this application avoids phase information loss due to photoelectric conversion, improving the fidelity of gradient updates. Through weight mirroring rules, the system learns a smooth and continuous physical phase profile, significantly improving physical robustness and fabrication friendliness. Experiments demonstrate that this application has extremely strong tolerance to phase noise and lateral misalignment, and can naturally adapt to the rounding effect in two-photon polymerization (TPP) fabrication, ensuring efficient migration of laboratory models to integrated chips. Furthermore, this application supports independent weight refresh at each physical layer in pipelined mode through software scheduling logic, without waiting for global feedback. This significantly improves the training efficiency and scalability of large-scale optical intelligent systems and supports high-throughput asynchronous parallel updates. See below for details.

[0044] The technical solutions of this application are described below through several embodiments. It should be understood that these embodiments can be implemented in many different forms and should not be construed as being limited to the embodiments set forth herein.

[0045] Example 1 Figure 1 A schematic diagram of an optical neural network training system according to Embodiment 1 of this application is shown.

[0046] The optical neural network training system 1 includes a coherent light source unit 10, a complex amplitude encoding unit 11, a physical modulation layer 12, a local error physical calculation unit 13, a parallel gradient interference acquisition unit 14, and a central controller 15.

[0047] The functions and coordination of the coherent light source unit 10, complex amplitude encoding unit 11, physical modulation layer 12, local error physical calculation unit 13, parallel gradient interferometry acquisition unit 14, and central controller 15 will be introduced below.

[0048] The coherent light source unit 10 is used to generate a coherent beam, providing a stable light source for subsequent interferometric calculations.

[0049] In this embodiment, the coherent light source unit 10 may include a laser and a beam expander collimator. The laser is used to generate laser light. The laser may be a continuous-wave laser with a narrow linewidth of 532 nm and a single longitudinal mode, with an adjustable output power of 10–50 mW. The beam expander collimator is used to enlarge the aperture of the narrow Gaussian beam output from the laser and convert it into a uniform parallel plane wave, ensuring the coherence length and spatial uniformity of the optical field, and providing a stable light source for subsequent calculations.

[0050] Complex amplitude encoding unit 11 is used to generate target light field based on virtual target field.

[0051] In this embodiment, the complex amplitude encoding unit 11 receives virtual target field data sent by the central controller 15 and converts the digital virtual target field into a target light field that can be stably transmitted in the optical path through optical field modulation, providing a standard reference light field for subsequent local error calculation.

[0052] In an optional implementation, the complex amplitude encoding unit 11 includes a spatial light modulator, and the central controller 15 is used to control the spatial light modulator to encode the virtual target field in order to output the target light field.

[0053] In this embodiment, the complex amplitude encoding unit 11 is configured with a spatial light modulator as the core modulation device. The spatial light modulator and the central controller achieve high-speed command interaction through the drive interface. The central controller converts the calculated virtual target field into a phase-coded grayscale image and loads it into the spatial light modulator in real time. The spatial light modulator performs phase modulation on the coherent beam incident from the coherent light source unit according to the encoding image, and outputs a target light field whose amplitude and phase are completely matched with the virtual target field, thereby realizing the accurate mapping from the digital target field to the optical target light field.

[0054] This embodiment achieves programmable generation of the target light field through a spatial light modulator, improving the precision and response speed of light field control and ensuring the baseline accuracy of local error calculation.

[0055] In an optional implementation, the complex amplitude encoding unit 11 includes a dual-phase holographic encoding spatial light modulator and a low-pass filter component, wherein the low-pass filter component is used to reconstruct the complex amplitude of the encoded light field.

[0056] In this embodiment, the complex amplitude encoding unit 11 is composed of a dual-phase holographic encoding spatial light modulator and a low-pass filter component. The dual-phase holographic encoding spatial light modulator converts the complex numerical virtual target field. The light field is decomposed into two pure phase components for encoding to avoid energy loss and phase distortion caused by direct modulation of complex amplitude. The encoded light field is then fed into the 4f system. The low-pass filter in the 4f system filters the diffracted light field, removes high-frequency diffraction noise, and completes the reconstruction of the complex amplitude light field, outputting a high-purity, low-distortion target light field.

[0057] This embodiment combines dual-phase encoding with low-pass complex amplitude reconstruction, which can completely preserve the phase and amplitude information of the virtual target field, thereby improving the quality of the target light field and the reliability of error calculation.

[0058] The physical modulation layer 12 is used to forward modulate the training input light field to output the actual output light field corresponding to the layer to be updated. The physical modulation layer includes at least a forward modulation component and an error convolution component.

[0059] In an optional implementation, the physical modulation layer 12 includes a plurality of cascaded diffraction modulation sub-layers, each of which is provided with the forward modulation component and the error convolution component.

[0060] In this embodiment, the physical modulation layer 12 of the optical neural network training system adopts a multi-cascade structure, setting up multiple diffraction modulation sub-layers arranged sequentially along the optical path. Each diffraction modulation sub-layer is independently configured with a forward modulation component and an error convolution component. The forward modulation component is responsible for feature transformation and spatial modulation of the light field, and the error convolution component is used to support the forward propagation of local targets between layers. Seamless transmission and spatial matching of the light field can be achieved between each modulation sub-layer through a 4f optical coupling system. The central controller can independently perform virtual target field calculation, local error extraction, gradient acquisition and parameter update for each modulation sub-layer, so that the deep network can still have the ability to learn independently in layers while maintaining strong feature expression capabilities.

[0061] This embodiment enhances the network's feature representation capability by cascading multi-level diffraction modulation sub-layers, while maintaining the independence of local learning in each layer, thereby improving the system's depth scalability and training stability.

[0062] In this embodiment, the physical modulation layer 12 integrates a forward modulation component and an error convolution component. The forward modulation component is used to perform spatial phase modulation on the training input light field to complete the extraction and transformation of light field features, and the error convolution component is used to assist in the forward transmission of local targets between layers. After modulation, the actual output light field corresponding to the layer to be updated is output.

[0063] In an alternative implementation, the physical modulation layer includes at least one of a programmable spatial light modulator and a three-dimensional diffraction chip formed by two-photon polymerization (TPP).

[0064] This embodiment uses a programmable spatial light modulator or a three-dimensional diffraction chip as the physical modulation layer, which can balance experimental flexibility and product deployment capabilities, and broaden the application scenarios of the system.

[0065] The local error physical calculation unit 13 is used to perform interference processing on the actual output light field and the target light field to obtain the local error field corresponding to the layer to be updated.

[0066] In this embodiment, the local error physical calculation unit 13 introduces the actual output light field and the target light field into the coherent interference structure, and performs destructive interference processing (differential operation) directly in the optical domain through coherent superposition of the light fields to obtain a local error field that can characterize the output deviation of the layer to be updated (current layer).

[0067] In an optional implementation, the local error physical calculation unit 13 includes a beam splitter, a phase shifter, and a detector. The phase shifter is used to introduce a π phase shift into one of the optical fields involved in the interference, so that the actual output optical field and the target optical field undergo destructive interference.

[0068] In this embodiment, the local error physical calculation unit 13 can adopt a Mach-Zehnder interferometer structure, which includes a beam splitter, a phase shifter, and a detector. The actual output light field and the target light field enter the two input arms of the interferometer, respectively. The phase shifter introduces a π phase shift in one arm to achieve the destructive interference condition. The two light fields converge at the output surface of the beam splitter, resulting in destructive interference. The following complex subtraction operation is directly performed in the optical domain: ,in, This represents the actual output light field of the m-th layer. Let m be the target light field of the m-th layer. Let be the local error field of the m-th layer.

[0069] The detector directly collects the interference light intensity distribution to obtain a local error field that is consistent with the difference between the actual output field and the target light field.

[0070] The phase shifter can be a piezoelectric actuator (PZT). The detector can be a CMOS detector, a CCD (charge-coupled device), etc.

[0071] This embodiment performs complex subtraction directly in the optical domain through π-phase shift destructive interference, eliminating the need for photoelectric conversion and digital computation, thereby improving the speed of acquiring local errors and preserving complete phase information.

[0072] The parallel gradient interference acquisition unit 14 is used to interfere the local error field with the modulation phase term representing the modulation state of the layer to be updated, and to extract the gradient information corresponding to the layer to be updated based on the interference pattern corresponding to the local error field.

[0073] In this embodiment, the parallel gradient interference acquisition unit 14 performs coherent interference between the local error field and the modulation phase term of the layer to be updated (current layer), and extracts gradient information that can be used for parameter updating of the layer to be updated (current layer) by acquiring multiple interference patterns under different phase shift conditions.

[0074] In an optional implementation, the parallel gradient interferometry acquisition unit 14 includes a phase shifter and a detector. The parallel gradient interferometry acquisition unit is used to cause the local error field to interfere with the modulation phase term under multiple phase shift conditions and record multiple frames of original interferograms under multiple phase shift conditions. The central controller is used to extract the gradient information based on the multiple frames of original interferograms.

[0075] In this embodiment, the parallel gradient interferometric acquisition unit 14 can be composed of a precision phase shifter and a high-resolution CMOS detector. The phase shifter sequentially applies multiple sets of phase shift conditions such as 0, π / 2, π, and 3π / 2, so that the local error field and the modulation phase term coherently interfere under different phase shifts. The CMOS detector synchronously records multiple frames of original interference patterns, and the central controller obtains complete gradient information based on the intensity changes of the original interference patterns.

[0076] The modulation phase term is the phase distribution currently being used by the forward modulation component of the layer to be updated.

[0077] This embodiment can simultaneously acquire the magnitude and phase information of the gradient through multi-phase-shift interferometric sampling, solving the problem that complex gradients cannot be extracted from single-frame intensity maps and improving the accuracy of gradient calculation.

[0078] The central controller 15 is used to determine the virtual target field of the layer to be updated (current layer) based on the local error information of the next layer and the error convolution component of the layer to be updated, and to update the forward modulation parameters of the layer to be updated according to the gradient information.

[0079] In this embodiment, the central controller 15 calculates and generates error processing information based on the local error information measured in the next layer and the error convolution component of the next layer. Then, based on the actual output light field and the error processing information, it determines the virtual target field exclusive to the layer to be updated (current layer) and iteratively updates the forward modulation parameters of the layer to be updated (current layer) according to the gradient information.

[0080] It should be noted that the "next layer" mentioned above refers to the next layer in the actual network space order. For example, if the layer to be updated is the m-th layer, then the next layer is the (m+1)-th layer.

[0081] The central controller 15 can calculate the virtual target field using the following formula: ,in, Let m be the virtual target field. This represents the actual output light field of the m-th layer. The error convolution component of the (m+1)th layer, Let m be the local error field of the (m+1)th layer. This is a preset correction factor (e.g., 0.9). right Nuclear correction factor Then, perform a Fourier transform (simulating diffraction propagation).

[0082] The local error information refers to the digital data obtained after detecting the local error field.

[0083] In an optional implementation, the central controller 15 is further configured to perform complex field recovery processing on the multi-frame original interferometric patterns to obtain a complex field recovery result, the complex field recovery result including a complex gradient field.

[0084] In this embodiment, the central controller 15 receives multiple frames of original interferometric patterns acquired by the parallel gradient interferometry acquisition unit 14, processes the patterns using phase demodulation and complex amplitude reconstruction algorithms, performs a complex field recovery operation, and reconstructs the complex gradient field containing the real and imaginary parts from the interferometric patterns. This provides a complete gradient basis for updating modulation parameters.

[0085] This embodiment can restore complete complex gradient information through complex field recovery processing, which makes up for the defect that light intensity detection cannot directly obtain the phase, and ensures that the parameter update direction and amplitude are accurate and reliable.

[0086] In an optional implementation, the central controller 15 is further configured to perform smoothing processing on the phase distribution corresponding to the complex gradient field in the complex field recovery result, so as to suppress high-frequency noise and induce the formation of a smooth phase distribution.

[0087] In this embodiment, after the complex field recovery is completed, the central controller 15 can also smooth the phase distribution corresponding to the complex gradient field to suppress high-frequency noise and induce a smooth phase distribution. For example, the central controller 15 can use Gaussian filtering and low-frequency constraint algorithms to suppress high-frequency abrupt changes, sharp jumps and fragmented structures, thereby inducing the phase distribution to form a smooth phase profile with strong robustness.

[0088] This embodiment reduces the system's sensitivity to micro / nano fabrication errors, optical path vibrations, and alignment deviations through phase smoothing, significantly improving the robustness of physical deployment.

[0089] In an optional implementation, the central controller 15 is further configured to simultaneously update the error convolution parameters of the layer to be updated while updating the forward modulation parameters of the layer to be updated.

[0090] In this embodiment, the central controller 15 uses weight mirroring to update the forward modulation parameters of the layer to be updated. Simultaneously, the error convolution parameters of the layer to be updated are updated. The specific weighted mirroring rules are as follows: ,and The learning rate η Set to 0.007, error kernel synchronization update rate γ It can be set to 0.1.

[0091] This embodiment enhances the consistency of local target propagation by updating forward and error parameters synchronously, thereby accelerating training convergence and improving the learning stability of deep networks.

[0092] In an optional implementation, the central controller 15 is also used to employ an asynchronous parallel pipeline scheduling method, so that the m-th layer and the m+1-th layer simultaneously perform local error calculation, gradient acquisition and parameter refresh, where m is an integer greater than or equal to 1.

[0093] In this embodiment, the central controller 15 adopts an asynchronous parallel pipeline scheduling strategy to control the m-th layer and the m+1-th layer to perform local error calculation, gradient acquisition and parameter refresh simultaneously, so that the multi-layer operations are executed overlappingly in the time dimension, eliminating global time delay.

[0094] In this embodiment, the central controller 15 determines the virtual target field of the layer to be updated based on the local error information of the next layer and the error convolution component of the layer to be updated. The complex amplitude encoding unit 11 generates the target light field, and the local error physical calculation unit 13 performs interference processing on the actual output light field and the target light field to obtain the local error field. The parallel gradient interference acquisition unit 14 acquires gradient information based on the interference pattern corresponding to the local error field and updates the forward modulation parameters of the layer to be updated accordingly. This enables the formation of a local training closed loop around the layer to be updated, from target reference construction, local error acquisition, gradient information acquisition to parameter update. Based on this local training closed loop, the dependence of the training process on the backpropagation of global error along the complete forward propagation path can be reduced, and it is beneficial to combine the parameter update process with the actual optical training object, thereby improving the adaptability of in-situ training of optical neural networks and the stability of training control.

[0095] Figure 2 A flowchart illustrating an optical neural network training method according to Embodiment 2 of this application is shown schematically.

[0096] like Figure 2 As shown, this optical neural network training method is applied to the optical neural network training system in Embodiment 1. The method may include steps S200~S210, wherein: Step S200: Forward modulation of the training input light field to output the actual output light field corresponding to the layer to be updated; Step S202: Based on the local error information of the next layer and the error convolution component of the next layer, determine the error processing information, and determine the virtual target field of the layer to be updated according to the actual output light field and the error processing information; Step S204: Generate a target light field based on the virtual target field; Step S206: Perform interference processing on the actual output light field and the target light field to obtain the local error field corresponding to the layer to be updated; Step S208: Based on the interferogram corresponding to the local error field, extract the gradient information corresponding to the layer to be updated; Step S210: Update the forward modulation parameters of the layer to be updated according to the gradient information.

[0097] The optical neural network training method provided in this embodiment first introduces the training input light field into the physical modulation layer for forward phase modulation, extracts the light field features, and outputs the actual output light field of the layer to be updated. Then, the central controller collects the local error information of the next layer and calculates the virtual target field of the current layer by combining it with the error convolution component of the layer to be updated. The complex amplitude encoding unit modulates and generates the corresponding target light field based on the virtual target field. The actual output light field and the target light field are then introduced into the interference optical path for coherent processing to obtain the local error field. The parallel gradient interference acquisition unit performs interference sampling on the local error field and the modulation phase term to obtain gradient information. Finally, the central controller updates the forward modulation parameters of the layer to be updated based on the gradient information. This training method replaces global backpropagation with local target propagation, achieving in-situ training across the entire optical domain, thereby reducing the requirements for optical reciprocity and time reversal symmetry.

[0098] In an optional implementation, see [link to relevant documentation]. Figure 3 The step of determining error processing information based on the local error information of the next layer and the error convolution component of the next layer may include: Step S300: Obtain the local error information of the next layer; Step S302: Apply correction coefficients to the error convolution component of the next layer, and perform a Fourier transform on the error convolution component after applying the correction coefficients.

[0099] Step S304: Perform calculations between the transformed result and the local error information of the next layer to obtain error processing information.

[0100] In this embodiment, the local error information of the next layer detection plane is first obtained through the detector. Then, the error convolution component of the next layer... Apply correction factor After applying a correction factor The error convolution component performs a Fourier transform, i.e. Finally, the transformed result is processed with the local error information of the next layer to obtain the error processing information, and the error processing information = × .

[0101] The local error information refers to the digital data obtained after detecting the local error field.

[0102] In an optional implementation, see [link to relevant documentation]. Figure 4 The step of generating a target light field based on the virtual target field includes: Step S400: Control the spatial light modulator to encode the virtual target field; Step S402: Output the encoded target light field.

[0103] In this embodiment, when generating the target light field based on the virtual target field, the central controller converts the virtual target field into a phase-coded signal, drives the spatial light modulator to perform real-time phase-coded modulation on the incident coherent beam, and outputs a target light field that matches the virtual target field.

[0104] This embodiment uses a spatial light modulator to convert digital targets into optical light fields, simplifying the generation process and improving modulation response speed.

[0105] In an optional implementation, see [link to relevant documentation]. Figure 5 Generating a target light field based on the virtual target field includes: Step S500: The virtual target field is decomposed into two phase components using dual-phase holographic coding technology; Step S502: The decomposed light field is reconstructed by performing complex amplitude reconstruction using a low-pass filter to obtain the target light field.

[0106] In this embodiment, the complex-valued virtual target field is decomposed into two orthogonal phase components using dual-phase holographic coding technology to avoid complex amplitude modulation loss. The decomposed light field is filtered by a low-pass filter in the 4f system to remove high-frequency stray components and complete complex amplitude reconstruction to obtain a standard target light field.

[0107] This embodiment combines dual-phase encoding with low-pass reconstruction to accurately generate a high-fidelity complex amplitude target light field, reducing error calculation interference.

[0108] In an optional implementation, see [link to relevant documentation]. Figure 6 The step of interferometric processing of the actual output light field and the target light field to obtain the local error field corresponding to the layer to be updated includes: Step S600: Connect the actual output light field and the target light field to the two input terminals of the local error physical calculation unit, respectively; Step S602: Control the phase shifter to introduce a π phase shift into one of the optical fields participating in the interference; Step S604: Cause the two coherent beams after the phase shift to undergo destructive interference; Step S606: Obtain the local error field after interference.

[0109] In this embodiment, the actual output light field and the target light field are respectively connected to the two input ports of the local error physical calculation unit. The central controller drives the phase shifter to introduce a π phase shift to one of the light fields, so that the two coherent lights form a destructive interference condition at the beam splitter. The detector directly collects the light field signal after interference as the local error field.

[0110] This embodiment directly realizes complex difference operations through optical domain destructive interference, without the need for digital circuit processing, preserving phase information and improving error acquisition efficiency.

[0111] In an optional implementation, see [link to relevant documentation]. Figure 7 The step of extracting gradient information corresponding to the layer to be updated based on the interferogram corresponding to the local error field includes: Step S700: Spatial interference is performed between the local error field and the modulation phase term using phase-shifting interferometry. Step S702: Record multiple frames of original interferometric patterns under multiple phase shift conditions; Step S704: Demodulate the multiple frames of original interferometric patterns to extract the gradient information.

[0112] In this embodiment, phase-shifting interferometry is used to make the local error field and the modulation phase term of the layer to be updated superimpose and interfere in space. Multiple sets of phase shifts are applied sequentially and multiple frames of original interferograms are recorded. The multiple frames of original interferograms are demodulated by phase demodulation and complex amplitude reconstruction algorithms to extract complete gradient information containing amplitude and phase.

[0113] This embodiment can simultaneously acquire the magnitude and phase information of the gradient through multi-phase-shift interferometric sampling, solving the problem that complex gradients cannot be extracted from single-frame intensity maps and improving the accuracy of gradient calculation.

[0114] In an optional implementation, the method further includes: The original interferograms of the multiple frames are subjected to complex field recovery processing to obtain a complex field recovery result, which includes a complex gradient field.

[0115] In this embodiment, the central controller can process the pattern using phase demodulation and complex amplitude reconstruction algorithms, perform a complex field recovery operation, and reconstruct the complex gradient field containing the real and imaginary parts from the interferogram. This provides a complete gradient basis for updating modulation parameters.

[0116] This embodiment can restore complete complex gradient information through complex field recovery processing, which makes up for the defect that light intensity detection cannot directly obtain the phase, and ensures that the parameter update direction and amplitude are accurate and reliable.

[0117] In an optional implementation, the method further includes: The complex field recovery result is smoothed to suppress high-frequency noise and induce a smooth phase distribution.

[0118] This embodiment has natural phase smoothing properties, which can reduce the system's sensitivity to micro-nano fabrication errors, optical path vibrations and alignment deviations, and significantly improve the robustness of physical deployment.

[0119] In an optional implementation, updating the forward modulation parameters of the layer to be updated based on the gradient information includes: Update the forward modulation parameters of the layer to be updated based on the gradient information; The error convolution parameters of the layer to be updated are updated synchronously.

[0120] In this embodiment, the central controller can use weight mirroring rules to update the forward modulation parameters of the layer to be updated. Simultaneously, the error convolution parameters of the layer to be updated are updated. The specific weighted mirroring rules are as follows: ,and The learning rate η Set to 0.007, error kernel synchronization update rate γ It can be set to 0.1.

[0121] This embodiment enhances the consistency of local target propagation by updating forward and error parameters synchronously, thereby accelerating training convergence and improving the learning stability of deep networks.

[0122] In an optional implementation, the method further includes: When performing gradient updates at layer m+1, gradient acquisition and parameter refresh are simultaneously triggered at layer m, where m is an integer greater than or equal to 1.

[0123] In this embodiment, the central controller can adopt an asynchronous parallel pipeline scheduling strategy to control the gradient update of the (m+1)th layer and simultaneously trigger the gradient acquisition and parameter refresh of the mth layer, so that the multi-layer operations are executed overlappingly in the time dimension, eliminating global time delay.

[0124] In an optional implementation, the forward modulation of the training input light field to output the actual output light field corresponding to the layer to be updated includes: The training input light field is spatially phase modulated by multiple cascaded diffraction modulation sublayers. The actual output light field is obtained on the detection plane corresponding to the layer to be updated.

[0125] In this embodiment, the training input light field is incident on multiple cascaded diffraction modulation sub-layers. Each sub-layer sequentially performs spatial phase modulation and feature extraction on the light field. When the light field propagates to the detection plane of the layer to be updated, the high-resolution detector collects the light field signal as the actual output light field.

[0126] This embodiment enhances feature representation capabilities through a multi-level diffraction structure.

[0127] In an optional implementation, see [link to relevant documentation]. Figure 8 Before performing forward modulation on the training input light field, the method further includes: Step S800: Initialize the physical modulation layer to a random phase distribution or a plane wave state; Step S802: Initialize the forward modulation parameters to zero; Step S804: Initialize the error convolution parameters to uniformly distributed random noise.

[0128] In this embodiment, before executing the forward modulation and training process, the system is standardized and initialized. The phase distribution of the physical modulation layer is initialized to a random phase or plane wave state, the forward modulation parameters are set to zero initial values, and the error convolution parameters are initialized to uniformly distributed random noise of 0~2π.

[0129] This embodiment avoids training divergence caused by abnormal parameters through initialization processing, ensuring stable start-up and rapid convergence of the training process.

[0130] To facilitate understanding of this application, the training system of this application is verified below through the MNIST handwritten digit recognition experiment.

[0131] The experimental conditions were as follows: a 532nm laser was used, the SLM resolution was 1920×1080, the pixel size was 8.0μm, and the image preprocessing was 112×112 pixels.

[0132] The experimental procedure was as follows: the system was trained in situ for 5 epochs on the MNIST dataset (4 classes from 0 to 3).

[0133] The experimental results are as follows: After the above steps, the system achieves an accuracy of 99.5% in simulation and 95% in physical experiments based on programmable SLM.

[0134] The observation results of this experiment are as follows: Figure 9 As shown, through Figure 9 As can be seen, the modulation phase profile finally learned in this application exhibits smooth and continuous macroscopic structural features, rather than fragmented high-frequency noise, which verifies the effectiveness of this implementation in improving the robustness of the physical system.

[0135] It should be noted that the above are merely preferred embodiments of this application and do not limit the scope of patent protection of this application. Any equivalent structural or procedural changes made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of this application.

Claims

1. An optical neural network training system, characterized in that, include: A coherent light source unit is used to generate coherent light beams; Complex amplitude encoding unit, used to generate target light field based on virtual target field; A physical modulation layer is used to forward modulate the training input light field to output the actual output light field corresponding to the layer to be updated. The physical modulation layer includes at least a forward modulation component and an error convolution component. The local error physical calculation unit is used to perform interference processing on the actual output light field and the target light field to obtain the local error field corresponding to the layer to be updated. A parallel gradient interferometry acquisition unit is used to interfere the local error field with the modulation phase term that characterizes the modulation state of the layer to be updated, and to extract the gradient information corresponding to the layer to be updated based on the interferogram corresponding to the local error field. The central controller is used to determine error processing information based on the local error information of the next layer and the error convolution unit of the next layer, determine the virtual target field of the layer to be updated according to the actual output light field and the error processing information, and update the forward modulation parameters of the layer to be updated according to the gradient information.

2. The optical neural network training system according to claim 1, characterized in that, The physical modulation layer includes multiple cascaded diffraction modulation sub-layers, each of which is provided with the forward modulation component and the error convolution component.

3. The optical neural network training system according to claim 1, characterized in that, The complex amplitude encoding unit includes a spatial light modulator, and the central controller is used to control the spatial light modulator to encode the virtual target field in order to output the target light field.

4. The optical neural network training system according to claim 3, characterized in that, The complex amplitude encoding unit includes a dual-phase holographic encoding spatial light modulator and a low-pass filter component. The low-pass filter component is used to reconstruct the complex amplitude of the encoded light field.

5. The optical neural network training system according to claim 1, characterized in that, The local error physical calculation unit includes a beam splitter, a phase shifter, and a detector. The phase shifter is used to introduce a π phase shift into one of the optical fields involved in the interference, so that the actual output optical field and the target optical field will undergo destructive interference.

6. The optical neural network training system according to claim 1, characterized in that, The parallel gradient interferometry acquisition unit includes a phase shifter and a detector. The parallel gradient interferometry acquisition unit is used to make the local error field interfere with the modulation phase term under multiple phase shift conditions, and record multiple frames of original interferograms under multiple phase shift conditions. The central controller is used to extract the gradient information based on the multiple frames of original interferograms.

7. The optical neural network training system according to claim 6, characterized in that, The central controller is also used to perform complex field recovery processing on the multi-frame original interferogram to obtain a complex field recovery result, which includes a complex gradient field.

8. The optical neural network training system according to claim 7, characterized in that, The central controller is also used to perform smoothing processing on the phase distribution corresponding to the complex gradient field in the complex field recovery result, so as to suppress high-frequency noise and induce the formation of a smooth phase distribution.

9. The optical neural network training system according to claim 1, characterized in that, The central controller is also used to simultaneously update the error convolution parameters of the layer to be updated while updating the forward modulation parameters of the layer to be updated.

10. The optical neural network training system according to claim 9, characterized in that, The central controller is also used to employ an asynchronous parallel pipeline scheduling method, so that the m-th layer and the (m+1)-th layer simultaneously perform local error calculation, gradient acquisition, and parameter refresh, where m is an integer greater than or equal to 1.

11. The optical neural network training system according to claim 1, characterized in that, The physical modulation layer includes at least one of a programmable spatial light modulator and a three-dimensional diffraction chip formed by two-photon polymerization.

12. A method for training an optical neural network, characterized in that, The method, applied to an optical neural network training system, includes: The training input light field is forward modulated to output the actual output light field corresponding to the layer to be updated; Error processing information is determined based on the local error information of the next layer and the error convolution component of the next layer, and the virtual target field of the layer to be updated is determined based on the actual output light field and the error processing information. Generate a target light field based on the virtual target field; The actual output light field and the target light field are subjected to interference processing to obtain the local error field corresponding to the layer to be updated; Based on the interferogram corresponding to the local error field, extract the gradient information corresponding to the layer to be updated. The forward modulation parameters of the layer to be updated are updated based on the gradient information.

13. The optical neural network training method according to claim 12, characterized in that, The determination of error processing information based on the local error information of the next layer and the error convolution component of the next layer includes: Obtain the local error information of the next layer; A correction coefficient is applied to the error convolution component of the next layer, and a Fourier transform is performed on the error convolution component after the correction coefficient is applied; The transformed result is then processed with the local error information of the next layer to obtain error processing information.

14. The optical neural network training method according to claim 12, characterized in that, The step of generating a target light field based on the virtual target field includes: The spatial light modulator is controlled to encode the virtual target field; Output the encoded target light field.

15. The optical neural network training method according to claim 14, characterized in that, The step of generating a target light field based on the virtual target field includes: The virtual target field is decomposed into two phase components using dual-phase holographic coding technology; The decomposed light field is reconstructed by using a low-pass filter to obtain the target light field.

16. The optical neural network training method according to claim 12, characterized in that, The step of interferometric processing of the actual output light field and the target light field to obtain the local error field corresponding to the layer to be updated includes: The actual output light field and the target light field are respectively connected to the two input terminals of the local error physical calculation unit; The phase shifter is controlled to introduce a π phase shift into one of the optical fields involved in the interference; This causes the two coherent beams, after being phase-shifted, to undergo destructive interference. Obtain the local error field after interference.

17. The optical neural network training method according to claim 12, characterized in that, The step of extracting gradient information corresponding to the layer to be updated based on the interferogram corresponding to the local error field includes: The local error field is spatially interfered with the modulation phase term using phase-shifting interferometry. Record multiple frames of original interferograms under various phase shift conditions; The original interferometric patterns of the multiple frames are demodulated to extract the gradient information.

18. The optical neural network training method according to claim 17, characterized in that, The method further includes: The original interferograms of the multiple frames are subjected to complex field recovery processing to obtain a complex field recovery result, which includes a complex gradient field.

19. The optical neural network training method according to claim 18, characterized in that, The method further includes: The complex field recovery result is smoothed to suppress high-frequency noise and induce a smooth phase distribution.

20. The optical neural network training method according to claim 12, characterized in that, Updating the forward modulation parameters of the layer to be updated based on the gradient information includes: Update the forward modulation parameters of the layer to be updated based on the gradient information; The error convolution parameters of the layer to be updated are updated synchronously.

21. The optical neural network training method according to claim 20, characterized in that, The method further includes: When performing gradient updates at layer m+1, gradient acquisition and parameter refresh are simultaneously triggered at layer m, where m is an integer greater than or equal to 1.

22. The optical neural network training method according to claim 12, characterized in that, The forward modulation of the training input light field to output the actual output light field corresponding to the layer to be updated includes: The training input light field is spatially phase modulated by multiple cascaded diffraction modulation sublayers. The actual output light field is obtained on the detection plane corresponding to the layer to be updated.

23. The optical neural network training method according to claim 12, characterized in that, Prior to the forward modulation of the training input optical field, the method further includes: The physical modulation layer is initialized to a random phase distribution or a plane wave state; Initialize the forward modulation parameters to zero; The error convolution parameters are initialized to uniformly distributed random noise.