Multitask continuous learning method based on optical diffraction deep neural network
By applying elastic weight holding method and error backpropagation training algorithm in optical diffraction deep neural networks, a modulated phase mask is generated, which solves the catastrophic forgetting problem in optical neural networks, and realizes continuous learning and efficient recognition of multitasking.
Patent Information
- Application Number
- CN202510122443.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-06-03
AI Technical Summary
Existing optical diffraction deep neural networks have catastrophic forgetting problems and cannot achieve continuous learning for multitasking without adding or changing the network structure.
Using a training algorithm based on elastic weight holding method and error backpropagation, a modulated phase mask is generated for hidden layer weights of optical diffraction deep neural networks to achieve continuous learning for multi-tasks.
Without increasing the network structure, the catastrophic forgetting problem in optical neural networks is effectively solved, the continuous learning of multi-tasks is realized, and the recognition accuracy and efficiency are improved.
Smart Images

Figure CN120087402A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of optoelectronic machine learning, and more particularly to a multi-task continuous learning method based on an optical diffraction deep neural network. Background Art
[0002] Artificial intelligence, as the most popular concept in information science today, has important applications in fields such as machine vision, autonomous driving, target tracking, and natural language processing. Among them, the artificial neural network (Artificial Neural Network, ANN) has been a research hotspot in the field of artificial intelligence since the 1980s. It abstracts the human brain neuron network from the perspective of information processing, establishes a simple model, and forms different networks according to different connection methods. As an interdisciplinary product of optoelectronic technology and artificial intelligence technology, the optical neural network can combine the advantages of both to construct a high-speed and low-power network structure, breaking through the bottleneck of traditional electronic neural networks. The optical neural network (Optical neural network, ONN) can effectively reduce the partial operations of both software and electronic hardware, and has characteristics such as high bandwidth, high interconnectivity, and inherent parallel processing, providing a promising method for replacing artificial neural networks. Since the optical diffraction deep neural network (Diffractive Deep Neural Networks, D2NN) implemented using 3D printing technology was proposed by Lin et al. in 2018, which is based on the diffraction principle of optics for network connection, the recognition accuracy of the 5-layer optical diffraction deep neural network for the MNIST dataset has reached as high as 97%. Catastrophic forgetting is an inevitable problem that occurs when artificial neural networks switch datasets. It refers to the fact that when learning a new task, the recognition accuracy of the original task will decrease significantly as the recognition accuracy of the new task increases. Because learning a new task will overwrite the weights learned in the past, thereby reducing the model performance of the past tasks. If this problem is not solved, a single neural network will not be able to adapt to continuous learning scenarios. Common solutions to catastrophic forgetting include regularization methods and replay methods. The replay method usually achieves this by maintaining a memory bank or experience pool during the training process. This memory bank stores the data or experience of past tasks, and then when training a new task, some old data or experience is randomly retrieved from the memory bank for re-learning, thereby helping the model maintain the memory of the old tasks. The disadvantage of the replay method is that it requires additional computational resources and storage space for recalling old knowledge. The representative of the regularization method is to impose a form of regularization on the loss function of the model when learning a new task, thereby retaining those parameters that are important for previous tasks. The representative of the regularization method is the elastic weight consolidation method. The above methods for solving the problem of catastrophic forgetting are only used in the scenario of electronic neural networks.
[0003] There are already many related literatures on optical diffraction deep neural networks, and there are also problems of catastrophic forgetting in optical diffraction deep neural networks. However, there is currently no optical method to solve the problem of catastrophic forgetting in optical diffraction deep neural networks. Therefore, how to solve the problem of catastrophic forgetting in optical neural networks with optical methods is a technical problem that needs to be solved urgently by those skilled in the art.
[0004] Lin X et al. "All-optical machine learning using diffractive deep neural networks", Ozcan A. Science. 2018 Sep 7; 361(6406): 1004-1008 proposed the optical diffraction deep network framework D2NN (Diffractive Deep Neural Network). The entire network model is based on the forward propagation of the optical field in free space. The connections of each layer of the network are through the diffraction formula. Each intermediate layer is composed of a diffraction plate printed by 3D printing technology. The weights of each intermediate layer can be phase or amplitude or a combination of the two. The training algorithm of the weights uses the Adam optimizer of deep learning. Once the network is trained, it can only perform a single task and cannot perform multiple tasks. Therefore, it also has the problem of catastrophic forgetting.
[0005] Kirkpatrick J et al. "Overcoming catastrophic forgetting in neural networks", national academy sciences 114, 3521–3526 (2017) proposed an elastic weight consolidation method (Elastic Weight Consolidation, EWC). This method solves the problem of catastrophic forgetting in the field of electronic neural networks. A regularization term is added to the original loss function of the electronic neural network, and the importance of the parameters is calculated using the Fisher information matrix, so that the parameters important for the previous task can be retained. The network verified the solution effect of the elastic weight consolidation method on the problem of catastrophic forgetting on the permuted MNIST dataset, proving that this method can protect the information of the previous task while learning new tasks.
[0006] "Experience replay for continual learning" by David Rolnick et al., Advances in Neural Information Processing Systems, vol. 32, 2019 proposes a method of experience replay. The network consists of a learning network and several behavior networks. The new experiences and replay experiences generated by the behavior networks are input into the learning network. The weights of the behavior networks are updated asynchronously. Online policy learning is used to process new experiences so as to quickly adapt to new tasks. At the same time, off-policy learning is used to combine behavior cloning of replay experiences to maintain and moderately improve the performance of past tasks. This method shows good performance and unique advantages in the catastrophic forgetting problem of multi-task reinforcement learning. Using the method of experience replay will additionally increase or change the structure of the electronic neural network, or require additional data or memory.
[0007] Existing optical diffraction neural networks have the problem of catastrophic forgetting. Although the method of experience replay shows good performance and unique advantages in the catastrophic forgetting problem of multi-task reinforcement learning, it will additionally increase or change the structure of the electronic neural network. There is no solution to the catastrophic forgetting problem of optical diffraction deep neural networks using optical methods.
[0008] How to efficiently and accurately achieve multi-task recognition without additionally increasing or changing the structure of the electronic neural network has become a technical problem to be solved. Summary of the Invention
[0009] The purpose of the present invention is to overcome the defects of the above-mentioned existing technologies and provide a multi-task continuous learning method based on an optical diffraction deep neural network.
[0010] The purpose of the present invention can be achieved by the following technical solutions:
[0011] According to one aspect of the present invention, there is provided a multi-task continuous learning method based on an optical diffraction deep neural network. The optical diffraction deep neural network includes an input layer, a hidden layer and an output layer. The hidden layer includes at least 1 phase modulation device. The weights of the hidden layer of the optical diffraction deep neural network are modulation phase masks loaded on the phase modulation devices. The modulation phase masks are generated by a training algorithm with elastic weight consolidation method and error backpropagation;
[0012] The method includes:
[0013] Step S1, initialize network parameters and modulation phases, and obtain n datasets to be recognized, where n is greater than 1;
[0014] Step S2: The input layer of the optical diffraction deep neural network sequentially loads each data set. The optical field of the l-th data set reaches the phase modulation device in the hidden layer through diffraction and is phase-modulated. The output layer detects the optical field intensity, and the output label of the l-th data set is obtained according to the brightness and darkness of the detection area, where 1 ≤ l ≤ n;
[0015] Step S3: Calculate the loss function of the l-th data set, calculate the phase gradient, and update the modulation phase mask matrix. Identify all data sets to obtain the recognition accuracy;
[0016] Step S4: Increase epoch by 1, and repeat Step S2 and Step S3 until epoch reaches the set number of iterations for the l-th data set, and calculate the Fisher information matrix;
[0017] Step S5: Adjust the parameters of the diffraction deep neural network, and repeat Steps S2 - S4 until the recognition accuracy of all data sets is higher than the set threshold.
[0018] Preferably, the input layer includes a laser, a collimator, and a phase modulation device;
[0019] The divergent light emitted by the laser becomes parallel light output after passing through the collimator;
[0020] The phase modulation device loads the data set onto the phase modulation device.
[0021] Preferably, the output layer includes a photodetector; the photodetector is used to capture the power image generated by the optical field diffracted by the last phase modulation device in the hidden layer. Different regions on the photodetector represent different labels, and the maximum brightness is used as the basis for judging the recognition result to obtain the output label of the optical diffraction deep neural network.
[0022] More preferably, in Step S2, for the i-th data set, after the light emitted by the laser passes through the collimator, the i-th data set is loaded onto the optical signal through the phase modulation device in the input layer; a random modulation phase mask is loaded onto the phase modulation device; the optical field of the i-th data set reaches the phase modulation device in the hidden layer through diffraction and is phase-modulated; the photodetector outputs the optical field intensity, and the output label is obtained according to the brightness and darkness of the detection area.
[0023] Preferably, the output of each phase modulation device in the hidden layer except the last layer is:
[0024] U k =M k W k U k-1
[0025]
[0026] where Uk is the vectorized output optical field of the phase modulation device for the k-th layer network, W k represents the diffraction weight matrix for the forward light propagating from the (k - 1)-th layer to the k-th layer, M k represents the vectorization of the phase mask of the phase modulation device in the k-th layer, and the phase weight of the hidden layer is j is the imaginary unit, and e is the base of the natural logarithm;
[0027]
[0028] Preferably, when training on the first data set, the loss function L of the optical diffraction deep neural network is the loss function generated by the forward propagation of the first data set, specifically:
[0029]
[0030] where T is the target field, i.e., the distribution of the labels; O is the power distribution detected by the output layer.
[0031] More preferably, when training on other data sets except the first data set, the loss function of the optical diffraction deep neural network is the loss function L plus the regularization term generated by the elastic weight consolidation method, specifically:
[0032]
[0033] where L EWC is the loss function of the elastic weight consolidation method, λ t is the hyperparameter of the elastic weight consolidation, F ti is the Fisher information matrix of the i-th phase weight in the t-th data set, i ∈ [1, M], M is the total number of phase weights, and L l is the loss function obtained by the forward propagation of the l-th data set, and are the i-th phase weights of the l-th data set and the t-th data set respectively.
[0034] Preferably, error backpropagation uses the Adam optimizer to update the modulation phase mask matrix.
[0035] Preferably, the adjustment of the parameters of the diffraction deep neural network includes adjusting the network size, the total number of iterations, the batch size, and the hyperparameter of the elastic weight consolidation.
[0036] Preferably, the data set includes optical signals with labels, where the optical signals include pictures and time-domain optical pulses.
[0037] Compared with the prior art, the present invention has the following beneficial effects:
[0038] 1) Without adding or changing the structure of the optical neural network device additionally, the present invention applies the training algorithm with elastic weight preservation to the optical diffraction deep neural network to solve the problem of catastrophic forgetting in the optical neural network, realizes continuous learning of multiple tasks, and improves the efficiency and accuracy of a single optical diffraction deep neural network in multi-task recognition.
[0039] 2) The optical diffraction deep neural network of the present invention has a simple structure, flexible tunability, good scalability, faster recognition speed and lower power consumption than the electronic neural network, and can be applied to optical computing in the future, playing the role of replacing electronic computers, which determines its huge potential market value. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 is a schematic flow chart of the multi-task continuous learning method in the present invention;
[0041] Figure 2 is a schematic structural diagram of the optical diffraction deep neural network in the present invention;
[0042] Figure 3 is a schematic diagram of the recognition rates of the existing optical diffraction deep neural network for the MNIST dataset and the Fashion MNIST dataset;
[0043] Figure 4 is a schematic diagram of the recognition rates of the optical diffraction deep neural network with elastic weight preservation in the present invention for the MNIST dataset and the Fashion MNIST dataset. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0045] This embodiment relates to a multi-task continuous learning method based on an optical diffraction deep neural network, innovatively uses the elastic weight preservation method on the optical neural network, solves the problem of catastrophic forgetting of the optical diffraction deep neural network without increasing the structure of the optical diffraction deep neural network, can realize continuous learning of multiple tasks, has a simple device, good scalability, flexible tunability, and has the advantages of faster recognition speed and lower power consumption than the electronic neural network, and has broad application prospects in the fields of optoelectronic technology and artificial intelligence.
[0046] This method is implemented based on an optical diffraction deep neural network with elastic weight preservation, as Figure 2, the optical diffraction deep neural network includes an input layer, a hidden layer, and an output layer in sequence.
[0047] The input layer includes a laser, a collimator, and a phase modulation device. The divergent light emitted by the laser becomes parallel light output after passing through the collimator; the phase modulation device loads the pictures of the data set onto the phase modulation device.
[0048] The hidden layer includes at least one phase modulation device. The light field after the laser passes through the collimator is diffracted through free space and incident on the phase modulation device. The phase mask generated by the training algorithm with elastic weight preservation method and error backpropagation is loaded onto the phase modulation device, and the modulation phase mask loaded on the phase modulation device serves as the weight of the hidden layer of the optical diffraction deep neural network.
[0049] The output layer includes a photodetector. The photodetector is used to capture the power image generated by the diffracted light field from the last phase modulation device in the hidden layer. Different regions on the photodetector represent different labels, and the maximum brightness is used as the basis for judging the recognition result to obtain the output label of the optical diffraction deep neural network.
[0050] Forward propagation: The light emitted by the laser loads the data set and is diffracted through free space and irradiated onto the phase modulation device.
[0051] The output of each layer of the phase modulation device is:
[0052] U k =M k W k U k-1
[0053]
[0054] where U k is the vectorized output light field of the phase modulation device of the k-th layer network, W k represents the diffraction weight matrix for the forward light propagating from the (k - 1)-th layer to the k-th layer, M k represents the vectorization of the phase mask of the k-th layer phase modulation device, and the phase weight of the hidden layer is j is the imaginary number and e is the natural logarithm.
[0055]
[0056] The formula for the network output field obtained by the photodetector is:
[0057] O=|U N+1 | 2
[0058] where O is the power distribution detected by the photodetector.
[0059] The loss function L of the network is as follows:
[0060]
[0061] Where T is the target field, that is, the distribution of labels.
[0062] Loss function formula based on elastic weight retention:
[0063]
[0064] Where L 2 is the loss function obtained by forward propagation of the second dataset, λ is the hyperparameter of elastic weight retention, and F is the Fisher information matrix M is the total number of phase weights, n is the number of dataset samples, L 1 is the loss function obtained by forward propagation of the first dataset, is the phase weight of the second dataset, is the phase weight of the first dataset.
[0065] Backpropagation: Error backpropagation training algorithm.
[0066] Adam optimizer formula:
[0067]
[0068] Where η is the learning rate; m t+1 is the first-order moment estimate of the gradient, v t+1 is the second-order moment estimate of the gradient; β 1 , β 2 are the exponential decay rates for controlling m t , v t respectively; ε is a very small constant to prevent division by zero; is the phase gradient; t is the current iteration number; ⊙ represents the Hadamard product; m t is the first-order moment of the previous gradient; v t is the second-order moment of the previous gradient.
[0069] Adjust the parameters according to the distribution of the recognition rate curve, and select a set of appropriate network parameters, including network size, total number of iterations iter, batch size, η, β 1 , β 2 , λ and ε.
[0070] In this embodiment, the elastic weight retention method is applied to the optical diffraction deep neural network. For the case of two datasets, such as Figure 1 , the specific operation steps are as follows:
[0071] Step 1: Build an optical diffraction deep neural network based on a phase modulation device, determine the number of intermediate layers of the network, the sizes of the input layer, output layer, and intermediate layers, and set the network model parameters (η, β 1 , β 2 , λ, ε), the total number of training iterations iter, and the batch size batchsize; Initialize the network parameters, randomly generate the modulation phase, and set epoch = 1;
[0072] Step 2: Obtain two datasets to be recognized, the first dataset and the second dataset. Each dataset includes optical signals with labels, where the optical signals include pictures and time-domain optical pulses.
[0073] When the currently loaded dataset is the first dataset, perform forward propagation on the optical diffraction deep neural network; In the optical diffraction deep neural network, the light emitted by the laser passes through the collimator and then the first dataset is loaded onto the optical signal through the phase modulation device of the input layer; Load the modulation phase mask onto the phase modulation device; The optical field loaded with the first dataset reaches the phase modulation device of the hidden layer through diffraction and is phase-modulated. The optical field diffracts every time it passes through the phase modulation device of a hidden layer; Obtain the intensity of the optical field output by the network through the photodetector, and obtain the output label according to the brightness and darkness of the detection area.
[0074] Step 3: Calculate the loss function of the first dataset and calculate the phase gradient. Based on the phase gradient, use the Adam optimizer to update the phase to obtain the updated modulation phase mask matrix, and identify the first dataset and the second dataset to obtain the recognition accuracy.
[0075] epoch = epoch + 1, repeat Steps 2 and 3 until epoch = iter / 2, and calculate the Fisher information matrix.
[0076] Step 4: Switch the loaded dataset to the second dataset and perform forward propagation on the optical diffraction deep neural network; Calculate the loss function of the second dataset plus the regularization term of elastic weight consolidation, and calculate the phase gradient; Based on the phase gradient, use the Adam optimizer to update the phase to obtain the updated modulation phase mask matrix, and identify the first dataset and the second dataset to obtain the recognition accuracy.
[0077] epoch = epoch + 1, repeat Step 4 until epoch = iter.
[0078] Step 5: Adjust the parameters of the diffraction deep neural network, and repeat Steps 1 - 4 until the recognition accuracy curves of both datasets maintain a high level, that is, the recognition rate is higher than the set threshold.
[0079] This embodiment also relates to a multi-task continuous learning method based on an optical diffraction deep neural network. For the case of three or more datasets, assuming the number of datasets is n, where n >= 3, the specific operation steps are as follows:
[0080] Step 1: Build an optical diffraction deep neural network based on a phase modulation device, determine the number of intermediate layers of the network, the sizes of the input layer, output layer, and intermediate layers, and set the network model parameters (η, β 1 , β 2 , λ, ε), the total number of training iterations iter, and the batch size batchsize; Initialize the network parameters, randomly generate the modulation phase, and set epoch = 1;
[0081] Step 2: Obtain n datasets to be recognized, namely the first dataset, the second dataset,..., the nth dataset in sequence.
[0082] The currently loaded dataset is the first dataset, and the optical diffraction deep neural network performs forward propagation; In the optical diffraction deep neural network, after the light emitted by the laser passes through the collimator, the first dataset is loaded onto the optical signal through the phase modulation device of the input layer; Load the modulation phase mask onto the phase modulation device; The light field loaded with the first dataset undergoes diffraction and reaches the phase modulation device of the hidden layer, and is phase-modulated. The light field undergoes diffraction every time it passes through the phase modulation device of a hidden layer; Obtain the light field intensity output by the network through the photodetector, and obtain the output label based on the bright and dark conditions of the detection area.
[0083] Step 3: Calculate the loss function of the first dataset and calculate the phase gradient. Based on the phase gradient, use the Adam optimizer to update the phase to obtain the updated modulation phase mask matrix, and perform recognition on all datasets to obtain the recognition accuracy.
[0084] epoch = epoch + 1, repeat Steps 2 and 3 until Calculate the Fisher information matrix.
[0085] Step 4: Switch the loaded dataset to the second dataset, and the optical diffraction deep neural network performs forward propagation; Calculate the loss function of the second dataset plus the regularization term of elastic weight consolidation, and calculate the phase gradient;
[0086] Step 5: Based on the phase gradient, use the Adam optimizer to update the phase to obtain the updated modulation phase mask matrix, and perform recognition on all datasets to obtain the recognition accuracy.
[0087] epoch = epoch + 1, the loss function is only the loss function of the second dataset, calculate the phase gradient; Repeat Step 5 until Calculate the Fisher information matrix.
[0088] Step 6: Switch the loaded dataset to the third dataset, and perform forward propagation on the optical diffraction deep neural network; calculate the loss function of the third dataset plus the regularization term with elastic weight preservation, and calculate the phase gradient.
[0089] Step 7: Based on the phase gradient, use the Adam optimizer to update the phase, obtain the updated modulation phase mask matrix, identify all datasets, and obtain the recognition accuracy rate.
[0090] epoch = epoch + 1, the loss function is only the loss function of the third dataset, calculate the phase gradient; repeat Step 7 until Calculate the Fisher information matrix.
[0091] Step 8: Continue to load the remaining datasets one by one according to the methods in Steps 6 - 7 until all datasets are processed.
[0092] Step 9: Adjust the parameters of the diffraction deep neural network, and repeat Steps 1 - 8 until the recognition accuracy rate curves of all datasets maintain a high level, that is, the recognition rate is higher than the set threshold.
[0093] This embodiment also relates to a multi - task continuous learning method based on an optical diffraction deep neural network, which is verified with a specific example. The working wavelength of the laser is 671 nm; the phase modulation device selects a spatial light modulator, which is of the reflective type, with a dimension of 800 * 600, and the size of each pixel is dx (2 μm); the photodetector selects a complementary metal - oxide - semiconductor sensor, with a size of 1.41312 * 0.7452 (cm²); the input image size is 28 * 28, which is changed to 64 * 64 through sampling, the output layer size is 128 * 128, there are 2 hidden layers, and the size of each layer is 128 * 128; the gradient update uses the Adam optimizer, the learning rate η is 0.001, β 1 is 0.9, β 2 is 0.999, λ is 500, ε is 1e - 8, the number of training Epochs is 20, and the batch size is 200. The first dataset is the MNIST dataset, and the second dataset is the Fashion MNIST dataset.
[0094] First, the network is trained. The epoch is divided into the first 10 generations and the last 10 generations. The training data of the MNIST dataset is trained in the first 10 generations, and the training data of the Fashion MNIST dataset is trained in the last 10 generations. After determining the network parameters, the modulation phase of each spatial light modulator is trained by the elastic weight consolidation method and the error backpropagation algorithm, and the parameters of all network models are saved. For example, the optical signal is transmitted forward from left to right. The Gaussian light emitted by the laser is collimated and loaded with the optical signals of the MNIST dataset and the Fashion MNIST dataset respectively, and passes through the spatial light modulator loaded with the trained modulation phase in turn through the diffraction in free space. Finally, the diffracted light field is measured by a complementary metal oxide semiconductor sensor, and the recognition result of the network is obtained by observing the label corresponding to the region with the highest brightness. The test data of the MNIST dataset and the Fashion MNIST dataset are recognized, and the recognition accuracy is output. When training on the first dataset, the loss function is only the loss function generated by the forward propagation of the first dataset, and the error backpropagation uses the Adam optimizer to update the phase weights. When training on the second dataset, the loss function is the loss function generated by the forward propagation of the second dataset plus the regularization term generated by the elastic weight consolidation method, and the error backpropagation uses the Adam optimizer to update the phase weights.
[0095] Figure 3 It is the recognition rate of the test data of the MNIST dataset and the Fashion MNIST dataset obtained by simulating the existing optical diffraction deep neural network without the elastic weight consolidation method. Figure 4 It is the recognition rate of the test data of the MNIST dataset and the Fashion MNIST dataset obtained by simulating the optical diffraction deep neural network system with elastic weight consolidation in the present invention.
[0096] The optical diffraction deep neural network with elastic weight consolidation uses the training method with elastic weight consolidation method and error backpropagation. Without additional increasing the network structure and data, it solves the catastrophic forgetting problem of the optical diffraction deep neural network, realizes the continuous learning of the optical diffraction neural network, replaces the electronic neural network calculation in an optical way, can effectively reduce the operation cost and time of the network, and has the advantages of high speed and energy saving. The present invention has good application prospects, simple equipment, and reconfigurability, which determines its huge potential market value in the fields of photon computing and artificial intelligence.
[0097] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A multi-task continuous learning method based on an optical diffraction deep neural network, wherein the optical diffraction deep neural network comprises an input layer, a hidden layer and an output layer, wherein the hidden layer comprises at least one phase modulation device, characterized in that: The weight of the hidden layer of the optical diffraction deep neural network is a modulated phase mask loaded on the phase modulation device, wherein the modulated phase mask is generated by a training algorithm with an elastic weight retention method and error back propagation; The method includes: Step S1, initializing network parameters and modulation phase, obtaining n data sets to be identified, where n is greater than 1; Step S2, the input layer of the optical diffraction deep neural network loads each data set in turn, the light field of the lth data set is diffracted to reach the phase modulation device of the hidden layer and is phase modulated, the output layer detects the light field intensity, and obtains the output label of the lth data set according to the brightness of the detection area, 1≤l≤n; Step S3, calculating the loss function of the first data set, calculating the phase gradient, and updating the modulation phase mask matrix, identifying all data sets to obtain the recognition accuracy; Step S4, epoch increases by 1, and steps S2 and S3 are repeated until epoch reaches the number of iterations set for the lth data set, and the Fisher information matrix is calculated; Step S5, adjust the parameters of the diffraction deep neural network and repeat steps S2 to S4 until the recognition accuracy of all data sets is higher than the set threshold.
2. The multi-task continuous learning method based on optical diffraction deep neural network according to claim 1 is characterized in that: The input layer includes a laser, a collimator and a phase modulation device; The divergent light emitted by the laser is converted into parallel light output after passing through the collimator; The phase modulation device loads the data set onto the phase modulation device.
3. The multi-task continuous learning method based on optical diffraction deep neural network according to claim 1 is characterized in that: The output layer includes a photodetector; The photodetector is used to capture the power image generated by the diffracted light field from the last layer of phase modulation devices in the hidden layer. Different areas on the photodetector represent different labels. The maximum brightness is used as the basis for judging the recognition result to obtain the output label of the optical diffraction deep neural network.
4. The multi-task continuous learning method based on optical diffraction deep neural network according to claim 2 is characterized in that: In step S2, for the i-th data set, after the light emitted by the laser passes through the collimator, the i-th data set is loaded into the optical signal through the phase modulation device of the input layer; a random modulation phase mask is loaded into the phase modulation device; The light field loaded with the i-th data set reaches the hidden layer phase modulation device after diffraction and is phase modulated; The photodetector outputs the light field intensity, and the output label is obtained according to the brightness of the detection area.
5. The multi-task continuous learning method based on optical diffraction deep neural network according to claim 1, characterized in that: The output of each phase modulation device in the hidden layer except the last layer is: U k =M k W k U k-1 Among them, U k is the vectorized output light field of the phase modulation device of the k-th layer network, W k represents the diffraction weight matrix of the forward light propagating from the k-1th layer to the kth layer, M k represents the vectorization of the phase mask of the k-th phase modulation device, and the phase weight of the hidden layer is j is an imaginary number, e is a natural logarithm; The output of the last phase modulation device in the hidden layer is: Among them, U0 is the data set of the input layer.
6. The multi-task continuous learning method based on optical diffraction deep neural network according to claim 1, characterized in that: When training the first data set, the loss function L of the optical diffraction deep neural network is the loss function generated by the forward propagation of the first data set, specifically: Among them, T is the target field, that is, the distribution of labels; O is the power distribution detected by the output layer.
7. The multi-task continuous learning method based on optical diffraction deep neural network according to claim 6, characterized in that: When training other data sets except the first data set, the loss function of the optical diffraction deep neural network is the loss function L plus the regularization term generated by the elastic weight retention method, specifically: Among them, L EWC is the loss function of the elastic weight preservation method, λ t is the hyperparameter maintained by the elastic weight, F ti is the Fisher information matrix of the t-th data set at the i-th phase weight, i∈[1,M], M is the total number of phase weights, L l is the loss function obtained by forward propagation of the lth data set, and are the i-th phase weights of the l-th data set and the t-th data set respectively.
8. The multi-task continuous learning method based on optical diffraction deep neural network according to claim 1, characterized in that: Error back-propagation uses the Adam optimizer to update the modulation phase mask matrix.
9. The multi-task continuous learning method based on optical diffraction deep neural network according to claim 1, characterized in that: The described adjustment of the diffractive deep neural network parameters includes adjusting the network size, the total number of iterations, the batch size, and the hyperparameters of elastic weight retention.
10. The multi-task continuous learning method based on optical diffraction deep neural network according to claim 1, characterized in that: The data set includes optical signals with labels, wherein the optical signals include images and time-domain optical pulses.