Deep optical neural network training method and system based on hybrid mutation strategy genetic algorithm

By training deep optical neural networks through a hybrid mutation strategy genetic algorithm, the problems of slow search speed and easy falling into local optimality in traditional algorithms are solved, and more efficient DONN training and optimization are achieved.

CN116596056BActive Publication Date: 2025-10-24HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310591505.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-24
Publication Date
2025-10-24
Estimated Expiration
2043-05-24

AI Technical Summary

Technical Problem

Traditional genetic algorithms have a slow search speed in deep optical neural network training and are prone to premature convergence, resulting in suboptimal performance.

Method used

A hybrid mutation strategy genetic algorithm (MSGA) is adopted, which combines uniform initialization, exponential sort selection, uniform crossover and three mutation operators, and adopts a double elite retention strategy to find the optimal DONN individual through iterative evolution.

Benefits of technology

The robustness and generalization ability of DONN individuals are enhanced, local optimality is avoided, and training efficiency and performance are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116596056B_ABST
    Figure CN116596056B_ABST
Patent Text Reader

Abstract

The application discloses a deep optical neural network training method and system based on a hybrid mutation strategy genetic algorithm, and the method comprises the following steps: S1, sequentially stacking a linear operation layer based on MZIs, a nonlinear activation layer based on EOA and a Dropmask based on a mask to build an N-layer deep DONN; S2, preprocessing a data set with different characteristic categories to conform to the data input size of the DONN; S3, uniformly initializing the DONN population, combining the MSE and the Accuracy between the real value and the predicted value as the fitness evaluation function of the individual; S4, taking the exponential ranking selection ERS and the uniform crossover UC as the selection operator and the crossover operator in the training process, adopting a hybrid mutation strategy, and distributing three operators, namely, the single-point mutation SM, the uniform mutation UM and the Gaussian mutation GM, to different individuals for mutation according to a dynamic game probability; and S5, adopting a double-elite reservation strategy, reserving two individuals with the optimal MSE and Accuracy performance to the next generation, and through iterative evolution, until a termination condition is met, and a DONN individual with the globally optimal network parameter is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of photonics design and the field of artificial intelligence technology, and in particular to an effective training method and system of a deep optical neural network based on a hybrid mutation strategy genetic algorithm. BACKGROUND

[0002] With the rapid development of artificial intelligence (AI), artificial neural networks (ANN) have been widely used in many fields, and have excellent performance for image classification, natural language processing, unmanned driving, robots and other learning tasks. However, traditional electronic chips have been unable to meet the growing demand for large-scale data processing in an information-based society, reaching the "von Neumann bottleneck". Because optical signal processing has excellent characteristics such as high parallelism, low power consumption and low latency, researchers have focused on deep optical neural networks (DONN), which make it possible to solve this problem.

[0003] For DONN, the network training algorithm determines the performance of the DONN, so developing an effective network training algorithm is a current problem to be solved. The traditional genetic algorithm (GA) is a gradient-free evolutionary algorithm, which is an algorithm model established by researchers by simulating the biological genetic evolution mechanism in nature. GA regards organisms as individuals, the adaptability of individuals to the environment as fitness, a limited number of individuals form a population, and the genes of individuals in the population are selected, crossed and mutated to generate a new population. However, the traditional GA has the following problems: slow search speed, easy to converge prematurely, i.e. "premature", thus falling into local optimum, resulting in the performance of the DONN obtained ultimately being unable to reach the best. SUMMARY

[0004] In view of the low efficiency of the existing deep optical neural network training algorithm, the present application provides a deep optical neural network training method and system based on a hybrid mutation strategy genetic algorithm.

[0005] On the basis of the prior art, the application adopts a mixed mutation strategy genetic algorithm (MSGA) to uniformly initialize a DONN population, combines mean square error (MSE) between a real value and a predicted value and accuracy as an individual fitness evaluation function, adopts exponential ranking selection (ERS) and uniform crossover (UC) as selection and crossover operators in a training process, adopts three mutation operators, namely, single-point mutation (SM), uniform mutation (UM) and Gaussian mutation (GM), to mutate different individuals according to dynamic game probabilities, simultaneously adopts a double-elite reservation strategy to generate a next-generation network population, and continuously iterates evolution until an optimal DONN individual is found. The MSGA enhances the optimization ability of the original GA, so that the robustness and generalization ability of the finally obtained DONN individual are stronger, and the performance is better.

[0006] In order to achieve the above object, the application adopts the following technical solutions:

[0007] The deep optical neural network training method based on the mixed mutation strategy genetic algorithm comprises the following steps:

[0008] S1. sequentially stacking a linear operation layer based on a cascaded Mach-Zehnder interferometer (MZI), a nonlinear activation layer based on an electro-optic activator (EOA) and an output layer (Dropmask) based on a mask, so as to build an N-layer deep optical neural network (DONN);

[0009] S2. pre-processing a data set with different characteristic categories to conform to a data input size of the DONN;

[0010] S3. uniformly initializing a DONN population, and combining mean square error (MSE) between a real value and a predicted value and accuracy as an individual fitness evaluation function;

[0011] S4. adopting exponential ranking selection (ERS) and uniform crossover (UC) as selection and crossover operators in a training process, and adopting a mixed mutation strategy to distribute three mutation operators, namely, single-point mutation (SM), uniform mutation (UM) and Gaussian mutation (GM), to mutate different individuals according to dynamic game probabilities;

[0012] S5. adopting a double-elite reservation strategy to reserve two individuals with optimal MSE and accuracy performance to a next generation, and continuously iteratively evolving until a termination condition is met, so as to obtain a DONN individual with globally optimal network parameters.

[0013] Further, in step S1, the cascaded Mach-Zehnder interferometers (MZIs) can realize large-scale optical matrix operations, the electro-optic activator (EOA) can realize the function of the nonlinear activation function, and the Dropmask layer removes the redundant neurons to meet the network output. A MZI is composed of two 3dB directional couplers in front and back and two adjustable phase shifters. The inner phase shifter controls the output splitting ratio, and the outer phase shifter controls the differential output phase. According to the singular value decomposition principle (SVD), any real matrix can be decomposed into the product of two unitary matrices and a diagonal matrix, that is

[0014] R = IΣJ *

[0015] where R is a real matrix, I is an m x m unitary matrix, Σ is an m x n diagonal matrix, and J * is the complex conjugate of an n x n unitary matrix J. The unitary matrices I and J * can be realized by the couplers and phase shifters in the MZI, and the diagonal matrix Σ can be realized by the optical attenuator. By configuring the network of cascaded MZIs, large-scale optical matrix operations can be realized. The EOA converts a small part of the optical input power into voltage and performs amplitude-phase modulation on the remaining part of the original optical signal. Assuming that the optical input signal is c, the generated nonlinear photoelectric activation function is f(c), and the specific expression is:

[0016]

[0017] where α is the tapping power ratio of the photodetector, is the responsivity of the photodetector, G is the gain of the transimpedance amplifier, H is the transfer function, V b is the static bias voltage, V π is the voltage required for the phase change π of the modulator.

[0018] Further, in step S2, data preprocessing includes data cleaning, i.e., checking whether there are missing values in the data set; data enhancement, i.e., adding Gaussian noise to the data set to make the generalization ability of the neural network stronger; and data division, i.e., dividing the data set into a training set and a test set to meet the data input size of the DONN.

[0019] Further, in step S3, the DONN population is uniformly initialized, so that the parameters of each network individual are uniformly distributed in the solution space. In addition, the mean square error (MSE) between the true value and the predicted value and the classification accuracy (Accuracy) are combined as the individual fitness evaluation function, and the specific expression is:

[0020]

[0021] wherein, MSE represents the mean square error between the true value and the predicted value, n represents the number of samples in the test set, Y i and Y and Y represent the true value and the predicted value of the i-th sample respectively, Accuracy represents the classification accuracy of the DONN, and q represents a weighted average factor, usually taking 0.5.

[0022] Further, in step S4, exponential ranking selection (ERS) and uniform crossover (UC) are selected as selection operators and crossover operators of the training process, and a mixed mutation strategy is adopted, in which single-point mutation (SM), uniform mutation (UM) and Gaussian mutation (GM) are distributed to different individuals for mutation according to dynamic game probability. wherein, r represents the r-th iteration, and i represents the i-th network individual in the current iteration;

[0023] If the network individual generates offspring through the mutation operator Q, and the offspring is selected to enter the next iteration evolution, the probability distribution of each operator is updated according to the following formula:

[0024]

[0025] If the network individual generates offspring through the mutation operator Q, and the offspring is not selected to enter the next iteration evolution, the probability distribution of each operator is updated according to the following formula:

[0026]

[0027] wherein, Q represents the current mutation operator, Q' represents other mutation operators, η ∈ (0, 1) is a coefficient for controlling the probability distribution of the mixed mutation strategy, which is equal to 1 / 3 here, and the probability distribution satisfies β represents the number of mutation operators.

[0028] Further, in step S5, a double-elite reservation strategy is adopted to reserve the two individuals with the best performances of MSE and Accuracy to the next generation, so as to prevent the loss of the current optimal individual and avoid the algorithm from failing to find the global optimal solution.

[0029] Further, in step S5, the termination condition is that the network is iterated 2000 times, and finally a DONN individual with global optimal network parameters is obtained.

[0030] Further, the hybrid mutation strategy genetic algorithm is as follows: uniformly initializing a DONN population, the population containing 50 individuals, inputting a test set into each DONN, and performing fitness evaluation; performing exponential sorting selection according to the obtained fitness values to obtain a selected population; extracting the weight and hyperparameter of each individual in the population for binary coding, 20-bit binary coding can be used to reduce precision loss; a uniform crossover operator is used, and the crossover probability is 0.8; a hybrid mutation strategy is used, and the mutation probability is 0.1; a new generation of network parameters is obtained, and after decoding, a double-elite reservation strategy is used to copy to generate the next generation population, and iteration is continuously performed until a termination condition is met, and a DONN individual with global optimal network parameters is obtained.

[0031] The application further discloses a system based on the deep optical neural network training method, which comprises the following modules.

[0032] A deep optical neural network building module: sequentially stacking a linear operation layer based on a cascade Mach-Zehnder interferometer (MZI), a nonlinear activation layer based on an electro-optic activator (EOA), and a Dropmask output layer based on a mask, so as to build an N-layer deep optical neural network (DONN).

[0033] A preprocessing module: preprocessing data sets with different characteristic categories to meet the data input size of the DONN.

[0034] A DONN population initialization module: uniformly initializing a DONN population, combining the mean square error (MSE) and the classification accuracy (Accuracy) between the true value and the predicted value as the fitness evaluation function of the individual.

[0035] A hybrid mutation strategy module: using exponential sorting selection (ERS) and uniform crossover (UC) as selection operators and crossover operators in the training process, using a hybrid mutation strategy, and distributing three kinds of operators, single-point mutation (SM), uniform mutation (UM) and Gaussian mutation (GM), to different individuals for mutation according to dynamic game probability.

[0036] A double-elite reservation strategy module: using a double-elite reservation strategy, reserving two individuals with the best MSE and Accuracy performance to the next generation, and through iterative evolution, until a termination condition is met, a DONN individual with global optimal network parameters is obtained.

[0037] Compared with the prior art, the application has the following beneficial effects:

[0038] 1. The application adopts optical signals as the carrier of information transmission, and has the characteristics of high parallelism, anti-interference, low delay, low power consumption and the like. Therefore, compared with traditional artificial neural networks, the N-layer deep optical neural network (DONN) built by basic optoelectronic devices including MZIs and EOAs and the like is more efficient and low in consumption, and can be widely applied to learning tasks such as image classification, natural language processing, unmanned driving, robots and the like.

[0039] 2. The hybrid mutation strategy genetic algorithm MSGA has good global search capability, can quickly find all solutions in the solution space, is less likely to fall into local optimum compared with the traditional genetic algorithm, has probability randomness, enhances the optimization ability of the algorithm, makes the robustness and generalization ability of the final DONN individual stronger, and has better performance. BRIEF DESCRIPTION OF DRAWINGS

[0040] Figure 1 is a flowchart of the deep optical neural network training method based on the hybrid mutation strategy genetic algorithm provided by the preferred embodiment of the application;

[0041] Figure 2 is a network framework structure diagram of the deep optical neural network training method based on the hybrid mutation strategy genetic algorithm provided by the preferred embodiment of the application;

[0042] Figure 3 is a flowchart of the hybrid mutation strategy genetic algorithm provided by the preferred embodiment of the application;

[0043] Figure 4 is a result diagram of the DONN performing an Iris classification task provided by the preferred embodiment of the application;

[0044] Figure 5 is a result diagram of the DONN performing a Seeds classification task provided by the preferred embodiment of the application;

[0045] Figure 6 is a result diagram of the DONN performing a Wine classification task provided by the preferred embodiment of the application;

[0046] Figure 7 is a system block diagram of the deep optical neural network training system based on the hybrid mutation strategy genetic algorithm provided by the preferred embodiment of the application. DETAILED DESCRIPTION

[0047] Following, the advantages and effects of the present application can be easily understood by those skilled in the art from the description. The present application can also be implemented or applied by different specific embodiments, and the details in the description can be modified or changed based on different views and applications without departing from the spirit of the present application. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict.

[0048] The purpose of the present application is to provide an effective training method and system for deep optical neural networks based on hybrid mutation strategy genetic algorithm in view of the low efficiency of existing deep optical neural network training algorithms.

[0049] As shown in Figure 1 The present embodiment provides an effective training method for optical neural networks based on hybrid mutation strategy genetic algorithm, which includes the following steps:

[0050] S1. Stack the linear operation layer based on cascaded Mach-Zehnder interferometer (MZIs), the nonlinear activation layer based on electro-optic activator (EOA), and the output layer based on mask (Dropmask) in sequence to build an N-layer deep optical neural network (DONN).

[0051] Specifically, in step S1, the cascaded Mach-Zehnder interferometer (MZIs) can realize large-scale optical matrix operation, the electro-optic activator (EOA) can realize the function of nonlinear activation function, and the Dropmask layer removes redundant neurons to meet the network output. A MZI is composed of two 3dB directional couplers in front and back and two adjustable phase shifters. The inner phase shifter controls the output splitting ratio, and the outer phase shifter controls the differential output phase. According to the singular value decomposition principle (SVD), any real matrix can be decomposed into the product of two unitary matrices and a diagonal matrix, that is,

[0052] R=IΣJ *

[0053] Wherein, R is a real matrix, I is a m×m unitary matrix, Σ is a m×n diagonal matrix, and J * is the complex conjugate of n×n unitary matrix J. Among them, the unitary matrices I and J *The coupler and phase shifter in the MZI can be realized, and the diagonal matrix Sigma can be realized by optical attenuator. By configuring the network of cascaded MZIs, large-scale optical matrix operation can be realized. EOA converts a small part of the optical input power into voltage, and the remaining part of the original optical signal is amplitude-phase modulated. Assuming that the optical input signal is c, the generated nonlinear photoelectric activation function is f(c), and the specific expression is:

[0054]

[0055] Wherein, alpha is the tapping power ratio of the photodetector, is the responsivity of the photodetector, G is the gain of the transimpedance amplifier, H is the transfer function, V b is the static bias voltage, V π is the voltage required for the phase change pi of the modulator.

[0056] S2. Preprocess the data set with different characteristic categories to meet the data input size of DONN.

[0057] Specifically, in this step S2, the data preprocessing includes data cleaning, i.e. checking whether there are missing values in the data set; data enhancement, i.e. adding Gaussian noise to the data set to make the generalization ability of the neural network stronger; data division, i.e. dividing the data set into training set and test set to meet the data input size of DONN.

[0058] S3. Uniformly initialize the DONN population, and combine the mean square error (MSE) between the true value and the predicted value and the classification accuracy (Accuracy) as the fitness evaluation function of the individual. In this step, the DONN population is uniformly initialized, so that the parameters of each network individual are uniformly distributed in the solution space, and the optimization speed of the network is accelerated. In addition, the mean square error (MSE) between the true value and the predicted value and the classification accuracy (Accuracy) are combined as the fitness evaluation function of the individual, so that the evaluation standard is more persuasive and can better reflect the overall performance of DONN. The specific expression is:

[0059]

[0060] Wherein, represents the mean square error between the true value and the predicted value, n represents the sample number of the test set, Y i and represent the true value and the predicted value of the i-th sample respectively, Accuracy represents the classification accuracy of DONN, and q represents the weighted average factor, which is usually 0.5.

[0061] S4. Exponential ranking selection (ERS) and uniform crossover (UC) are selected as selection operator and crossover operator of the training process, and a mixed mutation strategy is adopted, in which single-point mutation (SM), uniform mutation (UM) and Gaussian mutation (GM) are allocated to different individuals for mutation according to dynamic game probabilities; in this step, exponential ranking selection (ERS) and uniform crossover (UC) are selected as selection operator and crossover operator of the training process, which can increase the diversity of the population and speed up the convergence of the network. A mixed mutation strategy is adopted, in which single-point mutation (SM), uniform mutation (UM) and Gaussian mutation (GM) are allocated to different individuals for mutation according to dynamic game probabilities, which can enhance the optimization ability of the algorithm.

[0062] S5. A double-elite reservation strategy is adopted to reserve the two individuals with the best MSE and Accuracy performance to the next generation, and after multiple iterations of evolution, a DONN individual with global optimal network parameters is obtained until the termination condition is met. In this step, a double-elite reservation strategy is adopted to reserve the two individuals with the best MSE and Accuracy performance to the next generation, which prevents the loss of the current optimal individual and avoids the algorithm from failing to find the global optimal solution. The termination condition is that the network is iterated 2000 times, and finally a DONN individual with global optimal network parameters is obtained.

[0063] As shown in Figure 2 Fig. 1 is a preferred DONN framework structure diagram. MZIs and EOA complete matrix linear operation and nonlinear activation operation respectively, which can be regarded as a Layer in DONN. The whole network contains N layers of MZIs+EOA. The number of phase parameters in MZIs is related to the features of input data, and there are n phase parameters for n input features; there are three hyperparameters in EOA, which are photodetector tapping power α, phase gain g and bias phase θ; the last layer Dropmask is used to remove redundant neurons, so that the output conforms to the classification of the data set. 2

[0064] In this embodiment, the data set for the classification task is shown in Table 1, which is Iris (Iris), Seeds (wheat seeds) and Wine (wine). Iris has 150 samples, each sample contains 4 features and is divided into three categories; Seeds has 210 samples, each sample contains 7 features and is divided into three categories; Wine has 178 samples, each sample contains 13 features and is divided into three categories.

[0065] Table 1 is a data set description table provided in Example 1.

[0066] Table 1

[0067]

[0068] As​Figure 3 The figure shows a flow chart of the hybrid mutation strategy genetic algorithm, and the specific steps are as follows: uniformly initialize the DONN population. In this embodiment, the population contains 50 individuals, and the test set is input into each DONN for fitness evaluation; index sorting and selection are performed according to the obtained fitness value to obtain the selected population; the weight and hyperparameters of each individual in the population are extracted and binary encoded. To reduce the loss of accuracy, 20-bit binary encoding is used in this embodiment; a uniform crossover operator is used with a crossover probability of 0.8; a hybrid mutation strategy is used with a mutation probability of 0.1; the new generation network parameters are obtained, and after decoding, the double elite retention strategy is used to copy and generate the next generation population, and iterate continuously until the termination condition is met to obtain the DONN individual with the global optimal network parameters.

[0069] In step S4, the mixed mutation strategy refers to three different mutation operators with dynamic game probability distribution Among them, r refers to the rth iteration and i refers to the i-th network individual in this iteration.

[0070] If the network individuals produce offspring through the mutation operator Q, and the offspring are selected to enter the next iterative evolution, the probability distribution of each operator is iteratively updated according to the following formula:

[0071]

[0072] If a network individual produces offspring through the mutation operator Q, and the offspring is not selected to enter the next iterative evolution, the probability distribution of each operator is iteratively updated according to the following formula:

[0073]

[0074] Where Q represents the current mutation operator, Q' represents other mutation operators, η∈(0,1) is a coefficient that controls the probability distribution of the hybrid mutation strategy, which is equal to 1 / 3 here, and the probability distribution satisfies β represents the number of mutation operators.

[0075] In step S5, the dual-elite retention strategy aims to retain the two individuals with the best MSE and Accuracy performance to the next generation and continue to use them in the evolution process of the next generation. Specifically, when the best individual in the next generation is worse than the best individual in the previous generation, a portion of the better-performing individuals from the previous generation, i.e., individuals with high Accuracy and low MSE, will be selected to replace the poorer-performing individuals in the next generation. This ensures that there are at least some excellent individuals in the next generation, avoiding the situation where the global optimal solution is eliminated. In this embodiment, the termination condition is that the network is iterated 2000 times, and finally a DONN individual with the global optimal network parameters is obtained.

[0076] Compared with the prior art, the embodiment has the following beneficial effects:

[0077] 1. The light signal is used as the carrier for information transmission, and has the characteristics of high parallelism, anti-interference, low delay, low power consumption and the like. Therefore, the N-layer deep optical neural network (DONN) built by the basic optoelectronic devices including MZIs and EOA and the like is more efficient and low in consumption than the traditional artificial neural network, and can be widely applied in learning tasks such as image classification, natural language processing, unmanned driving and robots in the future.

[0078] 2. The MSGA has good global search capability, can quickly find all solutions in the solution space, is less likely to fall into local optimization compared with the traditional genetic algorithm, has probability randomness, enhances the optimization capability of the algorithm, and makes the robustness and generalization capability of the final DONN individual stronger and the performance better.

[0079] The preferred embodiment of the application is based on the effective training method of the deep optical neural network of the hybrid mutation strategy genetic algorithm, and the steps are as follows: S1. The linear operation layer based on the cascaded Mach-Zehnder interferometer (MZIs), the nonlinear activation layer based on the electro-optic activator (EOA) and the output layer based on the mask (Dropmask) are sequentially stacked, so as to build an N-layer deep optical neural network (DONN); S2. The data set with different characteristic categories is preprocessed to meet the data input size of the DONN; S3. The mean square error (MSE) between the true value and the predicted value and the classification accuracy (Accuracy) are combined as the fitness evaluation function of the individual; S4. The exponential ranking selection (ERS) and the uniform crossover (UC) are used as the selection operator and the crossover operator in the training process, the hybrid mutation strategy is adopted, three mutation operators including single-point mutation (SM), uniform mutation (UM) and Gaussian mutation (GM) are distributed to different individuals for mutation according to the dynamic game probability; S5. The double-elite reservation strategy is adopted, the two individuals with the best performance of MSE and Accuracy are reserved to the next generation, and after multiple iterations and evolution, the DONN individual with the global optimal network parameter is obtained until the termination condition is met.

[0080] As shown in Figure 7 The embodiment discloses a system based on the above deep optical neural network training method, which comprises the following modules:

[0081] The deep optical neural network building module sequentially stacks the linear operation layer based on the cascaded Mach-Zehnder interferometer MZIs, the nonlinear activation layer based on the electro-optic activator EOA and the output layer Dropmask based on the mask mask, so as to build an N-layer deep optical neural network DONN.

[0082] Preprocessing module: preprocessing the data set with different characteristic categories to meet the data input size of the DONN;

[0083] DONN population initialization module: uniformly initializing the DONN population, and combining the mean square error (MSE) and the classification accuracy (Accuracy) between the true value and the predicted value as the fitness evaluation function of the individual;

[0084] Mixed mutation strategy module: using the exponential ranking selection (ERS) and the uniform crossover (UC) as the selection operator and the crossover operator in the training process, and using the mixed mutation strategy to allocate different individuals to mutation by the single point mutation (SM), the uniform mutation (UM) and the Gaussian mutation (GM) according to the dynamic game probability;

[0085] Double-elite reservation strategy module: using the double-elite reservation strategy to reserve the two individuals with the best MSE and Accuracy to the next generation, and through the iterative evolution until the termination condition is met, to obtain the DONN individual with the global optimal network parameter.

[0086] Other contents of the embodiment can refer to the above deep optical neural network training method embodiment.

[0087] Note that the above are only the preferred embodiments of the present application and the technical principles applied. Those skilled in the art will understand that the present application is not limited to the specific embodiments described herein, and those skilled in the art can make various obvious changes, readjustments and substitutions without departing from the scope of the present application. Therefore, although the present application has been described in more detail through the above embodiments, the present application is not limited to the above embodiments, and can include more other equivalent embodiments without departing from the concept of the present application, and the scope of the present application is determined by the scope of the appended claims.

Claims

1. A deep optical neural network training method based on a hybrid mutation strategy genetic algorithm, characterized in that, Comprise the following steps: S1. Stack the linear operation layer based on cascaded Mach-Zehnder interferometer MZIs, the nonlinear activation layer based on electro-optic activator EOA, and the output layer Dropmask based on mask in sequence to build an N-layer deep optical neural network DONN; S2. Preprocess the data set with different feature categories to meet the data input size of the DONN; S3. Uniformly initialize the DONN population, and combine the mean square error MSE and the classification accuracy Accuracy between the true value and the predicted value as the fitness evaluation function of the individual; S4. Select the exponential ranking selection ERS and the uniform crossover UC as the selection operator and the crossover operator of the training process, and use a hybrid mutation strategy to allocate different individuals for mutation according to the dynamic game probability of three operators, namely single-point mutation SM, uniform mutation UM, and Gaussian mutation GM; S5. Use a double-elite reservation strategy to reserve the two individuals with the best MSE and Accuracy performance to the next generation, and iterate evolution until the termination condition is met to obtain a DONN individual with globally optimal network parameters; In step S1, a Mach-Zehnder interferometer MZI is composed of two 3dB directional couplers and two adjustable phase shifters, the inner phase shifter controls the output splitting ratio, and the outer phase shifter controls the differential output phase; according to the singular value decomposition principle, any real matrix can be decomposed into the product of two unitary matrices and a diagonal matrix, that is, R = I∑J * where R is a real-valued matrix, I is an m x m unitary matrix, Σ is an m x n diagonal matrix, J * is the complex conjugate of the n x n unitary matrix J; where the unitary matrices I and J * implemented by the coupler and phase shifter in the MZI, the diagonal matrix Σ is implemented by the optical attenuator, and a large-scale optical matrix operation is realized by configuring a network of cascaded MZIs; the EOA converts a small part of the optical input power into voltage and modulates the amplitude and phase of the remaining part of the original optical signal; assuming that the optical input signal is c, the generated nonlinear opto-electric activation function is f(c), and the specific expression is: where a is the tap power ratio of the photodetector, is the responsivity of the photodetector, G is the gain of the transimpedance amplifier, H is the transfer function, V b is the static bias voltage, V π is the voltage required for a phase change of p of the modulator.

2. The method of claim 1, wherein the method is based on a hybrid mutation strategy genetic algorithm. In step S2, data preprocessing includes: Data cleaning, that is, checking whether there are missing values in the data set; Data enhancement, that is, adding Gaussian noise to the data set to make the neural network have stronger generalization ability; Data division, that is, dividing the data set into a training set and a test set to meet the data input size of the DONN.

3. The method of claim 1 or 2, wherein the method further comprises: In step S3, the mean square error MSE and the classification accuracy Accuracy between the true value and the predicted value are combined as the fitness evaluation function of the individual, and the specific expression is: where, MSE represents the mean square error between the true value and the predicted value, n represents the number of samples in the test set, Y i and Yi and represent the true value and the predicted value of the i-th sample, respectively, Accuracy represents the classification accuracy of the DONN, and q represents the weighted average factor, which is taken as 0.

5.

4. The method of claim 3, wherein the method is characterized by, In step S4, the mixed mutation strategy refers to that three different mutation operators have dynamic game probability distribution wherein r refers to the rth iteration, and i refers to the ith network individual in the current iteration. If the network individual produces offspring through the mutation operator Q, and the offspring is selected into the next iteration evolution, the probability distribution of each operator is updated according to the following formula: If the network individual produces offspring through the mutation operator Q, and the offspring is not selected into the next iteration evolution, the probability distribution of each operator is updated according to the following formula: Wherein, Q represents the current mutation operator, Q' represents other mutation operators, η∈(0, 1) is a coefficient for controlling the probability distribution of the mixed mutation strategy, equal to 1 / 3, and the probability distribution satisfies β represents the number of mutation operators.

5. The method of claim 4, wherein the method further comprises: In step S5, the double-elite reservation strategy is as follows: when the optimal individual in the next generation is worse than the optimal individual in the last generation, a part of the individuals with better performance in the last generation are selected to replace the individuals with poor performance in the next generation.

6. The method of claim 5, wherein the method further comprises: In step S5, the termination condition is that the network is iterated 2000 times, and finally a DONN individual with globally optimal network parameters is obtained.

7. The method of claim 1-2, wherein, The hybrid mutation strategy genetic algorithm is as follows: uniformly initialize the DONN population, the population contains 50 individuals, input the test set into each DONN, and perform fitness evaluation; according to the obtained fitness value, perform exponential ranking selection to obtain the selected population; extract the weights and hyperparameters of each individual in the population for binary coding; The uniform crossover operator is adopted, the crossover probability is 0.8; the hybrid mutation strategy is adopted, the mutation probability is 0.1; the new generation of network parameters is obtained, after decoding, the double-elite reservation strategy is used to copy to generate the next generation population, and iteration is carried out constantly until the termination condition is met, and the DONN individual with global optimal network parameters is obtained.

8. A system for training a deep optical neural network according to any one of claims 1-7, characterized in that, Comprise the following modules: The deep optical neural network building module: the linear operation layer based on the cascade Mach-Zehnder interferometer MZI, the nonlinear activation layer based on the electro-optic activator EOA, the output layer Dropmask based on the mask are sequentially stacked, so as to build N-layer deep optical neural network DONN; The preprocessing module: the data set with different characteristic categories is preprocessed to meet the data input size of DONN; The DONN population initialization module: the DONN population is uniformly initialized, the mean square error MSE between the true value and the predicted value and the classification accuracy Accuracy are combined as the fitness evaluation function of the individual; The hybrid mutation strategy module: exponential ranking selection ERS and uniform crossover UC are used as selection operators and crossover operators in the training process, a hybrid mutation strategy is adopted, three operators of single-point mutation SM, uniform mutation UM and Gaussian mutation GM are distributed to different individuals for mutation according to dynamic game probability; The double-elite reservation strategy module: the double-elite reservation strategy is adopted, the two individuals with the best MSE and Accuracy are reserved to the next generation, through iterative evolution, until the termination condition is met, and the DONN individual with global optimal network parameters is obtained.

Citation Information

Patent Citations

  • Deployment method and device of CDN (Content Delivery Network) server

    CN112787833A

  • Optical fiber preform preparation process optimization method based on genetic algorithm and BP neural network

    CN114004341A