Laser fault diagnosis method and system based on neural network
By introducing the SE-LSTM algorithm in laser fault diagnosis, combining the SE module and LSTM network, it solves the problem of traditional diagnostic technology that difficult to detect and deal with laser failure in a timely manner, and realizes efficient feature extraction and classification of multi-dimensional signals of the laser, improving the accuracy and reliability of fault diagnosis.
Patent Information
- Application Number
- CN202411970346.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-06-03
AI Technical Summary
Traditional laser fault diagnosis technology is difficult to detect and effectively deal with the failure of semiconductor lasers in a timely manner, and cannot guarantee the long-term stable operation and high reliability of the laser.
Using a neural network-based laser fault diagnosis method, the SE-LSTM algorithm is introduced. By combining the SE module and the LSTM network, efficient feature extraction and classification of the laser multi-dimensional signals is realized, and the feature channel weight is adaptively adjusted to increase the model's attention to important features.
It realizes safety status monitoring for the entire life cycle of the laser, promptly detects abnormal status, warnings for potential faults, provides data support for preventive maintenance, and improves the training efficiency and performance of the model.
Smart Images

Figure CN120084523A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of laser fault detection, and more specifically, to a laser fault diagnosis method and system based on a neural network. Background Art
[0002] Semiconductor lasers, as the cornerstone of the contemporary optoelectronic technology field, since their emergence in the 1960s, with their significant advantages such as wide output wavelength coverage, compact structure design, and high integration, have shown great application potential and value in many key fields such as precision material processing, high-speed optical communication, advanced medical technology, high-precision laser sensing, and military aerospace. The continuous progress of this technology not only promotes the rapid development of related industries but also provides strong support for scientific and technological innovation. However, with the increasing requirements for the performance of semiconductor lasers in various fields, especially the significant increase in the output power demand, their reliability problems have gradually emerged, becoming the key factor restricting the further widespread application and technological upgrade of semiconductor lasers. However, traditional diagnostic technologies are not perfect in timely detecting and effectively dealing with the failure of semiconductor lasers, and cannot fully ensure the long-term stable operation and high reliability of lasers.
[0003] To solve these problems, some existing methods collect the output parameters of lasers and input them into a fault diagnosis model to achieve the positioning of faulty devices. Its advantage is that it does not require complex programming. Only by reasonably selecting measurement parameters, it can diagnose according to the laser structure parameters and is applicable to various models and specifications of lasers. However, its disadvantage is that it highly depends on a large amount of experimental data for model training, restricted by experimental conditions and costs; and although it can locate faulty devices, it is difficult to deeply analyze the fault mechanism, limiting the complex fault diagnosis ability. There are also some existing technologies based on model reinforcement learning. By constructing a world model of the laser state information, virtual data is generated for training. Its advantages are that it significantly shortens the training cycle, reduces costs, and at the same time reduces the troubleshooting time and the skill requirements for operators. However, constructing an accurate world model requires rich prior knowledge and data support, which poses a challenge to complex or unknown laser systems; in the face of emergencies such as rapid degradation or sudden failure, it may not be able to respond in a timely and effective manner. Summary of the Invention
[0004] Aiming at the above defects or improvement requirements of the prior art, the present invention provides a laser fault diagnosis method based on a neural network, introducing the SE-LSTM algorithm. This algorithm combines the SE module and the LSTM network to achieve efficient feature extraction and classification of multi-dimensional signals of lasers (such as drive current, optical power, temperature, etc.). This algorithm can adaptively adjust the weights of feature channels, significantly enhancing the model's attention to important features.
[0005] To achieve the above object, according to the first aspect of the present invention, a laser fault diagnosis method based on a neural network is provided, including the following steps:
[0006] S10, generate a drive current, set characteristic coefficients and basic coefficients in lasers of different degradation types and calculate the drive current;
[0007] S20, perform data normalization processing, use the normalized drive current to replace the drive current in S10 and compress the time length to a set value, organize to obtain several data groups that simultaneously include a set of optical power, normalized drive current, temperature, current threshold, and wavelength, and randomly divide each data group into a neural network training set and a test set;
[0008] S30, construct a neural network model, based on the CNN-LSTM model, introduce the SE module to construct the SE-LSTM algorithm, and use the Adam optimizer to optimize the parameters;
[0009] S40, network training and parameter optimization, input the neural network training set in S20 into the neural network model in S30 for training, use multi-class cross-entropy as the loss function for training, and output the degradation type;
[0010] S50, model evaluation and effect verification, evaluate the trained model through set evaluation indicators and use the test set in S20.
[0011] Further, the method for calculating the drive current in S10 is:
[0012] The lasers of different degradation types include normal, rapidly degrading, suddenly failing, and slowly degrading lasers;
[0013] The characteristic coefficients include optical power P, threshold current I 0 and temperature T;
[0014] The basic coefficients include activation energy E A , scaling parameter u 0 , degradation index n, and non-radiative current; where the activation energy E A , scaling parameter u 0 , degradation index n, and non-radiative current respectively conform to normal distributions, and each data of the basic coefficients is randomly combined to form several basic data groups that simultaneously include a set of E A , u 0 , n, and non-radiative current.
[0015] Further, randomly divide the obtained basic data groups into several parts, and use each basic data group to calculate the minimum drive current I(j) of the lasers of different degradation types changing with time j,
[0016] I(j) = I 0 + ΔI deg (j) + ΔI nr (j) + ΔI cnv (j) + ΔI th (j),
[0017] where j is time, I 0 is the threshold current, ΔI dcg (j) is the current change amount related to the degradation type, ΔI nr (j) is the current change amount caused by non-radiative current, ΔI cnv (j) is the current change amount caused by environmental factors, ΔI th (j) is the change amount of the threshold current I 0 with time;
[0018]
[0019] where α 1 is the degradation current change coefficient, g(t) is the degradation function, j lifc is the expected life time of the device, t is the time variable, m is the degradation exponent;
[0020] ΔInr(j) = βe kj ,
[0021]
[0022] where β is the non-radiative current, k is a constant, P is the optical power, n is the degradation index, u 0 is the scaling parameter, E A is the activation energy, T is the temperature, k B is the Boltzmann constant;
[0023]
[0024] where γ is the environmental factor influence coefficient, P is the optical power, P0 is the reference optical power obtained through experiments for normalizing the optical power, n is the degradation index, ω is the angular velocity of the periodic change of the environmental factor, φ is the initial phase, ΔT is the difference between the environmental temperature and the reference temperature, T ref is the set reference temperature, p is the temperature influence index;
[0025] ΔI th (j) = ζ·j q ·e -λ·j ,
[0026] where ζ is the threshold current change coefficient, q is the threshold current change exponent, λ is the threshold current decay constant.
[0027] Further, the method of using the normalized drive current to replace the drive current in S10 in S20 is as follows:
[0028] To control the differential effects caused by different laser thresholds, the normalized drive current I n is used to replace the drive current I(j):
[0029]
[0030] In the formula, In is the normalized drive current, I 0 is the threshold current, I(j) is the drive current, αT is the temperature influence coefficient, T is the temperature, Tref is the set reference temperature, λI is the saturation current influence coefficient, Isat is the saturation current of the laser, ζage is the aging influence coefficient, and Age is the usage age of the laser.
[0031] Further, the method of introducing the SE module to construct the SE-LSTM algorithm based on the CNN-LSTM model in S30 is as follows:
[0032] The CNN-LSTM model is a deep learning model that combines the convolutional neural network CNN and the long short-term memory network LSTM. The SE-LSTM algorithm is constructed by inserting the SE module after the CNN feature extraction and before the LSTM processing, where the SE module processes the features extracted by the CNN through the channel attention mechanism.
[0033] Further, the LSTM structure is as follows:
[0034] The LSTM structure includes three parts: an input gate, a forget gate, and an output gate, and processes data through the following formula to retain the set information:
[0035] f t = sigmoid(W f h t-l + W f x t + b f ),
[0036] i t = sigmoid(W i h t-1 + W i x t + b i ),
[0037] o t = sigmoid(W o h t-l + W o x t+b o ),
[0038]
[0039] h t =o t ×tanh(C t ),
[0040] where t represents the current discrete time step, t - 1 represents the previous time step of t, f t ,i t ,o t , C t ,h t ,x t are respectively the activation vector of the forget gate, the activation vector of the input gate, the activation vector of the output gate, the latent cell state, the cell state, the hidden state, and the input activation vector at the discrete time step t;
[0041] W f is the forget gate weight matrix, h t-1 is the hidden state at the previous time step of t, b f is the forget gate bias function, and sigmoid is the logistic function;
[0042] W i is the input gate weight matrix, and bi is the input gate bias function;
[0043] Wo is the output gate weight matrix, and bo is the output gate bias function;
[0044] Wc is the state update weight matrix, is the state update bias function;
[0045] C t-1 is the cell state at the previous time step of t, and tanh is the hyperbolic tangent function.
[0046] Furthermore, in S30, the Adam optimizer is used to optimize the parameters, and the parameters include the weight matrix of the convolutional kernel of the CNN, the forget gate weight matrix W f in the LSTM, the input gate weight matrix W i , the output gate weight matrix Wo, and the state update weight matrix Wc.
[0047] Furthermore, the method of using multi - class cross - entropy as the loss function for training in S40 is as follows:
[0048]
[0049] where, is the cross-entropy loss function, y is the variable corresponding to y c corresponding variable, is the probability that the neural network model predicts that the sample belongs to the c-th class, y c is the actual label of the c-th class, C is the total number of classes, c is the class serial number, α reg is the coefficient used to adjust the influence of the additional term on the cross-entropy loss. The adjustment additional term includes a regularization term and a mean square error term. λw is the weight factor of the regularization term, is the square of the Frobenius norm of the weight matrix W, βmse is the weight factor of the mean square error term, γ kl is the weight factor of the KL divergence term, represents the KL divergence used to measure the difference between probability distributions, P yc,smooth represents the actual label distribution after smoothing, yc,smooth represents the actual label of the c-th class after smoothing, represents the predicted probability distribution.
[0050] Furthermore, the method for evaluating the trained model by setting the evaluation index and using the test set in S20 is as follows:
[0051] The evaluation index is:
[0052]
[0053] In the formula, Recall is the recall rate, Precision is the precision rate, accuracy is the accuracy rate, F1 is the comprehensive evaluation index, TP, FN, FP, and TN all represent the relationship between the true result and the predicted result. Among them, TP means that both the predicted result and the true result are positive classes; FN means that the true result is a positive class but the predicted result is a negative class; FP means that the true result is a negative class but the predicted result is a positive class; TN means that both the true result and the predicted result are negative classes;
[0054] Based on the above evaluation index, a correction is made,
[0055] BalAcc = (Recall + TN_rate) / 2,
[0056] TN_rate = TN / (TN + FP),
[0057] In the formula, BalAcc is the balanced accuracy rate, and TN_rate is the true negative rate;
[0058]
[0059]
[0060] In the formula, Take F1 as the final evaluation index βw Let βw be the weighted F1 score, βw be the F1 weight coefficient, af be the accuracy factor, and b b be the balance addition coefficient, and B bv be the balance reference value;
[0061] The method for evaluating the trained model using the test set in S20 is to let the test set enter S40 and be brought into the evaluation index in S50 for calculation, and evaluate the accuracy of the trained model.
[0062] According to the second aspect of the present invention, a laser fault diagnosis system based on a neural network is provided, including:
[0063] A drive current generation module, which sets characteristic coefficients and basic coefficients in lasers of different degradation types and calculates the drive current;
[0064] A data normalization processing module, which uses the normalized drive current to replace the drive current in S10 and compresses the time length to a set value, and arranges to obtain several data groups that simultaneously include a set of optical power, normalized drive current, temperature, current threshold, and wavelength, and randomly divides each data group into a neural network training set and a test set;
[0065] A neural network model construction module, based on the CNN-LSTM model, introduces the SE module to construct the SE-LSTM algorithm, and uses the Adam optimizer to optimize the parameters;
[0066] A network training and parameter optimization module, which inputs the neural network training set in S20 into the neural network model in S30 for training, uses multi-class cross-entropy as the loss function for training, and outputs the degradation type;
[0067] A model evaluation and effect verification module, which evaluates the trained model through the set evaluation index and uses the test set in S20.
[0068] Generally speaking, compared with the prior art through the above technical solutions conceived by the present invention, the following beneficial effects can be achieved:
[0069] 1. The diagnostic method of the present invention introduces the SE-LSTM algorithm. By combining the SE module and the LSTM network, this algorithm realizes the efficient feature extraction and classification of multi-dimensional signals of lasers (such as drive current, optical power, temperature, etc.). This algorithm can adaptively adjust the weights of feature channels, significantly improving the model's attention to important features.
[0070] 2. The diagnostic method of the present invention realizes the monitoring of the safety state of the entire life cycle of the laser by real-time monitoring of the multi-dimensional operation data of the laser and using preprocessing and normalization techniques; it can timely detect the abnormal state of the laser, warn of potential faults, and provide strong data support for the preventive maintenance of the laser.
[0071] 3. The diagnostic method of the present invention uses the Adam optimizer to optimize the relevant parameters of the SE-LSTM algorithm, effectively improving the training efficiency and performance of the model; in addition, in order to accurately evaluate the fault diagnosis effect of the model, the present invention uses parameters such as F1-score to conduct a strict accuracy evaluation on the laser degradation model data, further improving the effectiveness and reliability of the fault diagnosis method proposed by the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] Figure 1 is a schematic flow chart provided by a preferred embodiment of the present invention;
[0073] Figure 2 is the drive current-time curve of the semiconductor laser provided by a preferred embodiment of the present invention under three degradation modes;
[0074] Figure 3 are the laser parameters generated by normal distribution provided by a preferred embodiment of the present invention;
[0075] Figure 4 is the In change trend of the laser under different degradation models after time window compression provided by a preferred embodiment of the present invention;
[0076] Figure 5 is the SE-LSTM architecture provided by a preferred embodiment of the present invention;
[0077] Figure 6 is the LSTM neuron model provided by a preferred embodiment of the present invention;
[0078] Figure 7 is the schematic diagram of the channel attention mechanism provided by a preferred embodiment of the present invention;
[0079] Figure 8 is the schematic diagram of the system structure provided by a preferred embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0080] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0081] Based on the above problems, the present invention takes calculating the drive currents of lasers with different degradation types as the starting step. In this step, the characteristic coefficients of the lasers (such as optical power, threshold current, and temperature) and the basic coefficients (such as activation energy, scaling parameter, degradation index, and non-radiative current) are considered, and these coefficients jointly determine the calculation of the drive current.
[0082] Next, data normalization is performed. Data such as drive current is standardized and organized into a dataset containing information such as degradation type, normalized drive current, time, and wavelength. This step helps to eliminate the differences between different lasers, improve the comparability of data, and the training effect of the neural network model.
[0083] Then, a neural network model based on the CNN-LSTM model with the SE module introduced is constructed. The CNN-LSTM model combines the advantages of the convolutional neural network (CNN) and the long short-term memory network (LSTM), and can extract the spatial features and temporal features of the data. The introduction of the SE module further enhances the model's attention to important features and improves the model's performance. At the same time, the Adam optimizer is used to optimize the model parameters to accelerate the training speed and improve the accuracy of the model.
[0084] In the network training and parameter optimization stage, the organized neural network training set is input into the model for training, and multi-class cross-entropy is used as the loss function for training. This step aims to enable the model to better fit the training data and have the ability to accurately predict new data through repeated iteration and adjustment of the model parameters.
[0085] Finally, the trained model is evaluated using the set evaluation metrics and the test set. The evaluation metrics include recall rate, precision rate, accuracy, and the comprehensive evaluation metric F1, etc. These metrics can comprehensively reflect the performance of the model. The introduction of the test set helps to verify the generalization ability of the model and its performance in practical applications.
[0086] In addition, this patent also proposes a neural network-based laser fault diagnosis method system, which includes parts such as a drive current calculation module, a data normalization processing module, a neural network model construction module, a network training and parameter optimization module, and a model evaluation and effect verification module. Each part works together to achieve efficient and accurate diagnosis of laser faults.
[0087] Please refer to Figure 2 , the degradation process of semiconductor lasers can be statistically summarized into three manifestations: rapid degradation, sudden failure, and slow degradation. As shown in the figure, these three forms are represented by the laser drive current-time curve at a constant output optical power.
[0088] Example 1
[0089] Please refer to Figure 1 , the present invention relates to a laser fault diagnosis method based on a neural network, comprising the following steps:
[0090] S10, generate a drive current, set characteristic coefficients and basic coefficients in lasers of different degradation types, and calculate the drive current;
[0091] The method for calculating the drive current in the S10 is as follows:
[0092] The lasers of different degradation types include normal, rapidly degrading, suddenly failing, and slowly degrading lasers;
[0093] The characteristic coefficients include optical power P, threshold current I 0 and temperature T;
[0094] The basic coefficients include activation energy E A , scaling parameter u 0 , degradation index n, and non-radiative current; wherein the activation energy E A , scaling parameter u 0 , degradation index n, and non-radiative current respectively conform to a normal distribution (please refer to Figure 3 , which is the schematic diagram of the basic coefficient distribution), and each data of the basic coefficients is randomly combined to form several basic data groups that simultaneously include a set of E A , u 0 , n, and non-radiative current.
[0095] Randomly divide the obtained basic data groups into several parts, and use each part of the basic data group to calculate the minimum drive current I(j) of the lasers of different degradation types changing with time j. Among them, j in I(j) is time. In this formula, I(j) represents the minimum drive current of the laser changing with time. Since the minimum drive current of the laser will slowly increase as the use time of the laser increases, part of the current is the energy source of the laser output, and part is the energy source of the heat generated during the operation of the laser.
[0096] I(j) = I 0 + ΔI deg (j) + ΔI nr (j) + ΔI cnv (j) + ΔI th (j),
[0097] In the formula, j is time, I 0 is the threshold current, ΔI dcg (j) is the current change amount related to the degradation type, ΔI nr(j) is the amount of current change caused by non-radiative current, ΔI cnv (j) is the amount of current change caused by environmental factors, ΔI th (j) is the threshold current I 0 Change amount over time;
[0098] In some preferred embodiments, 1500 samples are generated for each laser degradation mode and the multi-dimensional operation data of normal lasers, that is, a data set consisting of 6000 time series is constructed. Its laser parameters include optical power, current threshold, wavelength, temperature, and laser degradation type. Given that the lengths of the time series data of rapid degradation, gradual degradation, and diving degradation of lasers are different, the 6000 sets of generated time series data are preprocessed and normalized.
[0099]
[0100] In the formula, α 1 is the degradation current change coefficient, g(t) is the degradation function, j life is the expected life time of the device, t is the time variable, m is the degradation index;
[0101] ΔInr(j) = βe kj ,
[0102]
[0103] In the formula, β is the non-radiative current, k is a constant, P is the optical power, n is the degradation index, u 0 is the scaling parameter, E A is the activation energy, T is the temperature, k B is the Boltzmann constant;
[0104]
[0105] In the formula, γ is the environmental factor influence coefficient, P is the optical power, P0 is the reference optical power obtained through experiments for normalizing the optical power, n is the degradation index, ω is the angular velocity of the periodic change of environmental factors, φ is the initial phase, ΔT is the difference between the environmental temperature and the reference temperature, T ref is the set reference temperature, p is the temperature influence index;
[0106] ΔI th (j) = ζ·j q ·e -λ·j ,
[0107] In the formula, ζ is the threshold current change coefficient, q is the threshold current change index, and λ is the threshold current decay constant.
[0108] S20. Data normalization processing. Use the normalized drive current to replace the drive current in S10 and compress the time length to a set value. Sort out several data sets that simultaneously include a set of optical power, normalized drive current, temperature, current threshold, and wavelength. Randomly divide each data set into a neural network training set and a test set;
[0109] In some preferred embodiments, after shuffling the entire data set, 70% of it is taken as the neural network training set in the present invention. The input includes optical power, current threshold, wavelength, and temperature; the output is the degradation type, and the labels are 0, 1, 2, and 3 respectively. The remaining 30% is the test set.
[0110] The method of using the normalized drive current to replace the drive current in S20 is as follows:
[0111] To control the differential effects brought about by different laser thresholds, use the normalized drive current I n to replace the drive current I(j):
[0112]
[0113] In the formula, In is the normalized drive current, I 0 is the threshold current, I(j) is the drive current, αT is the temperature influence coefficient, T is the temperature, Tref is the set reference temperature, λI is the saturation current influence coefficient, Isat is the saturation current of the laser, ζage is the aging influence coefficient, and Age is the usage age of the laser.
[0114] Please refer to Figure 4 , in some preferred embodiments, compress the laser degradation data with different time lengths by compressing each degradation time series to a time window size of 100. Shown is the change trend of In after time compression for four different degradation modes of the laser.
[0115] Please refer to Figure 5 and Figure 6 , S30. Build a neural network model. Based on the CNN-LSTM model, introduce the SE module to build the SE-LSTM algorithm, and use the Adam optimizer to optimize the parameters;
[0116] This model introduces the channel attention mechanism by combining the SE module, enabling the model to adaptively adjust the feature channel weights, improving the model's attention and representation ability to important features, and helping the model capture the relationship between the optical power, wavelength, and temperature of the laser and the laser current; finally, LSTM, as a neural network structure suitable for processing sequence data, can capture the dependencies between features and effectively perform regression prediction on the laser fault mode.
[0117] The method of introducing the SE module into the SE-LSTM algorithm based on the CNN-LSTM model in S30 is as follows:
[0118] The CNN-LSTM model is a deep learning model that combines the Convolutional Neural Network (CNN) and the Long Short-Term Memory Network (LSTM). The SE-LSTM algorithm is constructed by inserting the SE module after CNN feature extraction and before LSTM processing, where the SE module processes the features extracted by CNN through a channel attention mechanism.
[0119] Convolutional Neural Networks (CNN) include convolutional layers, pooling layers, and fully connected layers. The main function of the convolutional layer is feature extraction. Its implementation is mainly achieved by several convolutional kernels, which are equivalent to weight matrices. After passing through the activation function, the feature map is output. The convolutional kernel can be regarded as an n×n weight matrix, whose function is to convert a submatrix of the previous layer into a unit matrix of the next layer through product calculation. The convolution operation refers to the convolutional kernel sliding from top to bottom and from left to right on the input feature map, and calculating the sum of the dot products with the submatrix at each position as the pixel value of the output at that position. The main parameters involved in this step are the sliding step of the convolutional kernel at each step, the number of convolutional kernels, and the size of the convolutional kernel.
[0120] To adjust the size of the feature map and retain more edge information, the padding operation fills the appropriate number of 0-value elements around the feature map to prevent the loss of edge pixel information during the convolution process and enables the input and output feature sizes to be consistent.
[0121] To reduce the output size of the convolutional layer, a pooling layer (Relu) is usually added after the convolutional layer to reduce network parameters, lower the model's computational complexity, suppress overfitting, and extract important features of the data to a certain extent. Different from the convolutional layer, the method of outputting nodes in the pooling layer is not to calculate the weighted sum of the submatrix but to select the maximum value or average value.
[0122] The fully connected layer (Fully Connection Layer, FC), also known as the affine layer, is an important hierarchical structure in deep learning neural networks, where each neuron is connected to the neurons of the previous layer. This connection method enables the fully connected layer to integrate all features of all layers, extract global information, and output as the final prediction result. In a convolutional neural network, the fully connected layer can integrate the local features extracted by the convolutional layer, pooling layer, etc. to form a global feature representation and change the data dimension.
[0123] The LSTM structure is as follows:
[0124] The LSTM structure consists of three parts: an input gate, a forget gate, and an output gate, and processes data through the following formula to retain the set information:
[0125] f t = sigmoid(W f h t-1 + W f x t + b f ),
[0126] i t = sigmoid(W i h t-1 + W i x t + b i ),
[0127] o t = sigmoid(W o h t-1 + W o x t + b o ),
[0128]
[0129] h t = o t × tanh(C t ),
[0130] In the formula, t represents the current discrete time step, t - 1 represents the previous time step before t, f t , i t , o t , C t , h t , x t are respectively the activation vector of the forget gate, the activation vector of the input gate, the activation vector of the output gate, the potential cell state, the cell state, the hidden state, and the input activation vector at the discrete time step t;
[0131] W f is the forget gate weight matrix, where the weight size determines which features can be forgotten, h t-1 is the hidden state at the previous time step before t, b f is the forget gate bias function, and sigmoid is the logistic function;
[0132] W i is the input gate weight matrix, and bi is the input gate bias function;
[0133] $W_o$ is the output gate weight matrix, which determines the weights of the input information of the current LSTM cell. Its output is $o$, that is, the input information is processed according to the weights, but it is not yet known which information should be forgotten or remembered. Specific information can be selected through experiments. $b_o$ is the output gate bias function;
[0134] $W_c$ is the state update weight matrix, and its weights determine the output of $C$ (carry state) from the previous cell of the LSTM to the next LSTM cell. is the state update bias function;
[0135] $C$ t-1 is the cell state at the previous time step $t$. Tanh is the hyperbolic tangent function.
[0136] In step S30, the Adam optimizer is used to optimize the parameters. The parameters include the weight matrix of the convolutional kernel of the CNN, the forget gate weight matrix $W_f$ in the LSTM f , the input gate weight matrix $W_i$ i , the output gate weight matrix $W_o$ and the state update weight matrix $W_c$.
[0137] The LSTM determines the update of the memory information through the forget gate, which enables the LSTM to exclude the interference of invalid information; determines the new information to be remembered through the input gate; and the output gate controls the timing of applying the memory to the current state. In the RNN, the weight matrix $W$ is relatively fixed, while $C_t$ can be updated with time steps, and the previous information can be passed on according to the degree of influence. Through this structure, the LSTM can capture long-range dependencies and alleviate the problems of gradient vanishing and gradient explosion. In the time series prediction task, a fully connected layer is connected after the LSTM to integrate the output information of the LSTM, extract key features, and provide a non-linear activation function for non-linear transformation, enhancing the demand for fitting non-linear time series relationships in the prediction task and serving as the output layer to output the prediction results of the final laser degradation mode.
[0138] S40, network training and parameter optimization. Input the neural network training set in S20 into the neural network model in S30 for training, and use multi-class cross-entropy as the loss function for training to output the degradation type;
[0139] The method of using multi-class cross-entropy as the loss function in S40 is as follows:
[0140]
[0141] In the formula, is the cross-entropy loss function, $y$ is the variable corresponding to $y$ c corresponding, is the probability that the neural network model predicts that the sample belongs to the $c$-th class, $y$ cis the actual label of the c-th category, C is the total number of categories, c is the category serial number, α reg is the coefficient for adjusting the influence of the additional term on the cross-entropy loss. The additional term includes a regularization term and a mean square error term. λw is the weight factor of the regularization term. is the square of the Frobenius norm of the weight matrix W. βmse is the weight factor of the mean square error term. γ kl is the weight factor of the KL divergence term. represents the KL divergence used to measure the difference between probability distributions. P yc,smooth represents the actual label distribution after smoothing. yc,smooth represents the actual label of the c-th category after smoothing. represents the predicted probability distribution, where all data can be obtained through experiments or selected according to experimental purposes.
[0142] S50, model evaluation and effect verification. The trained model is evaluated by the set evaluation metrics and using the test set in S20.
[0143] The method for evaluating the trained model by the set evaluation metrics and using the test set in S50 is as follows:
[0144] The evaluation metrics are:
[0145]
[0146] In the formula, Recall is the recall rate, Precision is the precision rate, accuracy is the accuracy, F1 is the comprehensive evaluation index. TP, FN, FP, and TN all represent the relationship between the true result and the predicted result. Among them, TP means that both the predicted result and the true result are positive classes; FN means that the true result is a positive class but the predicted result is a negative class; FP means that the true result is a negative class but the predicted result is a positive class; TN means that both the true result and the predicted result are negative classes.
[0147] The recall rate Recall describes the proportion of positive examples correctly identified by the classifier among all actual positive examples. The precision rate Precision describes the proportion of truly positive examples among those identified as positive examples by the classifier. In addition, the accuracy is the most common metric in classification tasks and can intuitively show the classification performance of the model. The F1-score can also be used to evaluate the algorithm accuracy and comprehensively evaluate the sensitivity and specificity of the classifier, which can be calculated by the harmonic mean of the two.
[0148] Based on the above evaluation metrics, make corrections.
[0149] BalAcc = (Recall + TN_rate) / 2
[0150] TN_rate = TN / (TN + FP),
[0151] where BalAcc is the balanced accuracy and TN_rate is the true negative rate;
[0152]
[0153]
[0154] wherein, is the final evaluation index, F1 βw is the weighted F1 score, βw is the F1 weight coefficient, af is the accuracy factor, b b is the balance addition coefficient, B bv is the balance reference value; The method for evaluating the trained model using the test set in S20 is to let the test set enter S40 and calculate it with the evaluation index in S50 to evaluate the accuracy of the trained model.
[0155] Please refer to Figure 7 , in some preferred embodiments, the SE-block is a lightweight network structure designed to improve the representation ability of convolutional neural networks (CNNs) by explicitly modeling the correlations between feature channels. It mainly includes two key operations: Squeeze and Excitation.
[0156] (1) Input data X: The leftmost box represents the input data X, which is a multi-dimensional tensor, usually having the shape [H, W, C], where H is the height, W is the width, and C is the number of channels.
[0157] (2) Squeeze operation: The input data X first passes through a global pooling layer. The role of the global pooling layer is to compress the global spatial features of each channel into a scalar, that is, to perform global average pooling or global max pooling on each channel. As shown in the figure, the two "1x1xC" boxes from "1x1xC" to "1x1xC" (actually C scalars) represent the same number of input and output channels. After the Squeeze operation, we obtain a compact channel descriptor vector that contains the global information of each channel.
[0158] (3) Excitation operation:
[0159] Following the Squeeze operation is the Excitation operation, which uses a fully connected layer (or called a bottleneck layer) to learn the correlations between channel descriptors.
[0160] "Fex(,W)" represents this fully connected layer, where W is the weight matrix. This fully connected layer usually has a relatively small dimension (for example, compressing C channels to C / r, where r is the reduction ratio, and then expanding back to C channels), and uses a non-linear activation function (such as ReLU) to increase the non-linearity of the model.
[0161] The output of the Excitation operation is a set of weights, which are used to recalibrate the responses of the original feature channels.
[0162] (4) Scale operation:
[0163] Finally, the weights output by the Excitation operation are multiplied element-wise with the original feature channels to achieve recalibration of the feature channels.
[0164] "Fscale()" represents this scaling operation. It takes the output of the Excitation operation and the original feature channels as inputs and outputs the recalibrated feature channels.
[0165] (5) Output data H':
[0166] After being processed by the SE-block, we obtain the output data H'. In the schematic diagram, the rightmost box represents the output data H', which is a multi-dimensional tensor with the same shape as the input data X.
[0167] Embodiment 2
[0168] As Figure 8 shown, as another aspect of the present invention, it also relates to a laser fault diagnosis system based on a neural network, including:
[0169] A drive current generation module, which sets characteristic coefficients and basic coefficients in lasers of different degradation types and calculates the drive current. By accurately calculating the drive current, the actual working state of the laser can be more accurately reflected, providing a reliable basis for subsequent data analysis and model training.
[0170] A data normalization processing module, which uses the normalized drive current to replace the drive current in S10 and compresses the time length to a set value, and organizes several data sets that simultaneously include a set of optical power, normalized drive current, temperature, current threshold, and wavelength. Each of the data sets is randomly divided into a neural network training set and a test set. Data normalization and time length compression help improve the training efficiency and prediction accuracy of the neural network, and at the same time avoid the dimensional interference between different data.
[0171] Build a neural network model module. Based on the CNN-LSTM model, introduce the SE module to construct the SE-LSTM algorithm, and use the Adam optimizer to optimize the parameters. The SE-LSTM algorithm combines the advantages of CNN and LSTM, and introduces the SE module to improve the representation ability of the model, so as to achieve more accurate prediction of the degradation types of lasers. And creatively introduce the Adam optimizer to make the model accuracy higher.
[0172] Network training and parameter optimization module. Input the neural network training set in S20 into the neural network model in S30 for training, use multi-class cross-entropy as the loss function for training, and output the degradation type. By training and optimizing the SE-LSTM algorithm, accurate prediction of the degradation types of lasers can be achieved, providing strong support for the maintenance and management of lasers.
[0173] Model evaluation and effect verification module. Evaluate the trained model through the set evaluation indicators and use the test set in S20. Through model evaluation and effect verification, the prediction performance of the SE-LSTM algorithm can be objectively evaluated, and strong support can be provided for subsequent model optimization and improvement. At the same time, it also helps to verify the effectiveness of data preprocessing and model construction.
[0174] Those skilled in the art can easily understand that the above are only preferred embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A laser fault diagnosis method based on neural network, characterized in that: The following steps are involved: S10, generating a driving current, setting characteristic coefficients and basic coefficients in lasers of different degradation types and calculating the driving current; S20, data normalization processing, using a normalized driving current to replace the driving current in S10 and compressing the time length to a set value, sorting out several data groups including a set of optical power, normalized driving current, temperature, current threshold and wavelength, and randomly dividing each of the data groups into a neural network training set and a test set; S30, build a neural network model. Based on the CNN-LSTM model, introduce the SE module to build the SE-LSTM algorithm, and use the Adam optimizer to optimize the parameters; S40, network training and parameter optimization, inputting the neural network training set in S20 into the neural network model in S30 for training, using multi-classification cross entropy as the loss function of training, and outputting the degradation type; S50, model evaluation and effect verification, evaluates the trained model through the set evaluation indicators and using the test set in S20.
2. The laser fault diagnosis method based on neural network according to claim 1 is characterized in that: The method for calculating the driving current in S10 is: The lasers of different degradation types include normal, fast degradation, sudden failure and slow degradation lasers; The characteristic coefficients include optical power P, threshold current I0 and temperature T; The basic coefficients include the activation energy E A , scaling parameter u0, degradation index n and non-radiative current; wherein the activation energy E A , scaling parameter u0, degradation index n and non-radiative current are in accordance with normal distribution, and the basic coefficients are randomly combined to form several data including a group of E A , u0, n and basic data set of non-radiative current.
3. The laser fault diagnosis method based on neural network according to claim 2 is characterized in that: The basic data set obtained is randomly divided into several parts, and each basic data set is used to calculate the minimum driving current I(j) of the lasers of different degradation types that varies with time j. I(j)=I0+ΔI deg (j)+ΔI nr (j)+ΔI cnv (j)+ΔI th (j), Where j is time, I0 is threshold current, ΔI deg (j) is the current change associated with the degradation type, ΔI nr (j) is the current change caused by non-radiative current, ΔI env (j) is the current change caused by environmental factors, ΔI th (j) is the change of threshold current I0 over time; Where α1 is the degradation current variation coefficient, g(t) is the degradation function, and j life is the expected life time of the device, t is the time variable, and m is the degradation index; ΔInr(j)=βe kj , Where β is the non-radiative current, k is a constant, P is the optical power, n is the degradation index, u0 is the scaling parameter, and E A is the activation energy, T is the temperature, k B is the Boltzmann constant; Where γ is the influence coefficient of environmental factors, P is the optical power, P0 is the reference optical power obtained through experiments for normalizing the optical power, n is the degradation index, ω is the angular velocity of the periodic change of environmental factors, φ is the initial phase, ΔT is the difference between the ambient temperature and the reference temperature, and T ref is the set reference temperature, p is the temperature influence index; ΔI th (j)=ζ·j q ·by -λ·j , Where ζ is the threshold current variation coefficient, q is the threshold current variation index, and λ is the threshold current attenuation constant.
4. The laser fault diagnosis method based on neural network according to claim 3 is characterized in that: The method of using the normalized driving current in S20 to replace the driving current in S10 is: In order to control the difference in impact caused by different laser thresholds, the normalized drive current I n Instead of driving current I(j): Where In is the normalized drive current, I0 is the threshold current, I(j) is the drive current, αT is the temperature influence coefficient, T is the temperature, Tref is the set reference temperature, λI is the saturation current influence coefficient, Isat is the saturation current of the laser, ζage is the aging influence coefficient, and Age is the service age of the laser.
5. The laser fault diagnosis method based on neural network according to claim 4 is characterized in that: In the S30, based on the CNN-LSTM model, the method of introducing the SE module to construct the SE-LSTM algorithm is as follows: The CNN-LSTM model is a deep learning model that combines a convolutional neural network (CNN) and a long short-term memory (LSTM) network. The SE-LSTM algorithm is constructed by inserting an SE module after CNN feature extraction and before LSTM processing, wherein the SE module performs a channel attention mechanism on the features extracted by CNN.
6. The laser fault diagnosis method based on neural network according to claim 5 is characterized in that: The LSTM structure is: The LSTM structure consists of three parts: input gate, forget gate, and output gate. It processes data through the following formula and retains the set information: f t =sigmoid(W f h t-1 +W f x t +b f ), i t =sigmoid(W i h t-1 +W i x t +b i ), o t =sigmoid(W o h t-1 +W o x t +b o ), h t =o t ×tanh(C t ), In the formula, t represents the current discrete time step, t-1 represents the time step before t, and f t ,i t , o t , C t ,h t , x t are the activation vector of the forget gate, the activation vector of the input gate, the activation vector of the output gate, the potential cell state, the cell state, the hidden state, and the input activation vector at the discrete time step t; W f is the forget gate weight matrix, h t-1 is the hidden state of the previous time step t, b f is the forget gate bias function, sigmoid is the logistic function; W i is the input gate weight matrix, bi is the input gate bias function; Wo is the output gate weight matrix, bo is the output gate bias function; Wc is the state update weight matrix, is the state update bias function; C t-1 is the cell state at the time step before t, and tanh is the hyperbolic tangent function.
7. The laser fault diagnosis method based on neural network according to claim 6 is characterized in that: In S30, the Adam optimizer is used to optimize the parameters, including the weight matrix of the convolution kernel of the CNN and the weight matrix W of the forget gate in the LSTM. f , input gate weight matrix W i , output gate weight matrix Wo and state update weight matrix Wc.
8. The laser fault diagnosis method based on neural network according to claim 7 is characterized in that: The method of using multi-classification cross entropy as the loss function for training in S40 is: In the formula, is the cross entropy loss function, y is the c The corresponding variable, The neural network model predicts the probability that the sample belongs to the cth category, y c is the actual label of the cth category, C is the total number of categories, c is the category number, α reg is a coefficient for adjusting the influence of the additional term on the cross entropy loss, wherein the adjusted additional term includes a regularization term and a mean square error term, λw is a weight factor of the regularization term, is the square of the Frobenius norm of the weight matrix W, βmse is the weight factor of the mean square error term, γ kl is the weight factor of the KL divergence term, represents the KL divergence used to measure the difference between probability distributions, P yc,smooth represents the actual label distribution after smoothing, yc,smooth represents the actual label of the cth category after smoothing, represents the predicted probability distribution.
9. The laser fault diagnosis method based on neural network according to claim 8, characterized in that: The method for evaluating the trained model by using the set evaluation index and the test set in S20 in S50 is: The evaluation indicators are: In the formula, Recall is the recall rate, Precision is the precision rate, accuracy is the accuracy, F1 is the comprehensive evaluation index, TP, FN, FP and TN all represent the relationship between the true result and the predicted result, among which TP means that both the predicted result and the true result are positive; FN means that the true result is positive but the predicted result is negative; FP means that the true result is negative but the predicted result is positive; TN means that both the true result and the predicted result are negative; Based on the above evaluation indicators, corrections are made. BalAcc=(Recall+TN_rate) / 2, TN_rate=TN / (TN+FP), Where BalAcc is the balanced accuracy, TN_rate is the true negative rate; In the formula, is the final evaluation index, F1 βw is the weighted F1 score, βw is the F1 weight coefficient, af is the accuracy factor, b b is the balance addition coefficient, B bv is the balance reference value; The method for evaluating the trained model using the test set in S20 is to allow the test set to enter S40 and be brought into the evaluation index in S50 for calculation, so as to evaluate the accuracy of the trained model.
10. A laser fault diagnosis method system based on neural network, characterized in that: include: Generate a drive current module, set characteristic coefficients and basic coefficients in lasers of different degradation types and calculate the drive current; A data normalization processing module, using a normalized driving current to replace the driving current in S10 and compressing the time length to a set value, to obtain several data groups including a set of optical power, normalized driving current, temperature, current threshold and wavelength, and randomly dividing each of the data groups into a neural network training set and a test set; Build a neural network model module. Based on the CNN-LSTM model, introduce the SE module to build the SE-LSTM algorithm, and use the Adam optimizer to optimize the parameters. A network training and parameter optimization module, which inputs the neural network training set in S20 into the neural network model in S30 for training, uses multi-classification cross entropy as the loss function of training, and outputs the degradation type; The model evaluation and effect verification module evaluates the trained model through the set evaluation indicators and using the test set in S20.
Citation Information
Cited By
Intelligent aging data analysis method and system for laser
CN121327755A
Industrial equipment fault diagnosis and early warning method and system based on intelligent equipment
CN121499130A