A Method for HRRP Sequence Recognition Based on Lightweight Transformer
By adopting lightweight Transformer in HRRP sequence recognition, rotating position coding, local aggregate attention unit and lightweight feedforward neural network, combined with Label Smoothing regularization, the problem of difficulty in extracting deep and global timing information in HRRP sequences in the prior art is solved, and efficient recognition performance and lightweight deployment are achieved.
Patent Information
- Application Number
- CN202310124590.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-16
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2043-02-16
AI Technical Summary
The prior art is difficult to effectively extract the deep and global timing information of HRRP sequences, resulting in poor recognition performance in real scenarios, and traditional methods have large parameters and complex calculations, making it difficult to deploy on edge devices.
A HRRP sequence recognition method based on lightweight Transformer is proposed, using rotating position encoding to embed position information, combined with local aggregation attention units and lightweight feedforward neural network, and using Label Smoothing regularization to enhance the generalization performance of the model.
The performance of HRRP sequence recognition is significantly improved, the number of parameters and computations of the model is reduced, making it suitable for deployment on edge devices, and shows better generalization performance and robustness under variant samples and finite sample conditions.
Smart Images

Figure CN116645615B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of radar, and specifically refers to a method for identifying HRRP sequences based on lightweight Transformer. Background Technique
[0002] Radar automatic target recognition has extensive applications in both military and civilian fields. Currently, high-resolution range profiles (HRRPs) are widely used in radar automatic target recognition because of their advantages of easy acquisition, convenient processing, and small storage space. HRRP is the vector sum of the target echo along the radar line of sight, containing rich target structure and scatter point distribution information, which has important application value for target recognition and has been widely used in the recognition of aircraft, ships, ballistic missiles, and ground targets.
[0003] When the radar detects a moving target, relative motion occurs between the radar and the target, and echo information at multiple azimuth angles can be obtained. The HRRPs at consecutive azimuth angles form an HRRP sequence. By modeling the correlation between HRRP sequences, the dynamic time features of the target can be effectively extracted for efficient target recognition. The HRRP sequence is a special multivariate time series. Due to the large dimension of HRRP features and the presence of a large amount of noise information in the real scene, the current HRRP recognition methods have limited ability to perceive global information and poor ability to represent target information. Therefore, enhancing and extracting features from global information is the key issue for further improving the performance of HRRP sequence recognition.
[0004] Traditional HRRP sequence recognition methods rely on manual extraction of highly distinguishable features. For example, Timothy et al. proposed an identification method based on the Hidden Markov Model (HMM), which extracted six power spectrum features from the high-resolution (HRR) radar signal amplitude and the target distance profile and used HMM for identification. Du et al. proposed an identification method based on a double-distribution composite statistical model. According to the number of dominant scattering points in the range cells of the scattering center model, the range cells were divided into three statistical types, and the echoes of different types of range cells were modeled as corresponding distribution forms to complete the identification task. Molchanov et al. proposed an identification method based on micro-Doppler double-coherence features. The proposed method extracted cepstral coefficients from the micro-Doppler contributions in the radar echo and used double-coherence estimation to calculate classification features. Manually extracted features are greatly affected by subjective factors of humans and have weak expression ability for features, so the recognition performance is limited. To overcome the limitations of manual feature extraction, machine learning methods are introduced into the HRRP recognition task. Lei et al. proposed an identification method based on support vector machines. According to the distances between classifiers given by the confusion matrix, different classifier confidences were defined, and then the values of support vector machines and posterior probabilities were integrated into the basic probability assignment to achieve an identification method combining support vector machines and evidence theory. Wang et al. proposed a network of Extreme Learning Autoencoders with One-Dimensional Local Receptive Fields (1DELM-LRF-AE) for meaningful representation learning of the local structure of HRRP, achieving efficient representation learning and recognition. The above methods overcome the influence of subjective factors of researchers and achieve more effective automatic feature extraction, but they ignore the correlation and temporal information between HRRPs, resulting in information loss, and the ability of machine learning methods to extract deep features is weak, and the recognition performance still needs to be improved.
[0005] With the development of deep learning, convolutional neural networks and recurrent neural networks are widely used in HRRP sequence recognition. For example, Xiang et al. proposed an identification method based on one dimensional CNN (1D-CNN), which extracts effective target structure information in HRRP through 1D-CNN and introduces an aggregation-perception-recalibration module for feature enhancement. 1D-CNN can effectively extract the local correlation of HRRP sequences, but ignores the temporal information between HRRP sequences. Due to the limitation of the convolutional kernel size in CNN, there are limitations in extracting the global features of long sequences, and the HRRP sequences in the actual scenario contain a large amount of noise information. The local information will damage the generalization of the model due to the influence of noise. Du et al. proposed an identification method based on Region-factorizeDrecurrent attentional network with deep clustering, which uses the time dependence of the recurrent neural network (RNN) to automatically find the information region in HRRP samples and weights the different recognition contributions of the hidden states at each time step. However, for long sequences, RNN will lose important target features due to the memory loss problem when extracting long-range information, and when the network is stacked deeper, the original important information may be lost. To further enhance the ability to extract global information, Pan et al. proposed an identification method based on CNN-Bi-RNNwith Attention Mechanism, which uses a convolutional neural network to obtain a richer embedding representation, and then uses an RNN based on the attention mechanism to extract temporal information, which can more effectively utilize local and global temporal features and still maintain a high recognition performance for limited samples. With the birth of the Transformer framework, the self-attention mechanism is used to model the long-range information of the sequence, and different weights are adaptively assigned to the sequence by calculating the correlation between sequences, paying more attention to the important information in the target area and effectively extracting the global information of the sequence. Zhang et al. proposed an identification method based on feature-guideDTransformer, which effectively enhances the model's extraction of global information and reduces the model's dependence on the number of samples by adding artificial features in the attention module to guide the model to focus on the distance units with more scattering information.
[0006] Current methods for HRRP sequence recognition cannot effectively extract the deep and global temporal information of HRRP sequences, so their performance in real scenarios is poor, such as environmental noise, variant targets, and limited samples. Moreover, most current methods ignore the lightweight nature of the method and improve the recognition performance by stacking a large number of modules, resulting in a large number of method parameters, complex calculations, and difficulty in meeting the deployment and application on edge devices. To address the above problems, the present invention proposes a lightweight Transformer (RLAT) based HRRP sequence recognition method using rotational position encoding, local aggregation attention unit (LAU), and lightweight feed-forward neural network (LW-FFN). Summary of the Invention
[0007] The present invention aims to solve the above technical problems and provides a method for HRRP sequence recognition based on a lightweight Transformer.
[0008] To solve the above technical problems, the technical solution provided by the present invention is: a method for HRRP sequence recognition based on a lightweight Transformer, comprising the following steps:
[0009] S1. Embed position information for the HRRP sequence using rotational position encoding;
[0010] Rotational position encoding is a multiplicative encoding that realizes relative position encoding through absolute position encoding. The calculation of rotational position encoding is as follows: Assume the HRRP sequence is X = [x1, x2,..., x N , where N represents the sequence length. Traditional absolute position encoding adds position information after the embedding layer. The calculation method for absolute position encoding to embed position information is
[0011] f(x i , i) = W(x i + p i )(1)
[0012] where i is the position of x i , is a trainable d-dimensional position vector, and d depends on x i . Relative position encoding adds relative position information in the self-attention mechanism. The calculation method for embedding relative position is
[0013]
[0014] where is a trainable relative position matrix, c = clip(m - n, r min , r max)It represents the relative position between positions m and n. As the distance between the relative positions increases, the correlation between the data weakens. Therefore, a restricted range r is set for c max , and if it exceeds the restricted range, the correlation is considered consistent; after adding the position matrix, the absolute position matrix in the process of calculating the attention weights in the self-attention mechanism is
[0015]
[0016] The core idea of relative position encoding is to replace the absolute position vectors p embedded in the third and fourth terms n with relative position vectors p m-n , and replace the in the third and fourth terms with two trainable vectors u T and v T . To distinguish content and position, the parameter matrices are represented by W k and W' k respectively. Therefore, the way of calculating the weights in the self-attention mechanism in relative position encoding becomes
[0017]
[0018] Rotary position encoding improves relative position encoding, and its calculation is
[0019]
[0020] where the rotation matrix is
[0021]
[0022] where When rotary position encoding calculates the attention weights in the self-attention mechanism, it can obtain
[0023]
[0024] S2. A local aggregation attention unit (LAU) is proposed. Using grouped linear transformation can effectively aggregate and extract local features, and using the self-attention mechanism can realize the perception and enhancement of global information;
[0025] LAU includes grouped linear transformation, non-linear activation, layer normalization, self-attention mechanism and residual connection. It effectively enhances features through local feature aggregation and global perception. The encoded vector I = g(X) is used as the input of LAU. After LAU performs local feature aggregation and global feature enhancement, the enhanced feature F is output. In the local feature aggregation stage, assume there are a total of L layers of grouped linear transformation. In the first layers, the dimension is increased, and the remaining The dimension of the layer is reduced, and the number of groups for the grouped linear transformation of each layer is calculated. The specific process is as follows
[0026]
[0027] where n l is the number of groups of the grouped linear transformation of the l-th layer, and n max is the maximum value of the number of groups of the grouped linear transformation. The calculation of the grouped linear transformation of each layer is as follows
[0028]
[0029] where M(·) represents performing residual concatenation, the non-linear activation function GELU, and layer normalization operations. The local feature aggregation obtains deep local features by first increasing the dimension and then reducing the dimension of the local features, and the residual concatenation operation is adopted in each layer of the grouped linear transformation to avoid losing the original information when the model is deepened. The low-dimensional aggregated features are more conducive to being processed by the self-attention mechanism. The low-dimensional aggregated features are used as the input of the self-attention mechanism to perform global perception operations. According to the relevant principles of information retrieval, the self-attention mechanism calculates the correlation of sequence data through the query vector and the key vector to obtain the attention matrix, and then calculates the globally enhanced features with the value matrix as follows
[0030]
[0031] where the query vector Q = YW Q , the key vector K = YW K and the value vector V = YW V , where W Q , W K , W V are learnable parameter matrices;
[0032] Finally, to ensure that the dimensions of the input and output are consistent, the vector needs to be first increased in dimension after being processed by the self-attention mechanism. At the same time, to avoid losing important features due to the excessive depth of the model, a residual connection is added at the end to retain the important information in the original features, and we can obtain
[0033] F = I + W s A(11)
[0034] where W s is the parameter matrix for increasing the dimension of the output of the self-attention mechanism;
[0035] S3. Use the lightweight feed-forward neural network (LW-FFN) to extract the enhanced features;
[0036] Assume that the dimension of the input feature F is d m , and the lightweight FFN first reduces the dimension to Raise the dimension to d again m , on the premise of unchanged performance, the number of parameters is reduced by 16 times. The calculation process of the lightweight FFN for feature extraction is
[0037] O m = Feedforward(F) = F + ((Relu(FW1 + b1))W2 + b2) (12)
[0038] In the formula, W1 and b1 are the parameter matrix and bias vector during dimensionality reduction respectively, W2 and b2 are the parameter matrix and bias vector during dimensionality reduction respectively, and Relu(·) is a non-linear activation function;
[0039] After being processed by the stacked M-layer LAT encoder, finally use SoftMax as the classifier, and the output result is
[0040]
[0041] In the formula, O M is the output of the M-layer LAT encoder, exp(·) represents the exponential function of e, and K represents the total number of categories;
[0042] S4. Adopt Label Smoothing regularization to introduce noise to the sample labels;
[0043] When encoding the labels of samples, the probability distribution of the traditional one-hot encoding is
[0044]
[0045] To enhance the generalization performance of the model, improve the one-hot encoding. Label Smoothing introduces soft one-hot encoding to add fuzzy noise to the labels. After adding Label Smoothing, the probability distribution of the soft one-hot labels is
[0046]
[0047] In the formula, K represents the total number of categories in the multi-classification task, i represents the category number, and ε is a hyperparameter;
[0048] When using the cross-entropy loss function to calculate the loss value between the predicted value and the true value, the calculation method of the cross-entropy loss function is
[0049]
[0050] In the formula, p i is the probability of the true value, and q i is the probability of the predicted value;
[0051] In Label Smoothing regularization, the loss functions for each category are calculated as
[0052]
[0053] where ε is a hyperparameter. During the neural network training process, when minimizing the cross-entropy loss value between the predicted value and the true value, the optimal predicted probability distribution is
[0054]
[0055] where K represents the total number of categories in the multi-classification task, ε is a hyperparameter, and α is an arbitrary real number.
[0056] The beneficial effects of the present invention are as follows:
[0057] 1. The present invention studies the application of Transformer in HRRP sequence recognition and proposes a method for HRRP sequence recognition based on lightweight Transformer, demonstrating that the Transformer-based method has more excellent performance in HRRP sequence recognition compared to the methods based on convolutional neural networks and recurrent neural networks;
[0058] 2. The present invention improves the traditional Transformer using local aggregation attention units, lightweight FFN, and Label Smoothing regularization. On the premise of effectively achieving lightweight, the recognition performance is further improved, which is beneficial for deployment on edge devices;
[0059] 3. The present invention verifies the effectiveness and generalization of the proposed method on variant samples and limited samples, and verifies the performance of position encoding and important hyperparameters in the HRRP sequence task.
[0060] The present invention first uses rotational position encoding to more efficiently embed relative position information, and proposes a lightweight local aggregation attention unit to perform local feature aggregation and global perception operations on high-dimensional HRRP sequence data. Feature aggregation can not only suppress the adverse effects of noise regions through group linear transformation, obtain a richer local feature representation, but also effectively reduce the number of parameters. The aggregated low-dimensional features are input into the self-attention mechanism, which can complete global information perception and enhancement, extract the long-range correlations in time and space of the HRRP sequence, effectively enhance the ability to extract important information in the target region, obtain deep temporal features with high separability in the HRRP sequence, and avoid the problem of information loss in deep networks through residual connections. Then, a lightweight LW-FFN is used to achieve feature extraction, which greatly reduces the number of parameters compared with traditional FFN. Finally, label smoothing is used to introduce label noise, avoid the model's over-reliance on limited training samples, and enhance the generalization performance of the proposed method in real scenarios. Experiments show that the method proposed by the present invention significantly improves the recognition performance on the premise of effectively reducing the number of parameters, and achieves better generalization performance and robustness in variant sample and limited sample experiments, and is more lightweight.
[0061] The above summary is only for the purpose of the specification and is not intended to be limiting in any way. In addition to the illustrative aspects, embodiments, and features described above, further aspects, embodiments, and features of the present invention will become apparent by reference to the drawings and the following detailed description. Brief Description of the Drawings
[0062] Figure 1 It is a schematic diagram of the implementation process of the present invention.
[0063] Figure 2 It is a schematic diagram of the structure of the LAU of the present invention.
[0064] Figure 3 It is a schematic diagram of the structure of the LW-FFN of the present invention.
[0065] Figure 4 It is a schematic diagram of the generation of the HRRP sequence of the present invention.
[0066] Figure 5 It is a partial sample diagram of the MSTAR dataset of the present invention.
[0067] Figure 6 It is a line graph of the recognition accuracy of 10 targets in the MSTAR dataset 1 of the present invention.
[0068] Figure 7 It is a line graph of the recognition accuracy of 10 targets in the MSTAR dataset 2 of the present invention.
[0069] Figure 8 It is a comparison chart of the recognition accuracy of different methods with limited sequence lengths in the experiments of the present invention.
[0070] Figure 9 It is a comparison chart of the recognition accuracy of different methods with limited training data in the experiments of the present invention.
[0071] Figure 10 It is a schematic diagram of the recognition performance of the five position encodings of the present invention for the standard dataset D1.
[0072] Figure 11 It is a schematic diagram of the recognition performance of the five position encodings of the present invention for the standard dataset D2.
[0073] Figure 12 It is a schematic diagram of the influence of important hyperparameters in the experiments of the present invention. Detailed implementation manners
[0074] An HRRP sequence recognition method based on a lightweight Transformer includes the following steps:
[0075] S1. Embed position information for the HRRP sequence by using rotational position encoding;
[0076] Rotational position encoding is a multiplicative encoding that realizes relative position encoding through absolute position encoding. The calculation of rotational position encoding is as follows: Assume that the HRRP sequence is X = [x1, x2,..., x N , where N represents the sequence length. The traditional absolute position encoding adds position information after the embedding layer. The calculation method for the absolute position encoding to embed position information is
[0077] f(x i , i) = W(x i + p i )(1)
[0078] where i is the position of x i , is a trainable d-dimensional position vector, and d depends on x i . The absolute position encoding adds independent position information to the sequence data at each position. The position vector requires a large amount of data training to obtain better performance. Since the HRRP sequence is usually non-cooperative target data and the data volume is limited, the absolute position encoding is not suitable for solving the HRRP sequence recognition problem.
[0079] Relative position encoding adds relative position information in the self-attention mechanism. The calculation method for embedding relative positions is
[0080]
[0081] where is a trainable relative position matrix, c = clip(m - n, r min , r max ) represents the relative position between positions m and n. As the distance between relative positions increases, the correlation between data weakens. Therefore, c sets a limit range r max , and if it exceeds the limit range, the correlation is considered consistent; after adding the position matrix, the absolute position matrix in the process of calculating the attention weights in the self-attention mechanism is
[0082]
[0083] The core idea of relative position encoding is to replace the absolute position vector p embedded in the third and fourth terms n with the relative position vector p m-n , and replace the in the third and fourth terms with two trainable vectors u T and v T . To distinguish content and position, the parameter matrices are represented by W k and W' k , so the way of calculating the weights in the self-attention mechanism in relative position encoding becomes
[0084]
[0085] To fully utilize the relative position information in the HRRP sequence to extract deep temporal information, rotational position encoding improves relative position encoding, and its calculation is
[0086]
[0087] where the rotation matrix is
[0088]
[0089] where When rotational position encoding calculates the attention weights in the self-attention mechanism, it can obtain
[0090]
[0091] Introduce relative position information using the rotation matrix and rotational position encoding avoids introducing more learnable parameter matrices and introduces relative position information in a more lightweight way.
[0092] S2. Propose a local aggregation attention unit (LAU), which can effectively aggregate and extract local features using grouped linear transformation, and use the self-attention mechanism to realize the perception and enhancement of global information;
[0093] Combined with the appendixFigure 2 , different from the way most methods perform high-dimensional mapping on input information, LAU performs dimensionality reduction on the input HRRP sequence, aggregates features of local information through grouped linear transformation, and then uses the self-attention mechanism to perform global perception on the aggregated low-dimensional features, effectively enhancing the ability to extract global information. Since the HRRP sequence contains a large amount of redundant noise information, shallow high-dimensional mapping will confuse noise information and target information. However, the present invention adopts multi-layer grouped linear transformation to perform local feature aggregation on local features by first increasing the dimension and then reducing the dimension. The low-dimensional aggregated features are more easily processed by the self-attention mechanism, and local feature aggregation can effectively improve the attention ability of the self-attention mechanism to global information.
[0094] LAU includes grouped linear transformation, non-linear activation, layer normalization, self-attention mechanism and residual connection, and effectively enhances features through local feature aggregation and global perception. The encoded vector I = g(X) is used as the input of LAU. After local feature aggregation and global feature enhancement by LAU, the enhanced feature F is output. In the local feature aggregation stage, assuming there are a total of L layers of grouped linear transformation, the first layers perform dimension increase operations, and the remaining layers perform dimension reduction. The process of calculating the number of groups for each layer of grouped linear transformation is as follows
[0095]
[0096] where n l is the number of groups of the l-th layer of grouped linear transformation, and n max is the maximum value of the number of groups of grouped linear transformation. The calculation of each layer of grouped linear transformation is
[0097]
[0098] where M(·) represents performing residual splicing, non-linear activation function GELU and layer normalization operations. Local feature aggregation obtains deep local features by first increasing the dimension and then reducing the dimension of local features, and residual splicing operations are adopted in each layer of grouped linear transformation to avoid losing the original information as the model deepens. The low-dimensional aggregated features are more easily processed by the self-attention mechanism. The low-dimensional aggregated features are used as the input of the self-attention mechanism for global perception operations. According to the relevant principles of information retrieval, the self-attention mechanism calculates the correlation of sequence data through query vectors and key vectors to obtain an attention matrix, and then calculates the globally enhanced feature with the value matrix as
[0099]
[0100] where the query vector Q = YW Q , the key vector K = YW KThe sum vector V = YW V , where W Q , W K , W V is a learnable parameter matrix;
[0101] Finally, to ensure that the dimensions of the input and output are consistent, the vector needs to be upsampled after the self-attention mechanism. At the same time, to avoid the loss of important features caused by an overly deep model, a residual connection is added at the end to retain the important information in the original features, and we can obtain
[0102] F = I + W s A(11)
[0103] In the formula, W s is the parameter matrix for upsampling the output of the self-attention mechanism;
[0104] S3. Use a lightweight feedforward neural network (LW-FFN) to extract the enhanced features;
[0105] LAU has a deeper network structure and more significant feature enhancement ability compared with the traditional multi-head attention mechanism. Therefore, the present invention uses a lightweight feedforward neural network to replace the traditional feedforward neural network for feature extraction. Assume that the dimension of the input feature F is d m , combined with the attached Figure 3 , the traditional FFN( Figure 3 shown on the left) uses the method of first upsampling to 2d m and then downsampling to d m for feature extraction, and the number of parameters is very large. Since LAU has a significant feature enhancement effect, the present invention uses a lightweight FFN( Figure 3 shown on the right) for feature extraction. The lightweight FFN consists of a fully connected layer, a non-linear activation function, and a residual connection, and is mainly used for feature extraction. The lightweight FFN first downsamples to and then upsamples to d m . On the premise of unchanged performance, the number of parameters is reduced by 16 times. The calculation process of the lightweight FFN for feature extraction is
[0106] O m = Feedforward(F) = F + ((Relu(FW1 + b1))W2 + b2)(12)
[0107] In the formula, W1 and b1 are the parameter matrix and bias vector during downsampling respectively, W2 and b2 are the parameter matrix and bias vector during upsampling respectively, and Relu(·) is a non-linear activation function;
[0108] After being processed by the stacked M-layer LAT encoder, finally, SoftMax is used as the classifier, and the output result is
[0109]
[0110] Wherein, O M is the output of the M-th layer LAT encoder, exp(·) represents the exponential function of e, and K represents the total number of categories;
[0111] S4. Introduce noise to the sample labels by using Label Smoothing regularization;
[0112] In the context of non-cooperative targets, the current number of HRRP sequence samples is limited, and the HRRP in the real scenario is in a complex noise environment. There are still certain differences in the HRRP of the same target, resulting in an overfitting problem prone to occur during HRRP sequence recognition. To solve the overfitting problem, a Label Smoothing regularization strategy is introduced. Label Smoothing adds label noise, avoids the model relying too much on limited training samples, and enhances the generalization performance of the proposed method in real scenarios.
[0113] When encoding the labels of samples, the probability distribution of the traditional one-hot encoding is
[0114]
[0115] To enhance the generalization performance of the model, the one-hot encoding is improved. Label Smoothing introduces soft one-hot encoding to add fuzzy noise to the labels, thereby reducing the proportion of real sample labels in calculating the loss, so that the model does not rely too much on limited samples and avoids falling into local optimal solutions, ultimately realizing the suppression of the overfitting problem. After adding Label Smoothing, the probability distribution of the soft one-hot label is
[0116]
[0117] Wherein, K represents the total number of categories in the multi-classification task, i represents the category number, and ε is a hyperparameter;
[0118] When using the cross-entropy loss function to calculate the loss value between the predicted value and the true value, the calculation method of the cross-entropy loss function is
[0119]
[0120] Wherein, p i is the probability of the true value, and q i is the probability of the predicted value;
[0121] During the training process of the neural network, the model is optimized towards the direction of smaller loss values. However, over-reliance on the training set data will reduce the generalization performance of the HRRP sequence recognition task in real scenarios. The Label Smoothing regularization strategy can prevent the network from being overconfident, slow down the penalty intensity for loss, and avoid the model falling into local optimal solutions. The loss functions for each category are calculated as
[0122]
[0123] where ε is a hyperparameter. During the neural network training process, when minimizing the cross-entropy loss value between the predicted value and the true value, the optimal predicted probability distribution is
[0124]
[0125] where K represents the total number of categories in the multi-classification task, ε is a hyperparameter, and α is an arbitrary real number.
[0126] From the predicted probability distribution, it can be concluded that Label Smoothing regularization can increase the tolerance for errors between the true value and the predicted value, prevent the model from over-relying on the training set samples, thereby avoiding the model falling into local optimal solutions, and enhancing the generalization performance of the model.
[0127] Experimental Results and Analysis:
[0128] a) The MSTAR dataset is a standard dataset widely used for SAR target recognition. Its data source is a high-resolution spotlight synthetic aperture radar, which operates in the X-band, has a resolution of 0.3m × 0.3m, and uses the HH polarization mode. The MSTAR dataset includes 10 types of targets, namely BMP2 (SN-9566), BRDM-2, BTR70 (SN-C71), D7, T62, T72 (SN-132), ZIL131, and ZSU23 / 4. The data with a pitch angle of 17° in the dataset is used as the training set, and the data with a pitch angle of 15° is used as the test set. The azimuth angles of all targets cover 0 to 360°. Dataset 1 includes 2747 SAR images in the original MSTAR dataset training set, and the test set includes 2348 SAR images. To further test the generalization performance of the model, 4 variant targets, namely BMP2 (SN-9563), BMP2 (SN-C21), T72 (SN-812), and T72 (SN-S7), are added to the test set to form Dataset 2. Dataset 2 includes 2747 SAR images in the original MSTAR dataset training set, and the test set includes 3203 SAR images. In the present invention, the SAR images are converted into HRRP sequences, and the composition of the MSTAR sequence dataset 2 is shown in Table 1.
[0129] Table 1 MSTAR sequence dataset2
[0130]
[0131] The conversion steps are as follows: First, convert the dataset into a complex SAR image, and then perform a one-dimensional inverse fast Fourier transform (IFFT) on the azimuth dimension of the complex SAR image. The data obtained along the range dimension is the HRRP complex sequence. Then, take the modulus of the HRRP complex sequence to obtain the HRRP sequence. Each complex SAR image can obtain 100 HRRP samples. Taking the average of every 10 HRRP samples can obtain 10 average HRRP samples. Therefore, the training set of the original MSTAR dataset includes 24,270 HRRP samples, and the test set includes 32,030 HRRP samples.
[0132] Assume that the length of the generated HRRP sequence is L (L ≤ 50). The sliding window algorithm for generating the HRRP sequence is shown in Algorithm 1 in Table 2:
[0133]
[0134] Combined with Appendix Figure 4 , use the sliding window method to process the HRRP data. Divide the azimuth angle of 360° into 50 azimuth blocks, each azimuth block contains 7.2°. The sampling interval of each SAR image is 1°. Each SAR image can be processed to obtain 10 average HRRP samples. Therefore, the sampling interval of each average HRRP sample is 0.1°. In many research works, denoised and enhanced samples are used as the model input, resulting in poor generalization of the model. The present invention only performs energy normalization on the HRRP data. Because in actual situations, data loss at individual angles often occurs due to aircraft movement, the dataset does not perform interpolation processing on the missing data in the MSTAR dataset, making the data more in line with the real situation.
[0135] After being processed by the sliding window method, Dataset 1 can obtain 24,270 HRRP sequence samples for the training set and 23,480 HRRP sequence samples for the test set; Dataset 2 can obtain 24,270 HRRP sequence samples for the training set and 32,030 HRRP sequence samples for the test set.
[0136] Combined with Appendix Figure 5 , the HRRP sequence samples of some targets are as Figure 5 (a)-(d) shown, Figure 5 (e)-(h) are the corresponding HRRP respectively. It can be seen that as a dataset in the real scene, the MSTAR samples contain a large amount of noise redundancy information, causing great difficulties and challenges for effective feature extraction during recognition.
[0137] b) Performance comparison experiment:
[0138] To verify the recognition performance of the proposed method, the present invention selects six commonly used baseline methods, namely LSTM, GRU, 1D-CNN, TCN, gMLP, and XCM, and Transformer as the comparison methods. To achieve the optimal model performance, the model architecture and parameters of the comparison experiment are designed according to the reference literature. The recognition performance of each method is verified on the dataset MSTAR Dataset 1. The recognition results of the comparison experiment on 10 types of targets are shown in Table 2.
[0139] Table 2 Recognition accuracy of compare experiments on dataset 1(%)
[0140]
[0141] As can be seen from Table 2, the RLAT proposed by the present invention has the highest average accuracy for the recognition of ten types of targets, reaching 99.86%, which is 17.09% higher than the traditional gMLP, more than 3.69% higher than the commonly used recurrent neural networks LSTM and GRU, 0.61% and 1.92% higher than the excellent convolutional neural networks 1D-CNN and TCN respectively, 0.96% higher than XCM, and 0.54% higher than Transformer with the same network structure. The RLAT proposed by the present invention achieves the optimal recognition performance for 6 types of targets, TCN achieves the optimal performance for 5 types of targets, and the other methods are less than 5, indicating that the proposed method has more stable recognition performance compared with other methods.
[0142] Combined with the attached Figure 6 , attached Figure 6 The accuracy of the 10 targets in the MSTAR dataset 1. The numbers on the x-axis respectively represent the 10 targets in dataset 1. It can be seen that the recognition performance of the proposed method for ten types of targets is more stable and balanced. Other methods all have certain recognition shortcomings. In particular, the recognition performance of gMLP and LSTM fluctuates greatly, with the fluctuation ranges being 31.17% and 17.50% respectively, the fluctuation range of GRU being 9.71%, the fluctuation ranges of 1D-CNN and TCN being 4.03% and 13.32% respectively, the fluctuation range of XCM being 5.20%, and the fluctuation range of Transformer being 2.15%. The maximum fluctuation range of the proposed method is less than 0.55%. This shows the effectiveness of the RLAT feature extraction proposed by the present invention, which can suppress the adverse effects of noise information and effectively extract highly distinguishable target information.
[0143] In addition, the parameter and computational amounts of the proposed RLAT have achieved significant lightweighting. Compared with the parameter and computational amounts of the original Transformer, the parameter amount has been reduced by 90.90%, and the computational amount has been reduced by 96.70%. Since RLAT uses the LAU module, the parameter and computational amounts are greatly reduced while ensuring performance. Except for GRU, the parameter amount of RLAT is less than that of other comparison models, and the computational amount of the proposed method is less than that of other comparison models. It can be seen that the computational amounts of 1D-CNN and XCM have increased significantly due to the introduction of the convolution module. The lightweight model is more conducive to being deployed on edge devices and in practical applications.
[0144] c) Robustness comparison experiment:
[0145] In actual application scenarios, HRRP data usually comes from non-cooperative targets. Currently, various non-cooperative targets usually include variant models, and there are certain differences in the shape and structure between the variant targets and the original targets. The recognition performance for variant targets is one of the important factors to measure the robustness of the method. By setting up a comparison experiment, the robustness of the proposed method on the variant dataset is verified on the dataset MSTAR Dataset 2. The training set of Dataset 2 is the same as that of Dataset 1, and only about 36% more variant samples are added to the test set, and they are concentrated in the BMP2 and T72 class targets, forming an unbalanced dataset. The experimental results are shown in Table 3:
[0146] Table 3 Recognition accuracy of compare experiments on dataset2(%)
[0147]
[0148] As can be seen from Table 3, due to the addition of variant samples, the generalization performance requirements for the model are relatively high, and the recognition performances of all methods have decreased. The proposed method still achieves the highest average recognition accuracy of 99.73%, which is only 0.13% lower than that of Dataset 1. The recognition accuracy of gMLP has decreased by 18.55%, and those of LSTM and GRU have decreased by 0.37% and 4.14% respectively, those of 1D-CNN and TCN have decreased by 1.90% and 1.16% respectively, and that of Transformer has decreased by 1.01%. In addition, the RLAT proposed in the present invention achieves the optimal recognition performance on 8 targets, and the number of targets for other comparison methods is less than 4, showing better performance than Dataset 1. The proposed method shows more significant stability and robustness on the variant dataset and has stronger generalization performance for variant data.
[0149] Combined with the appendix Figure 7 , appendix Figure 7The accuracy of the 10 targets in the MSTAR dataset 2, where the numbers on the x-axis represent the 10 targets in dataset 1. It can be seen that the recognition performance of RLAT for variant datasets is more stable and balanced, with a maximum fluctuation range of only 2.04%. The fluctuation ranges of the comparison methods LSTM and GRU are 22.76% and 16.64% respectively, those of 1D-CNN and TCN are 5.37% and 10.49% respectively, the fluctuation range of Transformer is 7.49%, and the fluctuation range of XCM is 11.58%. The fluctuation range of gMLP reaches 60.19%. This shows that the global temporal features extracted by RLAT are more robust and stable, with higher separability. At the same time, Label Smoothing can avoid over-reliance on training samples and further improve the generalization performance for variant samples.
[0150] d) Comparative experiment with limited samples:
[0151] The recognition of HRRP sequences under the condition of limited samples is one of the major challenges currently faced. Regarding the sequence length and the number of training samples of HRRP sequences, the recognition performance for limited samples is verified. A limited sequence length can improve the real-time performance of model recognition and enable target recognition to be completed several time periods in advance according to requirements. Limited training samples can effectively verify the generalization performance of the model under non-cooperative target conditions and can effectively shorten the training time. At the same time, limited samples mean less target information and pose higher requirements for the feature extraction ability of the model. A comparative experiment is set up to verify the recognition performance of the proposed method under limited sample conditions using HRRP sequence samples with limited sequence lengths and limited training samples. Using the HRRP sequence generation algorithm, the lengths of the HRRP sequences are set to {1, 2, 4, 8, 16}, and the training samples are set to {1%, 2%, 5%, 10%, 50%, 90%} of the original training dataset. Relatively strict experimental conditions are set to verify the recognition performance of the proposed method under limited sample conditions. The experimental results are shown in Tables 4 and 5.
[0152] Table 4 Recognition accuracy for limited sequence length(%)
[0153]
[0154] As can be seen from Table 4, the recognition accuracy of HRRP sequences increases with the increase of sequence length. The proposed method shows more significant performance than other comparative experiments on most short sequences. On the standard dataset D1 and the variant dataset D2, the recognition performance of RLAT is better than that of methods other than Transformer. When the sequence length is 1 in the D1 dataset and the sequence lengths are 1 and 2 in the D2 dataset, since Transformer has a more complex model structure than RLAT, and the information contained in sequences of length 1 and 2 is less, the recognition performance is slightly higher. However, the recognition performance of RLAT is better than that of other comparative methods. As the sequence length increases, FLAT can perform feature selection more effectively, eliminating the adverse effects of redundant information contained in HRRP sequences, so it can achieve more significant recognition performance than Transformer.
[0155] Combined with the attached Figure 8 , the traditional Transformer relies on a large number of training samples to improve the recognition performance, which severely restricts the application of Transformer in the field of HRRP recognition. The ability of RLAT in feature enhancement and feature extraction is more prominent, and it can effectively improve the recognition performance for limited training samples. To verify the recognition performance of the proposed method under the condition of limited training samples, the sequence length is set to 32, and the number of training samples is {1%, 2%, 5%, 10%, 50%, 90%} of the original training dataset respectively. Among them, 1% of the original training set only contains 274 training samples, and the recognition performance is shown in Table 5.
[0156] Table 5 Recognition accuracy for limiteDtraining data(%)
[0157]
[0158] As can be seen from Table 5, the recognition accuracy of HRRP sequences increases with the increase in the number of training samples. On the standard dataset D1 and the variant dataset D2, RLAT achieves more prominent recognition performance than other methods on datasets with a limited number of training samples. Notably, when the training set is only 1% of the original training set, RLAT still achieves an accuracy of 95.83% on the MSTAR standard dataset D1, which is 15.41% higher than that of Transformer, 76.18% higher than that of the traditional gMLP, more than 35.8% higher than that of the LSTM and GRU recurrent neural network methods, and more than 2.42% higher than that of the TCN, 1D-CNN, and XCM convolutional neural network methods. On the variant dataset D2, it achieves an accuracy of 93.36%, which is 18.15% higher than that of Transformer, 75.75% higher than that of the traditional gMLP, more than 44.01% higher than that of the LSTM and GRU recurrent neural network methods, and more than 2.42% higher than that of the TCN, 1D-CNN, and XCM convolutional neural network methods.
[0159] Combined with the attached Figure 9 , RLAT can achieve significant recognition performance when the number of training samples is only 274, while other methods are more dependent on the number of samples and are severely affected in training performance when the number of training samples drops sharply. This shows that RLAT has a more prominent generalization performance under the condition of limited samples and can more effectively complete the recognition of HRRP sequences under non-cooperative targets.
[0160] 3.5 Positional Encoding Experiment
[0161] RLAT is a network based on the self-attention mechanism and is thus insensitive to the position information of HRRP sequences. Therefore, positional encoding is needed to add position information to HRRP sequences in order to extract temporal features more efficiently. Currently, the commonly used positional encodings mainly include absolute positional encoding and relative positional encoding. The absolute positional encoding LAPE and the relative positional encodings Bias, C1, and C2, as well as Rope, are selected for comparative experiments to verify the effectiveness of different positional encodings. The experimental results are shown in Table 6:
[0162] Table 6 Recognition accuracy of different positional encoding methods(%)
[0163]
[0164] As can be seen from Table 6, in the standard dataset D1, Rope achieved an identification accuracy of 99.86%, which is more than 0.12% higher than other relative position encoding methods and 0.45% higher than absolute position encoding. The identification accuracy of relative position encoding is slightly higher than that of absolute position encoding. In the variant dataset D2, Rope achieved an identification accuracy of 99.73%, which is more than 1.34% higher than other relative position encoding methods and 1.23% higher than absolute position encoding. At this time, the identification accuracy of absolute position encoding is slightly higher than that of relative position encoding.
[0165] Combined with the attached Figure 10 , the recognition performance of different position encodings can be analyzed more intuitively. For the MSTAR standard dataset D1, which contains 10 types of military targets, RoPE achieved the optimal recognition performance for 8 types of targets, and the maximum difference in recognition accuracy was 0.55%. C1 achieved the optimal performance for 6 types of targets, and the maximum difference in recognition accuracy was 2.30%. Bias achieved the optimal performance for 5 types of targets, and the maximum difference in recognition accuracy was 1.33%. C2 achieved the optimal performance for 5 types of targets, and the maximum difference in recognition accuracy was 1.43%. LAPE achieved the optimal performance for 5 types of targets, and the maximum difference in recognition accuracy was 3.58%. It can be analyzed that because RoPE combines the advantages of absolute position encoding and relative position encoding and is more conducive to the extraction of temporal features, the recognition performance of RoPE is significantly better than other methods. At the same time, the relative position encoding not only has a higher accuracy than the absolute position encoding method, but also has a smaller fluctuation difference in the recognition accuracy for 10 types of targets, and the recognition performance is more balanced. Because relative position encoding can more effectively extract the relative information between HRRP sequences, it is beneficial to extract temporal correlation.
[0166] Combined with the attached Figure 11, for the MSTAR variant dataset D2, variant targets are added to the test set. RoPE achieves the optimal recognition performance for 8 types of targets, and the maximum difference in recognition accuracy is 2.04%. C1 achieves the optimal performance for 3 types of targets, and the maximum difference in recognition accuracy is 5.73%. Bias achieves the optimal performance for 3 types of targets, and the maximum difference in recognition accuracy is 3.58%. C2 achieves the optimal performance for 3 types of targets, and the maximum difference in recognition accuracy is 6.13%. LAE achieves the optimal performance for 3 types of targets, and the maximum difference in recognition accuracy is 5.47%. It can be analyzed that the recognition performance of RoPE is significantly better than other methods because RoPE combines the advantages of absolute position encoding and relative position encoding, which is more conducive to the extraction of temporal features and has stronger robustness to variant samples. At the same time, it is worth noting that the recognition performance of relative position encoding decreases significantly for the variant dataset, and the average recognition accuracy of absolute position encoding is higher than that of relative position encoding. Because the variant dataset adds 36% variant samples, which requires higher generalization ability for recognition methods, and relative position encoding introduces more parameters in the self-attention mechanism, resulting in overfitting problems in the recognition of variant samples.
[0167] Influence of hyperparameters:
[0168] Hyperparameters play a key role in deep learning models. In RLAT, the Embedding of the LAU mapping dimension, the Layer of the LAT stacking layer, and the Depth of the LAU maximum depth, these three hyperparameters have important influences on the model performance. Among them, Embedding can effectively control the width of a single LAU, and the stacking layer and the LAU maximum depth can affect the feature extraction ability of the model from the aspect of model depth. At the same time, the three hyperparameters are coupled with each other, so the three hyperparameters are combined to verify their influence on the model. Set the range of the mapping dimension Embedding = {64, 128, 256, 512}, the range of the stacking layer Layer = {1, 2, 3, 4, 5, 6, 7, 8}, and the LAU maximum depth Depth = {2, 4, 6, 8, 10}, where, Figure 12 (a)-(c) show the experimental results on the MSTAR standard dataset D1, Figure 12 (d)-(f) show the experimental results on the MSTAR standard dataset D2.
[0169] For the MSTAR standard dataset D1, as shown in Figure 12 (a), the recognition performance of the proposed method increases with the increase of Embedding and Depth. Because with the increase of Embedding and Depth, the width and depth of LAU increase, and more rich and abstract features can be extracted. As shown in Figure 12(As shown in (b), the recognition performance of the proposed method shows an increasing trend with the increase of Embedding, and shows a trend of increasing first and then decreasing with the increase of Layer. It can be seen that too many stacked LATs will instead damage the recognition performance, and too deep a model will lead to a sharp increase in the number of parameters, resulting in a serious overfitting problem of the model. From Figure 12 (c), the recognition performance of the proposed method shows a trend of increasing first and then decreasing with the increase of Layer, and shows a slow growth trend with the increase of Depth. It can be analyzed that Embedding and Depth mainly affect the width and depth of a single LAU, so they show a positive correlation trend, while Layer controls the depth of the whole model, so it shows a trend of increasing first and then decreasing. Among them, the recognition performance is the best when Layer = 2, and too deep a model will lead to an exacerbation of the overfitting problem.
[0170] For the MSTAR variant dataset D2, due to the addition of a large number of variant samples, the requirement for the generalization performance of the model is higher. From Figure 12 (d), the recognition performance of the proposed method shows a trend of increasing first and then decreasing with the increase of Embedding and Depth. Because with the increase of Embedding and Depth, the width and depth of LAU increase, and more rich and abstract features can be extracted. However, with the increase in the number of parameters, the overfitting problem will be more significant on the variant dataset. From Figure 12 (e), the recognition performance of the proposed method shows a trend of increasing first and then decreasing with the increase of Embedding and Layer. Appropriate model width and depth are beneficial to enhancing the feature extraction ability, but with the increase of the model width and depth, it will lead to a sharp increase in the number of parameters, exacerbating the overfitting problem for the variant dataset. Figure 12 (f), the recognition performance of the proposed method shows a trend of increasing first and then decreasing with the increase of Layer and Depth. It can be analyzed that different from the standard dataset, there are certain differences between the training set samples and the test set samples in the variant dataset, and the overfitting problem is very likely to occur. Therefore, when the width and depth of the model increase, it shows a trend of increasing first and then decreasing. Among them, the recognition performance is the best when Embedding = 128, Layer = 2, and Depth = 6. Too wide and too deep models will extract a large amount of redundant noise information, leading to an exacerbation of the overfitting problem. Studying the variant dataset also more conforms to the actual needs.
[0171] The present invention explores the application of Transformer in HRRP sequence recognition, and proposes a method for HRRP sequence recognition based on lightweight Transformer called RLAT. This method uses more lightweight rotary position encoding, local aggregation attention unit and lightweight feed-forward neural network, and its recognition performance in real scenarios is significantly better than other methods, and effectively reduces the number of model parameters and computational complexity, which helps the application and deployment on edge devices. The present invention also explores the recognition performance of the proposed method under variant targets and limited sample conditions, and verifies that the generalization performance of the proposed method is significantly better than other methods. Finally, the present invention further studies the influence of position encoding on HRRP sequence recognition, as well as the influence of other important hyperparameters of the proposed method.
[0172] The above describes the present invention and its implementation manners. Such description is not restrictive. What is shown throughout the text is only one of the implementation manners of the present invention, and the actual structure is not limited thereto. Generally speaking, if those of ordinary skill in the art are inspired by it and design similar structural modes and embodiments to this technical solution without creative efforts without departing from the purpose of the present invention creation, they shall fall within the protection scope of the present invention.
Claims
1. A method for HRRP sequence recognition based on lightweight Transformer, characterized in that, It includes the following steps: S1. Embed position information into the HRRP sequence using rotational position encoding; Rotary position encoding is a multiplicative encoding that achieves relative position encoding through absolute position encoding. The calculation of rotary position encoding is as follows: Suppose the HRRP sequence is X = [x1, x2,..., x N , where N represents the sequence length. Traditional absolute position encoding adds position information after the embedding layer. The calculation method for embedding position information by absolute position encoding is f(x i , i) = W(x i + p i )(1) where i is the position of x i position, is a trainable d-dimensional positional vector, where d depends on x i The relative position encoding adds relative position information to the self-attention mechanism. The calculation method for embedding the relative position is Among them is a trainable relative position matrix, c = clip(m - n, r min , r max ) represents the relative position between positions m and n. As the distance between relative positions increases, the correlation between data weakens. Therefore, c sets a limit range r max . If it exceeds the limit range, the correlation is considered consistent; after adding the position matrix, the absolute position matrix is in the process of calculating the attention weights in the self-attention mechanism as The core idea of relative position encoding is to replace the absolute position vectors p embedded in the third and fourth terms n with the relative position vector p m-n , and replace the in the third and fourth terms with two trainable vectors u T and v T . To distinguish content and position, the parameter matrices are denoted by W k and W k ′ respectively. Therefore, the way of calculating weights in the self-attention mechanism in relative position encoding becomes The rotational position encoding improves the relative position encoding, and its calculation is where the rotation matrix is Among them When the rotation position encoding calculates the attention weights in the self-attention mechanism, it can obtain S2. Propose a local aggregation attention unit, which can effectively aggregate and extract local features using grouped linear transformation, and use the self-attention mechanism to realize the perception and enhancement of global information; The local aggregation attention unit includes grouped linear transformation, non-linear activation, layer normalization, self-attention mechanism and residual connection, and effectively enhances features through local feature aggregation and global perception. The encoded vector I = g(X) is used as the input of the local aggregation attention unit. After passing through the local aggregation attention unit for local feature aggregation and global feature enhancement, the enhanced feature F is output. In the local feature aggregation stage, assuming there are a total of L layers of grouped linear transformation, the dimensionality increase operation is performed in the first layers, and the remaining layers are used for dimensionality reduction. The specific process of calculating the number of groups for each layer of grouped linear transformation is where n l is the number of groups of the l-th layer of grouped linear transformation, and n max is the maximum value of the number of groups of grouped linear transformation. The calculation of each layer of grouped linear transformation is In the formula, M(·) represents performing residual concatenation, the nonlinear activation function GELU, and layer normalization operations. The local feature aggregation obtains deep local features from local features by first increasing the dimension and then reducing the dimension, and residual concatenation operations are adopted in each layer of grouped linear transformation to avoid the loss of original information due to the deepening of the model. The low-dimensional aggregated features are more conducive to being processed by the self-attention mechanism. The low-dimensional aggregated features are used as the input of the self-attention mechanism for global perception operations. According to the relevant principles of information retrieval, the self-attention mechanism calculates the correlation of sequence data through the query vector and the key vector to obtain the attention matrix, and then calculates the globally enhanced features with the value matrix as where the query vector Q = YW Q , the key vector K = YW K and the value vector V = YW V , where W Q , W K , W V are learnable parameter matrices; Finally, to ensure that the dimensions of the input and output are consistent, the vector needs to be first upsampled after being processed by the self-attention mechanism. At the same time, to avoid the loss of important features caused by the model being too deep, a residual connection is added at the end to retain the important information in the original features, and we can get F = I + W s A (11) where, W s is a parameter matrix for dimension elevation processing of the output of the self-attention mechanism; S3. Extract the enhanced features using a lightweight feed-forward neural network; Assume that the dimension of the input feature F is d m , the lightweight FFN first reduces the dimension to and then increases the dimension to d m , on the premise of unchanged performance, the number of parameters is reduced by 16 times. The calculation process of the lightweight FFN for feature extraction is O m = Feedforward(F) = F + ((Relu(FW1 + b1))W2 + b2)(12) In the formula, W1 and b1 are the parameter matrix and bias vector during dimensionality reduction respectively, W2 and b2 are the parameter matrix and bias vector during dimensionality reduction respectively, and Relu(·) is the nonlinear activation function; After being processed by the stacked M-layer LAT encoder, finally use SoftMax as the classifier, and the output result is where O M is the output of the LAT encoder of the M-th layer, exp(·) represents the exponential function of e, and K represents the total number of categories; S4. Introduce noise to the sample labels using Label Smoothing regularization; When encoding the labels of samples, the probability distribution of the traditional one-hot encoding is To enhance the generalization performance of the model, the one-hot encoding is improved. Label Smoothing introduces soft one-hot encoding to add fuzzy noise to the labels. After adding Label Smoothing, the probability distribution of the soft one-hot labels is In the formula, K represents the total number of categories in the multi-classification task, i represents the category number, and ε is the hyperparameter; When using the cross-entropy loss function to calculate the loss value between the predicted value and the true value, the calculation method of the cross-entropy loss function is where p i is the probability of the true value, and q i is the probability of the predicted value; The loss function calculation for each category in Label Smoothing regularization is In the formula, ε is the hyperparameter. During the neural network training process, when minimizing the cross-entropy loss value between the predicted value and the true value, the optimal predicted probability distribution is In the formula, K represents the total number of categories in the multi-classification task, ε is the hyperparameter, and α is an arbitrary real number.