A Radar Target Recognition Method Based on Optimized Capsules

By applying a deep learning model of SimSiam-CapsNet structure on small embedded devices, combining multi-scale CNN and CapsNet, using self-attention routing mechanism and lightweight SE layer, the problem of device computing and storage limitations is solved, and efficient radar target recognition is achieved.

CN115062754BActive Publication Date: 2025-05-27HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210391546.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-14
Publication Date
2025-05-27
Estimated Expiration
2042-04-14

AI Technical Summary

Technical Problem

Existing radar target recognition methods are difficult to achieve efficient identification on small embedded devices, mainly due to limited computing power, storage capacity and power consumption, and the traditional methods require high degree of completeness of target data, so they cannot effectively process incomplete data that is not cooperative.

Method used

A deep learning model based on SimSiam-CapsNet structure is proposed, combining multi-scale CNN structure and CapsNet to optimize network parameters through self-attention routing mechanism and lightweight SE layer, reduce computing and storage requirements, and improve identification performance.

Benefits of technology

It realizes efficient identification of radar targets on small embedded devices, reduces model complexity and parameter volume, improves identification performance and computing efficiency, and is suitable for practical radar identification engineering applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115062754B_ABST
    Figure CN115062754B_ABST
Patent Text Reader

Abstract

The present invention discloses a radar target recognition method based on optimized capsules. First, the original HRRP sample set is preprocessed for data augmentation; feature extraction is performed through the SimSiam module, and the extracted high-dimensional features are fed into the convolutional module. The features extracted by the convolutional module are input into the basic capsule network based on the SE layer. Finally, a classifier is built to classify HRRP targets. The recognition of HRRP targets is achieved through a new routing mechanism, realizing faster convergence and higher recognition performance, capturing the separable information at different levels of HRRP targets in detail from multiple perspectives, and increasing the sparsity of the network. Optimizing the quality of basic capsules and reducing the number of capsule parameters greatly reduces the number of trainable parameters in the network and also improves the recognition performance to a certain extent, providing a feasible method for the actual radar recognition engineering application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of radar target recognition. Specifically, it relates to a radar target recognition method for optimizing capsules. Background Art

[0002] With the continuous maturity of radar technology, most of the early traditional HRRP recognition methods based on, such as statistical models, manifold learning, and kernel methods, have been able to obtain the distribution of target strong scattering points and perform recognition and classification. However, most of the traditional methods perform frame-by-frame modeling based on a fully connected structure, missing the inter-frame correlation information, unable to capture the structural information reflecting the HRRP characteristics, and at the same time having relatively high requirements for the completeness of target data. However, the recognition objects in the actual environment are usually non-cooperative targets, and the collected sample angular domain information is incomplete, and the data distribution cannot reach the ideal state of the experiment, thus increasing the dependence on the experience of researchers in the feature extraction process. In recent years, with the rise of deep learning algorithms, the traditional fully connected structure has been changed, and the deep features contained in HRRP data can be automatically obtained.

[0003] Deep learning has developed rapidly in recent years, and a large number of models have been proposed to improve the performance and efficiency of the network. However, most of the development environments are deployed on large server sides, which require a large amount of memory resources and strong computing capabilities. With the continuous improvement of the computing power of small mobile embedded devices, it is more practical to apply various deep learning methods to these small embedded development devices with strong real-time performance, high stability, and convenient portability in production and life. Compared with large-scale servers, the computing power, storage capacity, and power consumption of general small embedded development platforms are relatively limited. In order to have higher reliability on small embedded development devices, we usually need to optimize the adapted model methods, such as reducing the model complexity, reducing the number of network parameters and computational amounts, and reducing the memory occupation and power consumption. Summary of the Invention

[0004] Aiming at the deficiencies existing in the prior art, the present invention provides a radar target recognition method based on optimizing capsules.

[0005] We propose an efficient and lightweight deep learning model based on the SimSiam-CapsNet structure, and achieve the recognition of HRRP targets through a new routing mechanism. We integrate a network with strong generalization ability to fully capture the effective structural information in HRRP. While constructing a complex model, we also optimize the parameters in the network to achieve faster convergence and higher recognition performance. Based on the ability of CNN to obtain local structural information and the Inception structure, we propose a multi-scale CNN structure to capture the separable information at different levels of HRRP targets in detail from multiple perspectives, increasing the sparsity of the network. At the same time, we use the more robust CapsNet as a supplement to CNN to make up for the deficiency that CNN is difficult to reflect the structural information of the target HRRP sequence, and play a greater advantage in the field of radar HRRP recognition with stronger fitting ability. At the same time, we propose to optimize the quality of basic capsules and reduce the number of capsule parameters, which greatly reduces the number of trainable parameters in the network and improves the recognition performance to a certain extent, providing a feasible method for the actual radar recognition engineering application.

[0006] A radar target recognition method based on optimized capsules, comprising the following steps:

[0007] S1: Preprocess the original HRRP sample set.

[0008] Through l 2 Process the original HRRP echo by intensity normalization to improve the intensity sensitivity problem of HRRP. HRRP is intercepted from radar echo data through a distance window. During the interception process, the position of the recorded range image in the range gate is not fixed, resulting in the translational sensitivity of HRRP. To make the training and testing have a unified standard, the centroid alignment method is used to eliminate the translational sensitivity.

[0009] S2: Perform translational processing on the processed HRRP samples to achieve data augmentation.

[0010] S3: Input the HRRP samples after data augmentation into the SimSiam module for feature extraction.

[0011] S4: Input the high-dimensional features extracted by the SimSiam module into the convolutional module, which consists of a convolutional adjustment layer and an SE layer. Input the output features obtained after the convolutional adjustment layer into the lightweight SE layer to enhance the sensitivity of the network to the HRRP feature channels.

[0012] S5: Input the features extracted by the convolutional module into the basic capsule network based on the SE layer. Adopt the method of depthwise separable convolution for the effective channel features of the SE layer to construct effective basic capsules, and extract higher-quality features by combining the spatial features and channel feature information in the HRRP data. At the same time, a self-attention routing mechanism (Self-Attention Routing) is also used between the capsules to construct a non-iterative routing mechanism, capture the connection between local and global features, route the effective HRRP features from the low-level capsules to the high-level capsules corresponding to the HRRP target categories, obtain rich features and position dependence relationships in the low-level capsules while reducing the number of parameters of the capsules.

[0013] S6: Build a classifier to classify the HRRP targets, and use the attention mechanism in the classifier to retain more effective features.

[0014] S7: Design a loss function, initialize all the weights and biases to be trained in the SimSiam module, convolutional module, and capsule network, and set training parameters, including the learning rate, batch_size, and number of training batches, and then perform training.

[0015] Further, the detailed steps of S1 are as follows:

[0016] S1.1: Intensity normalization. Represent the original HRRP as where L 1 represents the total number of range cells included in the HRRP. Then the intensity-normalized HRRP can be expressed as:

[0017]

[0018] S1.2: Sample alignment. Translate the HRRP so that its centroid g 1 is moved to nearby, so that the range cells containing information in the HRRP will be distributed near the center. The calculation method of the centroid g 1 of the HRRP is as follows:

[0019]

[0020] Further, the detailed steps of S2 are as follows:

[0021] In order to avoid overfitting and obtain important semantic information in the HRRP data during the pre-training process, data augmentation is performed by translating the centroid of each HRRP sample after sensitivity processing 1 - 4 range cells to the left and right respectively. Then the number of samples available for unsupervised pre-training can increase by 8 times based on the previous training set, thereby improving the generalization ability of the network to new samples to a certain extent.

[0022] Further, the detailed steps of S3 are as follows:

[0023] S3.1: Input the HRRP samples obtained by different data augmentation methods into the SimSiam module composed of random data augmentation, backbone network, projector, and predictor. After random data augmentation, enter the encoder for encoding, perform feature matching by maximizing the consistency between the feature vectors of different views from the same HRRP sample, and then extract the high-quality main feature representation of the HRRP sample through the projector. Input the output features into the predictor to optimize the network parameters through backpropagation.

[0024] SimSiam working mechanism: Define the loss of SimSiam using the EM algorithm, and the expression is as follows:

[0025]

[0026] where represents the encoder network for feature extraction, θ is a learnable parameter, x is the HRRP sample, represents the random data augmentation function before the HRPP data input, and the expectation represents the distribution of the HRRP sample x and the random data augmentation method In other words, is equivalent to the sum of the loss expectations of all HRRP samples and random data augmentations. η x is the feature representation of the HRRP sample x, that is, the feature vector z output by the encoder i . Use the Mean Squared Error (MSE) to calculate the similarity. At this time, the working mode of SimSiam is similar to the K-means clustering algorithm, fixing one variable to solve the other variable, which is the EM iteration algorithm. Convert it into the following two sub-problems:

[0027]

[0028]

[0029] where ← represents the assignment operation, and r represents the number of algorithm iteration updates. Use the Stochastic Gradient Descent (SGD) algorithm to calculate the solution of θ r in the first sub-problem. Through formula (3-5-2), stop the backpropagation of the gradient to η r-1 , then η r-1 is a constant in formula (3-5-2). If the backpropagation of the gradient is not stopped, there are two variables in the formula, resulting in inability to solve.

[0030] Obtain θ r After obtaining the solution of θ, substitute it into the second sub-problem. At this time, there is only one variable η in formula (3-5-3), and we need to minimize the expectation of each HRRP sample x Substitute formula (3-5-1) into formula (3-5-3) again, then the solution of the second sub-problem is transformed into:

[0031]

[0032] According to the expectation formula transformation, we get:

[0033]

[0034] At this time, it means that the feature representation of a certain HRRP sample x during the r-th iteration update is obtained by the expectation of random data augmentation of the sample x

[0035] Perform a random data augmentation on the transformed second sub-problem according to formula (3-5-5) The formula is as follows:

[0036]

[0037] Substitute it into formula (3-5-2) again, and we get:

[0038]

[0039] where θ r is the solution of the equation of formula (3-5-2), and represent two different data augmentation methods acting on a certain HRRP sample, then this formula can be regarded as a twin two-tower architecture

[0040] Add a predictor to one side of the simsiam module branch and define it as h 1 , according to the expectation formula, transform formula (3-5-4) into:

[0041] h 1 (z 1 ) = E z [z 1 = E T [f(T(x))] (3-5-8)

[0042] Since it is difficult to directly calculate the expectation value of the random augmentation , for the convenience of analysis, the expectation after augmentation is equivalent to the expectation of itself

[0043] Furthermore, the detailed steps of S4 are as follows:

[0044] S4.1: The convolution adjustment layer of the first part of the convolution module contains three convolution groups, and each convolution group contains three parts: a convolution layer, a batch normalization layer, and a ReLU activation function. First, the first convolution group performs dimensional conversion on the high-dimensional features extracted by the SimSiam module, and then maps them to a higher-dimensional space more suitable for constructing capsules through the latter two convolution groups. At the same time, the local feature information in the HRRP features is fully extracted.

[0045] S4.2: The second part of the convolution module is the SE layer. The SE layer is used to strengthen the effective HRRP feature information contained in the channels, weaken the invalid features in the channels, and improve the HRRP feature representation ability of the network. Denote the three-dimensional feature map obtained through the convolution module calculation as F. For the input HRRP feature F, first use the squeeze operation to average map the global features of K channels into a scalar where k = 1, 2, …, K, and regard these values together as a vector X sq , then The calculation expression for the global response value of the k-th channel in the input HRRP feature is as follows:

[0046]

[0047] where l 1 represents the elements in each channel, and F(k, l 1 ) represents the l-th 1 element of the k-th channel in the feature. Then, pass the vector X sq through two fully connected layers and perform activation to learn the weight parameter s k of each channel to model the inter-channel dependence relationship of the HRRP features, which is obtained by the following calculation:

[0048]

[0049] where δ represents the activation function Sigmoid of the second fully connected layer FC2, whose role is to normalize the output channel weights and distribute them between 0 and 1, σ represents the activation function ReLU of the first fully connected layer FC1, and W FC1 and W FC2 are the weight matrices in these two fully connected layers respectively.

[0050] Finally, through the recalibration operation, multiply the channel weight s k by the three-dimensional feature map F after convolution calculation to obtain the feature F se , and the adjusted feature of the k-th channel is calculated by the following formula:

[0051]

[0052] where ⊙ represents multiplying the weight s k by each element of the corresponding channel. Therefore, the final output HRRP channel features total K channels.

[0053] Furthermore, the detailed steps of S5 are as follows:

[0054] S5.1: Depth convolutional capsule layer, using depthwise separable convolution operations on each channel of the feature F se to extract more effective features and transfer them to the base capsules, which helps the network achieve more accurate prediction and classification. The specific method of constructing the vector neurons of the base capsules is to merge every 8 channels of the HRRP features of K channels into a capsule s represented by a vector. The obtained group of base capsules is denoted as S n,d , where n represents the number of base capsules and d represents the dimension of each base capsule. The group of base capsules is the input of the routing layer and has various attributes of HRRP learned during the training process. After the capsules are constructed, the position information of the HRRP features is no longer "position encoding", but "rate encoding" in the attributes of the capsules. Therefore, the basic element of the capsule network is no longer a single neuron, but a capsule with vector output. To encode the relationship between the probability of the existence of a certain type of HRRP target and the length of the vector capsule, and let the high-level prediction capsules perform prediction and classification, a special Squash function is used to nonlinearly activate the base capsules, so that the vector maintains its original direction while shrinking its modulus length to between 0 and 1. The Squash compression operation is shown in the following formula:

[0055]

[0056] where s n represents the nth base capsule in S n,d , The new vector unit obtained after compression is denoted as U n,d , which contains n capsules u n,d , and has the same dimension and attributes as s n,d .

[0057] S5.2: Self-attention routing layer: The self-attention routing layer contains all the affine transformations embedded between two adjacent capsule layers. Each input capsule predicts the attributes of the next-layer capsules according to the transformation matrix. Using the self-attention routing layer between the capsules to capture the connection between local and global features, routing the effective HRRP features from the low-level capsules to the high-level capsules corresponding to the HRRP target classes, obtaining rich features and position dependence relationships in the low-level capsules while reducing the number of parameters of the capsules.

[0058] Furthermore, the detailed steps of S6 are as follows:

[0059] S6.1: After the self-attention routing layer, the output shape is [batch_size, M, N]. The classifier consists of a fully connected layer and a softmax layer. In the classifier, an attention mechanism is applied:

[0060]

[0061] is the weight parameter for the dimension M, and L(i 2 ) is the feature of each dimension. Different weights are learned according to the different importance levels of each dimension feature;

[0062] S6.2: Classification is performed through the softmax layer. If the total number of targets included in the training set is C, the probability that the test HRRP sample X test corresponds to the i-th target in the target set is expressed as:

[0063]

[0064] where exp(.) represents the exponential operation, and F s (i) refers to the i-th element in the vector F s , F s =W s F ATT , W s is the weight matrix of the vector F s , and F ATT is the feature vector output through the fully connected layer. The test HRRP sample X test is classified into the maximum target probability c 0 by the maximum a posteriori probability:

[0065]

[0066] Furthermore, the detailed steps of S7 are as follows:

[0067] S7.1: The loss function is designed as cross-entropy. The parameters are learned by calculating the gradient of the loss function with respect to the parameters using the training data, and the learned parameters are fixed when the model converges. The present invention adopts a cost function based on cross-entropy, which can be expressed as:

[0068]

[0069] where N 1 represents the number of training samples in a batch, is a one-hot vector, and n 2 is used to represent the n 2 th training sample, and P(i 3 |x train)Indicates the probability that the training sample corresponds to the i-th 3 target.

[0070] S7.2: Initialize all the weights and biases to be trained in the SimSiam module, convolutional module, and capsule network, and set the training parameters, including the learning rate, batch_size, and number of training batches, and then conduct the training.

[0071] The beneficial effects of the present invention are as follows:

[0072] 1. In the present invention, a lightweight SE layer is applied to enhance the sensitivity of the network to the HRRP feature channels. It not only does not increase excessive parameters and complexity but also can obtain the global feature representation of HRRP from the perspective of channel dependence relationships, improve the feature representation according to the importance levels of different channels, strengthen the effective features, and at the same time suppress the ineffective features that have no effect on the task. It can not only remove the redundant features in the HRRP samples but also help to construct higher-quality basic capsules and reduce the computational amount in the routing process.

[0073] 2. The present invention applies a depthwise separable convolutional network. In order to reduce the number of capsule parameters and make the model more lightweight for practical device implementation, the present invention adopts the method of depthwise separable convolution based on the effective channel features of the SE layer to construct effective basic capsules, combines the spatial and channel feature information in the HRRP samples to extract higher-quality features, not only increases the receptive field of each channel but also reduces the dimension of the input of the lower-layer capsules, extremely simplifies and reduces the number of parameters required for constructing the basic capsules, and improves the convergence efficiency of the model.

[0074] 3. The present invention applies a self-attention routing layer, which can be trained in parallel and shows good performance in many classification tasks. The main feature of self-attention is that it can learn the internal structure information of vectors, which can not only improve the efficiency through parallel computing but also improve the stability and accuracy. This mechanism weakens the irrelevant features based on feature analysis and highlights the significant features useful for the current task.

[0075] 4. Aiming at the high computational consumption of the vector capsule network, the present invention combines two methods of optimizing capsule construction and routing method to reduce the number of network parameters and computational overhead, providing convenience for the future engineering application of the model in the field of energy-saving computing. First, use depthwise separable convolution to replace the conventional CNN to construct basic capsules, which greatly reduces the number of capsule parameters and improves the quality of the capsules at the same time; finally, a self-attention routing algorithm mechanism is proposed to update the capsule transfer process, achieving a better classification effect. Description of the Drawings

[0076] Figure 1 is the step flow chart of the embodiment of the present invention;

[0077] Figure 2 Schematic diagram of the self-attention routing process in the embodiment of the present invention; Detailed implementation manners

[0078] Referring to Figure 1 , which is a flowchart of a radar high-resolution range profile recognition technology based on optimized capsules according to the present invention. The specific implementation steps are as follows:

[0079] S1: Preprocess the original HRRP sample set.

[0080] Since the intensity of HRRP is jointly determined by factors such as radar transmission power, target distance, radar antenna gain, and radar receiver gain, before using HRRP for target recognition, we use l 2 The method of intensity normalization is used to process the original HRRP echo, so as to improve the intensity sensitivity problem of HRRP. HRRP is intercepted from the radar echo data through a distance window. During the interception process, the position of the intercepted range profile in the range gate is not fixed, resulting in the translational sensitivity of HRRP. In order to make the training and testing have a unified standard, the centroid alignment method is used to eliminate the translational sensitivity.

[0081] S1.1: Intensity normalization. Represent the original HRRP as where L 1 represents the total number of range cells included in HRRP. Then the HRRP after intensity normalization can be expressed as:

[0082]

[0083] S1.2: Sample alignment. Translate the HRRP so that its centroid g 1 moves to near , so that the range cells containing information in HRRP will be distributed near the center. The calculation method of the centroid g 1 of HRRP is as follows:

[0084]

[0085] After the original HRRP sample is processed by intensity normalization and the centroid alignment method, the amplitude has been limited between 0 and 1, which not only unifies the scale, but also the values between 0 and 1 are very beneficial for subsequent neural network processing; the HRRP echo signals with a distribution biased to the right or left are all adjusted to near the center point.

[0086] S2: Perform translational processing on the processed HRRP samples to achieve data augmentation, provide a large number of rich and diverse training samples for contrastive learning unsupervised pre-training, and thus improve the generalization ability of the network to new samples to a certain extent.

[0087] To avoid overfitting during the pre-training process and obtain important semantic information in HRRP data, we perform data augmentation by shifting the center of gravity of each sensitivity-processed HRRP sample 1-4 distance units to the left and right respectively. At this time, the number of samples available for unsupervised pre-training can be increased by 8 times based on the previous training set, thereby improving the generalization ability of the network to new samples to a certain extent.

[0088] S3: Input the HRRP samples after data augmentation into the SimSiam module for feature extraction.

[0089] S3.1: Input the HRRP samples obtained by different data augmentation methods into the SimSiam module composed of random data augmentation, backbone network, projector, and predictor. After random data augmentation processing, it enters the encoder for encoding, performs feature matching by maximizing the consistency between the feature vectors of different views from the same HRRP sample, and then extracts the high-quality main feature representation of the HRRP sample through the projector. The output features are input into the predictor to optimize the network parameters through backpropagation.

[0090] SimSiam working mechanism: Use the EM algorithm to define the loss of SimSiam, and the expression is as follows:

[0091]

[0092] where represents the encoder network for feature extraction, θ is a learnable parameter, x is the HRRP sample, represents the random data augmentation function before the HRPP data input, and the expectation represents the distribution of the HRRP sample x and the random data augmentation method In other words, is equivalent to the sum of the loss expectations of all HRRP samples and random data augmentations. η x is the feature representation of the HRRP sample x, that is, the feature vector z output by the encoder i . The mean squared error (MSE) is used to calculate the similarity. At this time, the working method of SimSiam is similar to the K-means clustering algorithm, fixing one variable and solving the other variable, which is the EM iteration algorithm. It is converted into the following two sub-problems:

[0093]

[0094]

[0095] where ← represents the assignment operation, and r represents the number of iterations of the algorithm update. The Stochastic Gradient Descent (SGD) algorithm is used to calculate θ in the first sub-problem r of the solution, through formula (3-5-2), stop the backpropagation of the gradient to η r-1 , then η r-1 in formula (3-5-2) is a constant. If the backpropagation of the gradient is not stopped, there are two variables in the formula, resulting in an unsolvable situation. Therefore, the stop-grad operation in SimSiam is reasonably explained

[0096] After obtaining the solution of θ r , substitute it into the second sub-problem. At this time, there is only one variable η in formula (3-5-3). We need to minimize the expectation of each HRRP sample x Substitute formula (3-5-1) into formula (3-5-3) again, then the solution of the second sub-problem is transformed into

[0097]

[0098] According to the expectation formula transformation, we get

[0099]

[0100] At this time, it represents that the feature representation of a certain HRRP sample x at the r-th iteration update is obtained by the expectation of the sample x through random data augmentation

[0101] Perform a random data augmentation on the transformed second sub-problem according to formula (3-5-5) The formula is as follows

[0102]

[0103] Substitute it into formula (3-5-2) again, and we get

[0104]

[0105] where θ r is the solution of the equation in formula (3-5-2), and represent two different data augmentation methods acting on a certain HRRP sample, then this formula can be regarded as a twin two-tower architecture

[0106] Add a predictor to one side of the simsiam module branch and define it as h 1 , according to the expectation formula, transform formula (3-5-4) into

[0107] h 1 (z 1 ) = E z [z 1 = E T [f(T(x))] (3-5-8)

[0108] Since it is difficult to directly calculate the expected value of random augmentation for the sake of easy analysis, the expected value after augmentation is made equivalent to its own expected value. At this time, the existence of the predictor can compensate for the difficult-to-directly-relate gap between the feature representation z and the expected value. Because the sampling using the random augmentation method in several different epochs conforms to a stable uniform distribution, such a distribution is relatively easy for the network to learn and remember, so that the expected value can be predicted through the representation z 1 and the expected value. Because the sampling using the random augmentation method in several different epochs conforms to a stable uniform distribution, such a distribution is relatively easy for the network to learn and remember, so that the expected value can be predicted through the representation z 1 to predict the expected value.

[0109] S4: Input the high-dimensional features extracted by the SimSiam module into the convolutional module, which is composed of a convolutional adjustment layer and an SE layer. Input the output features obtained after passing through the convolutional adjustment layer into the lightweight SE layer to enhance the network's sensitivity to the HRRP feature channels, without adding too many parameters and complexity and being able to obtain the global feature representation of HRRP from the perspective of channel dependence relationship, thereby improving the network's extraction performance.

[0110] S4.1: The convolutional adjustment layer in the first part of the convolutional module contains three convolutional groups, and each convolutional group contains three parts: a convolutional layer, a batch normalization layer, and a ReLU activation function. First, the first convolutional group performs dimensional conversion on the high-dimensional features extracted by the SimSiam module, and then maps them to a higher-dimensional space more suitable for constructing capsules through the latter two convolutional groups, while also fully extracting the local feature information in the HRRP features.

[0111] S4.2: The second part of the convolutional module is the SE layer, which is used to strengthen the effective HRRP feature information contained in the channels and weaken the invalid features in the channels, improving the network's HRRP feature representation ability. Denote the three-dimensional feature map obtained after calculation by the convolutional module as F. For the input HRRP feature F, first use the squeeze operation to respectively map the global features of K channels to a scalar where k = 1, 2,... K, and regard these values as a vector X sq , then The calculation expression of the global response value of the k-th channel in the input HRRP feature is as follows:

[0112]

[0113] where l 1Denote the elements in each channel, F(k, l 1 ) represents the l-th element in the k-th channel of the feature 1 . Then, the vector X sq passes through two fully-connected layers and is activated to learn the weight parameter s k for each channel to model the inter-channel dependencies of HRRP features, which is calculated as follows:

[0114]

[0115] where δ represents the activation function Sigmoid of the second fully-connected layer FC2, whose role is to normalize the output channel weights to be distributed between 0 and 1, σ represents the activation function ReLU of the first fully-connected layer FC1, and W FC1 and W FC2 are the weight matrices in these two fully-connected layers respectively.

[0116] Finally, through a recalibration operation, the channel weight s k is multiplied by the three-dimensional feature map F after convolution calculation to obtain the feature F se , and the adjusted feature of the k-th channel is calculated by the following formula:

[0117]

[0118] where ⊙ represents multiplying the weight s k by each element of the corresponding channel. Therefore, the finally output HRRP channel features have a total of K channels.

[0119] S5: Input the features extracted by the convolution module into the basic capsule network based on the SE layer. The effective channel features based on the SE layer adopt the method of depthwise separable convolution to construct effective basic capsules, combine the spatial features and channel feature information in the HRRP data to extract higher-quality features, increase the receptive field of each channel, reduce the dimension of the input of the low-level capsules, simplify and reduce the number of parameters required to construct the basic capsules, and improve the convergence efficiency of the model. At the same time, a self-attention routing mechanism (Self-Attention Routing) is also used between the capsules to construct a non-iterative routing mechanism, capture the connection between local and global features, route the HRRP effective features from the low-level capsules to the high-level capsules corresponding to the HRRP target categories, obtain the rich features and position dependence relationships in the low-level capsules while reducing the number of parameters of the capsules.

[0120] S5.1: Depth convolutional capsule layer, using the feature F seDepthwise separable convolution operations are performed on each channel to extract more effective features and pass them to the basic capsules, which helps the network achieve more accurate prediction and classification. The specific method for constructing the vector neurons of the basic capsules is to combine every 8 channels out of the HRRP features of K channels into a capsule s represented by a vector. The resulting group of basic capsules is denoted as S n,d , where n represents the number of basic capsules and d represents the dimension of each basic capsule. The group of basic capsules serves as the input to the routing layer and possesses various attributes of HRRP learned during the training process. After the capsules are constructed, the position information of the HRRP features is no longer "position encoding" but "rate encoding" in the attributes of the capsules. Therefore, the basic element of the capsule network is no longer a single neuron but a capsule with vector output. To encode the relationship between the probability of the existence of a certain type of HRRP target and the length of the vector capsule and enable the high-level prediction capsules to perform prediction and classification, a special Squash function is used to nonlinearly activate the basic capsules, keeping the vector in its original direction while shrinking its modulus length to between 0 and 1. The Squash compression operation is shown in the following formula:

[0121]

[0122] where s n represents the nth basic capsule in S n,d , and the new vector unit obtained after compression is denoted as U n,d , which contains n capsules u n,d and has the same dimension and attributes as s n,d . The non-linear activation operation ensures that short capsule vectors are compressed to a length almost equal to 0, while long vectors are compressed to a length slightly less than 1.

[0123] S5.2: Self-attention routing layer: The self-attention routing layer includes all affine transformations embedded between two adjacent capsule layers. Each input capsule predicts the attributes of the capsules in the next layer according to the transformation matrix. In the present invention, a self-attention routing layer is used between the capsules to capture the connection between local and global features, route the effective HRRP features from the low-level capsules to the high-level capsules corresponding to the HRRP target categories, obtain rich features and position-dependent relationships in the low-level capsules while reducing the number of parameters of the capsules. Figure 2 is a schematic diagram of the self-attention routing process in an embodiment of the present invention;

[0124] S6: Build a classifier to classify HRRP targets. Use the attention mechanism in the classifier to retain more effective features.

[0125] S6.1: After the self-attention routing layer, the output shape is [batch_size, M, N]. The classifier consists of a fully connected layer and a softmax layer. In the classifier, an attention mechanism is applied:

[0126]

[0127] is the weight parameter for the dimension M, and L(i 2 ) is the feature of each dimension. Different weights are learned according to the different importance degrees of each dimension feature;

[0128] S6.2: Classification is performed through the softmax layer. If the total number of targets included in the training set is C, the probability that the test HRRP sample X test corresponds to the i-th target in the target set is expressed as:

[0129]

[0130] where exp(.) represents the exponential operation, and F s (i) refers to the i-th element in the vector F s , F s = W s F ATT , W s is the weight matrix of the vector F s , and F ATT is the feature vector output through the fully connected layer. The test HRRP sample X test is classified into the maximum target probability c 0 by the maximum a posteriori probability:

[0131]

[0132] S7: Design a loss function, initialize all the weights and biases to be trained in the SimSiam module, the convolutional module, and the capsule network, and set the training parameters, including the learning rate, batch_size, and number of training batches, and then perform training.

[0133] S7.1: The loss function is designed as cross-entropy. The parameters are learned by calculating the gradient of the loss function with respect to the parameters using the training data, and the learned parameters are fixed when the model converges. The cost function based on cross-entropy adopted in the present invention can be expressed as:

[0134]

[0135] where, N 1 represents the number of training samples in a batch, is a one-hot vector, and n 2 is used to represent the n2 training samples, P(i 3 |x train ) represents the probability that the training sample corresponds to the i-th 3 target.

[0136] S7.2: Initialize all the weights and biases to be trained in the SimSiam module, convolutional module, and capsule network, and set the training parameters, including the learning rate, batch_size, and number of training batches, and then conduct the training.

[0137] Embodiment

[0138] Training phase:

[0139] S1: Collect the data set and preprocess the samples. The specific operation steps are as follows:

[0140] S1.1: Collect the data set. Merge the HRRP data set collected by the radar according to the types of targets. For each type of sample, select the training samples and test samples in different data segments. During the selection process of the training set and test set, ensure that the poses of the training set samples selected are covered by the poses of the test set samples with respect to the radar. The ratio of the number of training set samples to test set samples for each type of target is 7:3.

[0141] S1.2: Intensity normalization. Represent the original HRRP as where L 1 represents the total number of range cells included in the HRRP. Then the intensity-normalized HRRP can be represented as:

[0142]

[0143] S1.3: Sample alignment. Translate the HRRP so that its center of gravity g 1 moves to near such that the range cells containing information in the HRRP will be distributed near the center. The calculation method of the center of gravity g 1 of the HRRP is as follows:

[0144]

[0145] S2: Conduct data augmentation on the samples. The specific steps are as follows:

[0146] S2.1: Translate the center of gravity of each HRRP sample after sensitivity processing 1 - 4 range cells to the left and right respectively for data augmentation. In this way, the number of samples available for unsupervised pre-training can be increased by 8 times on the basis of the previous training set, thereby improving the generalization ability of the network to new samples to a certain extent.

[0147] S3: Input the data-augmented samples into the SimSiam module, and the steps are as follows:

[0148] Input the data-augmented HRRP samples into the SimSiam module. First, enter the encoder network for encoding. Feature matching is performed by maximizing the consistency between the feature vectors of different views from the same HRRP sample. Then, use the projector to extract the high-quality main feature representations of the HRRP samples, which helps the contrast prediction task to obtain a general consistency representation. Input the output features into the predictor to estimate the overall expectation of the SimSiam module, that is, optimize the network parameters through backpropagation, and at the same time filter out some invalid feature information in the high-level part, so that the feature information in the HRRP samples is fully retained for contrast prediction optimization.

[0149] S4: Input the output features of the SimSiam module into the convolution module, and the specific steps are as follows:

[0150] S4.1: Each convolution group in the convolution adjustment layer of the convolution module contains three processes: the convolution layer, the batch normalization layer, and the ReLU activation function. The input X passes through N 2 convolution kernels with a size of (1, 5) to obtain the output N 2 represents the total number of channels, and i 4 represents the i 4 th channel. Although the convolution kernel sizes are the same, the weight initializations are different, so these N 2 channels are also different, and different local features are extracted:

[0151]

[0152] The convolved data needs to be further processed. To make the model easier to converge and the network training process more stable, batch normalization is added after convolution. By calculating the mean and variance of the data in each mini-batch, assuming there are N m HRRP samples in a small batch, then the output is defined as where F n represents the convolution output corresponding to the nth HRRP sample. In each small batch, perform batch normalization on the HRRP data in to obtain which is expressed as:

[0153]

[0154] where, F n (k, l) represents the lth element in the kth channel of the convolution layer output corresponding to the HRRP sample before batch normalization, Namely, the HRRP data after batch normalization, α k and β k are trainable parameters corresponding to the k-th channel. ε is a very small number to prevent division by zero. E(.) represents the mean operation, and Var(.) represents the variance operation;

[0155] After that, the activation function ReLU is used to perform non-linear activation on each element in If the input is then the corresponding output after ReLU is expressed as:

[0156]

[0157] S5: Input the output result of the previous step into the capsule network. The specific steps are as follows:

[0158] Utilize the depthwise separable convolution operation acting on the feature F se for each channel to extract more effective features and transfer them to the basic capsules, which helps the network achieve more accurate prediction and classification. The specific method for constructing the vector neurons of the basic capsules is to combine every 8 channels out of the HRRP features of K channels into a capsule s represented by a vector. Denote the obtained group of basic capsules as S n,d , where n represents the number of basic capsules and d represents the dimension of each basic capsule. The basic capsule units serve as the input to the self-attention routing layer and possess various attributes of HRRP learned during the training process.

[0159] After the capsules are constructed, the position information of the HRRP features is no longer "position encoding", but "rate encoding" in the attributes of the capsules. Therefore, the basic elements of the capsule network are no longer single neurons, but capsules with vector outputs. To encode the relationship between the probability of the existence of a certain type of HRRP target and the length of the vector capsule, and to enable the high-level prediction capsules to predict and classify the parameters, a special Squash function is used to perform non-linear activation on the basic capsules, keeping the vector in its original direction while shrinking its modulus to between 0 and 1. The Squash compression operation is shown in the following formula:

[0160]

[0161] where s n represents the n-th basic capsule in S n,d , Denote the new vector unit obtained after compression as U n,d , which contains n capsules u n,d , and is related to s n,d have the same dimensions and attributes. The non-linear activation operation ensures that short capsule vectors are compressed to a length close to 0, while vectors with longer lengths are compressed to a length slightly below 1.

[0162] S6: Build the self-attention routing layer, and the specific steps are as follows:

[0163] S6.1: The self-attention routing layer contains all the affine transformations embedded between two adjacent capsule layers. Each input capsule predicts the attributes of the capsules in the next layer according to the transformation matrix.

[0164] S6.2: Concatenate the output features passing through the self-attention routing layer, and then connect a fully-connected layer with the number of nodes equal to the number of radar categories. That is, the output of the fully-connected layer is the prediction result of the model. Finally, obtain the probability through the softmax function, and the output can be expressed as:

[0165] output = f(C(c ATT )W o )

[0166] where C(·) is the concatenation operation, c represents the number of categories, and f(·) represents the softmax function.

[0167] S7: Design the loss function and train the model, and the specific steps are as follows:

[0168] S7.1: The loss function is designed as cross-entropy.

[0169] Learn the parameters by calculating the gradient of the loss function with respect to the parameters using the training data, and fix the learned parameters when the model converges. The loss function based on cross-entropy is adopted in the present invention and can be expressed as:

[0170]

[0171] where, N 1 represents the number of training samples in a batch, is a one-hot vector, n 2 is used to represent the true label of the nth 2 training sample, and P(i 3 |x train ) represents the probability that the training sample corresponds to the ith 3 target.

[0172] S7.2: Initialize all the weights and biases to be trained in the above model, set the training parameters, including the learning rate, batch_size, and number of training batches, and start training the model.

[0173] Testing phase:

[0174] S8: Perform the preprocessing operations of steps S1 and S2 in the training phase on the test data collected by S1.

[0175] S9: Send the samples processed by S8 into the model constructed and trained by S3, S4, S5, S6, and S7 for testing to obtain the results, that is, the output after the attention mechanism will be classified through the softmax layer. The HRRP test sample x test The probability corresponding to the k-th type of radar target in the target set can be calculated as:

[0176]

[0177] where exp(·) represents the exponential operation and c represents the number of categories.

[0178] We use the maximum a posteriori probability to classify the test HRRP sample x test into the k with the maximum target probability 0 as follows:

[0179]

[0180] After the above 9 steps, a radar high-resolution range profile recognition method based on a deep neural network and an attention mechanism proposed by the present invention can be obtained.

Claims

1. A radar target recognition method based on optimized capsules, characterized in that, it includes the following steps: S1: Preprocess the original HRRP sample set; By l 2 The original HRRP echo is processed by a method of intensity normalization to improve the intensity sensitivity problem of HRRP; HRRP is intercepted from radar echo data through a range window, and the position of the intercepted range image in the range gate is not fixed during the interception process, resulting in the translational sensitivity of HRRP; in order to make the training and testing have a unified standard, the centroid alignment method is used to eliminate the translational sensitivity; S2: Perform translation processing on the processed HRRP samples to achieve data augmentation; S3: Input the HRRP samples after data augmentation into the SimSiam module for feature extraction; S4: Input the high-dimensional features extracted by the SimSiam module into the convolution module, which consists of a convolution adjustment layer and an SE layer; input the output features obtained after the convolution adjustment layer into the lightweight SE layer to enhance the network's sensitivity to the HRRP feature channels; S5: Input the features extracted by the convolution module into the basic capsule network based on the SE layer. The effective channel features based on the SE layer adopt the method of depthwise separable convolution to construct effective basic capsules, and combine the spatial and channel feature information in the HRRP data to extract higher-quality features; at the same time, a self-attention routing mechanism is also used between the capsules to construct a non-iterative routing mechanism, capture the connection between local and global features, route the HRRP effective features from the low-level capsules to the high-level capsules corresponding to the HRRP target categories, obtain the rich features and position dependence relationships in the low-level capsules while reducing the number of parameters of the capsules; S6: Build a classifier to classify HRRP targets, and use an attention mechanism in the classifier to retain more effective features; S7: Design a loss function, initialize all the weights and biases to be trained in the SimSiam module, convolution module, and capsule network, and set training parameters, including learning rate, batch_size, and number of training batches, and perform training.

2. The radar target recognition method based on optimized capsules according to claim 1, characterized in that, the detailed steps of S1 are: S1.1: Intensity normalization; Represent the original HRRP as where L 1 represents the total number of range cells included in the HRRP, then the HRRP after intensity normalization is represented as: S1.2: Sample alignment; translate the HRRP to move its center of gravity \(g\) 1 to nearby, so that the range cells containing information in the HRRP will be distributed near the center; where the calculation method of the HRRP center of gravity \(g\) 1 is as follows:

3. The radar target recognition method based on optimized capsules according to claim 2, characterized in that, the detailed steps of S2 are: In order to avoid overfitting during the pre-training process and obtain important semantic information in the HRRP data, data augmentation is performed by translating the center of gravity of each HRRP sample after sensitivity processing 1-4 distance units to the left and right respectively. Then, the samples available for unsupervised pre-training can increase the data volume by 8 times on the basis of the previous training set, thereby improving the network's generalization ability to new samples to a certain extent.

4. The radar target recognition method based on optimized capsules according to claim 3, characterized in that, the detailed steps of S3 are: S3.1: Input the HRRP samples obtained by different data augmentation methods into the SimSiam module composed of random data augmentation, backbone network, projector, and predictor. After passing through the random data augmentation process, enter the encoder for encoding, perform feature matching by maximizing the consistency between the feature vectors of different views of the same HRRP sample, and then extract the high-quality main feature representation of the HRRP sample through the projector; input the output features into the predictor to optimize the network parameters through backpropagation; Working mechanism of SimSiam: Define the loss of SimSiam using the EM algorithm, and the expression is as follows: Among them represents the encoder network for feature extraction, θ is a learnable parameter, and x is the HRRP sample represents the random data augmentation function before the HRPP data input, and the expectation represents the distribution with respect to the HRRP sample x and the random data augmentation method In other words is equivalent to the sum of the loss expectations of all HRRP samples and random data augmentations; η x is the feature representation of the HRRP sample x, that is, the feature vector z output by the encoder i ; The mean squared error MSE is used to calculate the similarity. At this time, the working mode of SimSiam is similar to the K-means clustering algorithm, fixing one variable and solving the other variable, which is the EM iteration algorithm; it is converted into the following two sub-problems where ← represents the assignment operation, and r represents the number of times the algorithm iteratively updates; the solution of θ in the first sub-problem is calculated using the stochastic gradient descent algorithm, and the backpropagation of the gradient is stopped at η through formula (3-5-2). r In formula (3-5-2), η becomes a constant. If the backpropagation of the gradient is not stopped, there are two variables in the formula, making it impossible to solve. r-1 Then η r-1 in formula (3-5-2) is a constant. If the backpropagation of the gradient is not stopped, there are two variables in the formula, making it impossible to solve. Obtain θ r After obtaining the solution of, substitute it into the second sub-problem. At this time, there is only one variable η in formula (3-5-3), and we need to minimize the expectation of each HRRP sample x Substitute formula (3-5-1) into formula (3-5-3) again, then the solution of the second sub-problem is converted to: Obtained by transforming according to the expectation formula: At this time, it represents that the feature representation of a certain HRRP sample x during the r-th iteration update is obtained by the expectation of the sample x through random data augmentation; Perform a random data augmentation on the transformed second sub-problem according to formula (3-5-5). The formula is as follows: Substitute it into formula (3-5-2) again to get: where θ r is the solution of the equation in formula (3-5-2), and represent two different data augmentation methods acting on a certain HRRP sample, then formula (3-5-7) can be regarded as a twin two-tower architecture; Add a predictor to one side of the simsiam module branch and define it as h 1 , according to the expected formula, transform formula (3-5-4) into: h 1 (z 1 ) = E z [z 1 = E T [f(T(x))](3 - 5 - 8) Since it is difficult to directly calculate the expected value of the random augmentation it is difficult, for the sake of easy analysis, to equate the expectation after augmentation to its own expectation.

5. A radar target recognition method based on an optimized capsule according to claim 4, characterized in that The detailed steps of S4 are as follows: S4.1: The convolution adjustment layer in the first part of the convolution module contains three convolution groups, and each convolution group contains three parts: a convolution layer, a batch normalization layer, and a ReLU activation function; first, the first convolution group performs dimensional transformation on the high-dimensional features extracted by the SimSiam module, and then maps them to a higher-dimensional space more suitable for constructing capsules through the latter two convolution groups, while also fully extracting the local feature information in the HRRP features; S4.2: The second part of the convolutional module is the SE layer. The SE layer is used to strengthen the effective HRRP feature information contained in the channels, weaken the invalid features in the channels, and improve the HRRP feature representation ability of the network. Denote the three-dimensional feature map obtained by the calculation of the convolutional module as F. For the input HRRP feature F, first use the squeeze operation to average the global features of K channels into a scalar respectively where k = 1, 2, …, K, and regard these values together as a vector X sq , then The calculation expression of the global response value of the k-th channel in the input HRRP feature is as follows: where l 1 represents the elements in each channel, and F(k, l 1 ) represents the l 1 -th element in the k-th channel of the feature; then the vector X sq is passed through two fully connected layers and activated to learn the weight parameter s k for each channel to model the inter-channel dependencies of HRRP features, which is obtained by the following calculation: where δ represents the Sigmoid activation function of the second fully connected layer FC2, whose role is to normalize the output channel weights and distribute them between 0 and 1, σ represents the ReLU activation function of the first fully connected layer FC1, and W FC1 and W FC2 are the weight matrices in these two fully connected layers respectively; Finally, the channel weight s is multiplied by the three-dimensional feature map F after convolution calculation to obtain the feature F k The adjusted feature of the k-th channel is calculated by the following formula: se ​ Among them, ⊙ means multiplying the weight s k by each element of the corresponding channel; thus, the final output HRRP channel features have a total of K channels.

6. A radar target recognition method based on an optimized capsule according to claim 5, characterized in that The detailed steps of S5 are as follows: S5.1: Deep convolutional capsule layer, using the feature F se Depthwise separable convolutional operations for each channel are used to extract more effective features and pass them to the primary capsules, which helps the network achieve more accurate prediction and classification; the specific method for constructing the vector neurons of the primary capsules is to merge every 8 channels out of the HRRP features of K channels into a capsule s represented by a vector, and the resulting set of primary capsules is denoted as S n,d , where n represents the number of primary capsules and d represents the dimension of each primary capsule; the set of primary capsules is the input to the routing layer and has various attributes of HRRP learned during the training process; after the capsules are constructed, the position information of the HRRP features is no longer "position encoding", but "rate encoding" in the attributes of the capsules; therefore, the basic element of the capsule network is no longer a single neuron, but a capsule with vector output; in order to encode the relationship between the probability of the existence of a certain type of HRRP target and the length of the vector capsule, and let the higher-level prediction capsules perform prediction and classification, a special Squash function is used to nonlinearly activate the primary capsules, keeping the vector in the original direction while shrinking its magnitude to between 0 and 1; the Squash compression operation is shown in the following formula: where s n represents the nth basic capsule in S n,d , Denote the new vector unit obtained after compression as U n,d , which contains n capsules u n,d , and has the same dimension and attributes as s n,d ; S5.2: Self-attention routing layer: The self-attention routing layer contains all the affine transformations embedded between two adjacent capsule layers; each input capsule predicts the attributes of the next-layer capsules according to the transformation matrix; the self-attention routing layer is used between the capsules to capture the connection between local and global features, route the effective HRRP features from the low-layer capsules to the high-layer capsules corresponding to the HRRP target classes, obtain the rich features and position dependence relationships in the low-layer capsules while reducing the number of parameters of the capsules.

7. A radar target recognition method based on an optimized capsule according to claim 6, characterized in that The detailed steps of S6 are as follows: S6.1: After the self-attention routing layer, the output shape is [batch_size, M, N], and the classifier consists of a fully connected layer and a softmax layer; in the classifier, an attention mechanism is applied: α i2 is the weight parameter for the dimension M, and L(i 2 ) is the feature of each dimension. Different weights are learned according to the different importance levels of the features of each dimension; S6.2: Classification is performed through the softmax layer. If the total number of targets included in the training set is C, and the test HRRP sample X test The probability corresponding to the i-th target in the target set is expressed as: where exp(.) represents the exponential operation, F s (i) refers to the i-th element of the vector F s in, F s = W s F ATT , W s is the weight matrix of the vector F s ; F ATT is the feature vector output by the fully connected layer; the test HRRP sample X test is classified into the maximum target probability c 0 by maximum a posteriori probability as follows:

8. A radar target recognition method based on an optimized capsule according to claim 7, characterized in that The detailed steps of S7 are as follows: S7.1: The loss function is designed as cross-entropy; the parameters are learned by calculating the gradient of the loss function with respect to the parameters using the training data, and the learned parameters are fixed when the model converges; the present invention adopts a cost function based on cross-entropy, expressed as: Among them, N 1 represents the number of training samples in a batch, is a one - hot vector, and n 2 is used to represent the n 2 th training sample. P(i 3 |x train ) represents the probability that the training sample corresponds to the i 3 th target; S7.2: Initialize all the weights and biases to be trained in the SimSiam module, the convolution module, and the capsule network, set the training parameters, including the learning rate, batch_size, and number of training batches, and perform training.

Citation Information

Patent Citations

  • Radar HRRP target recognition method based on multi-scale convolutional neural network

    CN111580058A

  • Radar target identification method based on convolutional neural network and Bert

    CN112764024A