A Small-Sample SAR Target Recognition Method Based on Siamese Neural Network
By using a combination of twin neural networks and DenseNet in SAR target recognition, using Inception and jump connections, the problem of poor recognition performance of small-sample SAR targets is solved, and higher recognition accuracy and generalization capabilities are achieved.
Patent Information
- Application Number
- CN202310942393.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-30
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2043-07-30
AI Technical Summary
Existing SAR target recognition methods perform poorly under small sample data conditions, making it difficult to effectively identify SAR targets, especially when image recognition is recognized in small data sets, single categories and complex scenarios.
A small sample SAR target recognition method based on twin neural networks is used, and two symmetric branches are used, each branch uses DenseNet as the backbone to extract the network, and uses Inception to replace the conventional convolutional layer in the backbone network, adding jump connections to enhance feature utilization capabilities.
By utilizing features and jump connections at different scales, the gradient disappearance problem in small sample learning is alleviated, the generalization ability and recognition accuracy of the model are improved, and the SAR target recognition performance under small sample data is effectively improved.
Smart Images

Figure CN117173556B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to radar technology, and in particular to a small-sample SAR target recognition method based on a twin neural network. Background Art
[0002] Synthetic Aperture Radar (SAR) is an all-weather, all-day radar with high resolution and high penetrability. SAR plays an important role in both military and civilian fields. With its observation advantage of ground object penetration ability, SAR has achieved rapid development in recent decades and is widely used in fields such as agriculture and forestry, hydrological surveys, disaster warnings, resource exploration, and marine monitoring.
[0003] Traditional SAR target recognition methods rely on a large number of manually designed features. In addition, traditional methods usually have a large amount of computation and poor generalization performance for new categories. With the rapid development of machine learning technology, Convolutional Neural Network (CNN) has become a very popular and indispensable architecture in SAR target recognition and is largely superior to traditional methods.
[0004] In recent years, CNN has become the mainstream feature extraction method in SAR automatic target recognition. Its variants have greatly improved the recognition performance by methods such as deleting fully connected layers, constructing multi-channel structures, adopting batch normalization, adding information recorders, constructing cascade networks, and introducing attention mechanisms. At the same time, multi-feature fusion networks have been further designed using important features such as geometry, space, and time. However, due to the simplicity of the network structure and the deformation of the input data resulting in the inability to extract geometric structures, the feature representation ability of these methods is limited. Traditionally, deep CNNs usually require a large amount of training data to avoid overfitting, but due to the limitations of observation conditions and the high cost and long time of manual annotation, the training data is insufficient to achieve the automatic recognition of SAR targets. Since the data acquisition of SAR is more difficult than natural scene images and the annotation recognition of SAR images is very time-consuming and laborious, obtaining a large-scale SAR target sample and annotation recognition for SAR remains a challenge.
[0005] Most of the existing SAR target recognition methods apply the target recognition methods of natural images to SAR images, ignoring the characteristics of SAR data itself, that is, SAR images are essentially complex-valued images containing both amplitude and phase information. When traditional feature extraction methods are used to recognize SAR targets, due to the fact that SAR images (or feature vectors) are very sensitive to the change of the target's attitude angle, in addition, changes in the structure of the target itself, occlusion, concealment, and changes in the background and parameters, etc., will all cause changes in the target SAR image or SAR feature vector. Therefore, the recognition effect may not be ideal. Since the acquisition and sample labeling of SAR data may be very expensive and laborious, and a large number of samples cannot be obtained, some target categories may only have a few or dozens of labeled samples. The heavy dependence of deep learning on large datasets limits the application of deep learning in the field of synthetic aperture radar (SAR) automatic target recognition (ATR) where the target sample set is generally small. In this case, the performance of CNN drops significantly, and the problem of small-sample SAR target recognition appears. Summary of the Invention
[0006] The main object of the present invention is to provide a small-sample SAR target recognition method based on a siamese neural network. This method includes two symmetric branches, and each branch uses DenseNet as the backbone extraction network. Aiming at the problem of gradient disappearance that occurs when the DenseNet network model performs small-sample learning, Inception is used in the backbone network to replace the conventional convolutional layer, which not only strengthens feature propagation but also can utilize features of different scales. Skip connections are added to the network to enhance the network's ability to utilize features of different scales in the final prediction. This network structure can reduce the number of learning parameters in the model and alleviate the overfitting problem at the same time. Then, similarity calculation is performed on the two inputs after feature extraction to determine whether the two inputs belong to the same category. Finally, the sigmoid function is used to compress the output into the interval [0, 1].
[0007] The technical solution adopted by the present invention is: a small-sample SAR target recognition method based on a siamese neural network, including two symmetric branches, and each branch uses DenseNet as the backbone extraction network, and Inception is used in the backbone network to replace the conventional convolutional layer;
[0008] Each branch contains 3 DenseBlocks, and each DenseBlock contains 4 Inception units. As Figure 2 shown, each Inception unit outputs 12 channels of feature maps; the ratio of the output numbers of feature map channels of different convolutional kernels is 1:3:8.
[0009] Further, where Hl represents the output of the l-th layer in the DenseBlock, the output Hl+1 of the (l + 1)-th layer can be expressed as formula (1)
[0010] Hl+1 = [[f7(Hl), f5(Hl), f3(Hl)], Hl-1, ..., H1] (1)
[0012] fn() represents a convolutional layer with a convolutional kernel size of n, and [] represents the concatenation operation on the channels; the convolutional layer consists of four basic operations, batch normalization, bias-free convolution, ReLU activation function, and dropout with a probability of 0.2;
[0013] The Inception unit consists of several convolutional layers. Its first 1×1 convolution is introduced as a bottleneck layer, which generates a 4-fold increase in the feature map dimension;
[0014] Then, convolutional layers with different kernel sizes are used to process the feature map;
[0015] The feature maps generated by the first DenseBlock and the second DenseBlock are added to the final fully connected layer through a 1×1 convolutional layer and global average pooling.
[0016] Even further, the feature extraction module is mainly composed of 3 DenseBlock modules;
[0017] After the first two DenseBlock blocks, a transition layer is connected. Through the transition layer between adjacent DenseBlocks, the grid size and the number of channels of the feature map are reduced to half of the original size;
[0018] The transition layer mainly includes a 1×1 convolution and average pooling with a stride of 2.
[0019] Even further, during training, the binary cross-entropy loss function is used for model training to measure the similarity between two input samples;
[0020] Binary cross-entropy is a commonly used Loss function. p represents the distribution of the true labels, and q represents the predicted label distribution of the trained model. The cross-entropy loss function can measure the similarity between p and q. Binary cross-entropy is shown in formula (2):
[0021]
[0022] Advantages of the present invention:
[0023] By utilizing features of different scales, the present invention strengthens feature propagation and effectively alleviates the problem of gradient disappearance that occurs during small-sample learning;
[0024] Skip connections are added to the network to enhance the network's ability to utilize features of different scales in the final prediction. This network structure can reduce the number of learning parameters in the model and alleviate the overfitting problem at the same time;
[0025] In the case of limited training samples, the generalization ability of the model is further improved;
[0026] When performing image recognition in scenarios with fewer datasets, single categories, and complex scenes, the recognition accuracy is effectively improved;
[0027] By sharing the local connection weights to generate different feature maps to represent features, the computational complexity of the network is reduced to a large extent.
[0028] In addition to the purposes, features, and advantages described above, the present invention has other purposes, features, and advantages. The present invention will be further described in detail below with reference to the drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The drawings forming a part of this application are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.
[0030] Figure 1 is the network structure diagram of the small-sample SAR target recognition method based on the siamese neural network of the present invention;
[0031] Figure 2 is the DenseBlock module diagram of the present invention;
[0032] Figure 3 is the Inception unit structure diagram of the present invention;
[0033] Figure 4 is the transition layer diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] In order to make the purposes, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0035] (1) In the multi-scale dense connection model based on Siamese neural network for few shot SAR target recognition (MDSN), a multi-scale dense connection network structure is used to solve the overfitting problem, and the idea of Siamese neural network is utilized. The MDSN network adopts dual inputs. As Figure 1 shown, the weights of the dashed part are shared. After two inputs obtain two outputs, the loss is calculated. For an image, a convolutional kernel scans each pixel of the image in turn, and the weights for processing each pixel point remain unchanged, that is, weight sharing is achieved. The network output obtained through this convolutional method is called a feature map of the image, that is, one convolutional kernel corresponds to one feature map. In order to extract more features from the original natural data, the input data can be processed using multiple convolutional kernels to obtain the same number of differentiated feature maps. Therefore, after introducing the idea of weight sharing, the network no longer expresses features through a large number of connection weights, but generates different feature maps by sharing local connection weights to express features. This can greatly reduce the computational complexity of the network and facilitate the actual implementation of the network. At the same time, various features of the local structure of the image are mined, which is beneficial to the recognition of image targets.
[0036] (2) The few shot recognition method based on the multi-scale dense connection model of Siamese neural network contains two symmetric branches. Each branch uses DenseNet as the backbone extraction network. Aiming at the problem of gradient disappearance that occurs when the DenseNet network model conducts few shot learning, Inception is used to replace the conventional convolutional layer in the backbone network. Each branch contains 3 DenseBlocks, and each DenseBlock contains 4 Inception units. As Figure 2 shown, each Inception unit outputs 12 channels of feature maps. The ratio of the output numbers of feature map channels of different convolutional kernels (3×3, 5×5, 7×7) is 1:3:8.
[0037] (3) Let \(H_l\) represent the output of the \(l\)-th layer in the DenseBlock, then the output \(H_{l + 1}\) of the \((l + 1)\)-th layer can be expressed as formula (1)
[0038] \(H_{l + 1} = [[f_7(H_l), f_5(H_l), f_3(H_l)], H_{l - 1},..., H_1]\) (1)
[0040] The fn() represents a convolutional layer with a convolution kernel size of n, and [] represents the connection operation on the channels. The convolutional layer consists of four basic operations: batch normalization, bias-free convolution, ReLU activation function, and dropout with a probability of 0.2. The Inception unit consists of several convolutional layers. Its first 1×1 convolution is introduced as a bottleneck layer, which produces a 4-fold increase in the feature map dimension. Then, convolutional layers with different kernel sizes are used to process the feature maps. For a large-size convolution kernel, such as 7×7, two convolution kernels with sizes of 7×1 and 1×7 can be used to replace this 7×7 kernel, as Figure 3 shown. The bottleneck layer and the replacement operation help reduce the model complexity and alleviate overfitting. In addition, in order to utilize features with different grid sizes, the feature maps generated by the first DenseBlock and the second DenseBlock are added to the final fully connected layer through 1×1 convolutional layers and global average pooling.
[0041] (4) The feature extraction module is mainly composed of 3 DenseBlock modules. After the first two DenseBlock blocks, a transition layer is connected. Through the transition layer between adjacent DenseBlocks, the grid size and the number of channels of the feature map are reduced to half of the original size. The transition layer mainly includes a 1×1 convolution and an average pooling with a stride of 2, as Figure 4 shown. The transition layer mainly plays two roles:
[0042] 1) Prevent the number of features from increasing infinitely and further compress the data.
[0043] 2) Downsample to reduce the resolution of the feature map.
[0044] (5) During training, the binary cross-entropy loss function is used to train the model to measure the similarity between two input samples. Binary cross-entropy is a commonly used loss function. p represents the distribution of the true labels, and q represents the predicted label distribution of the trained model. The cross-entropy loss function can measure the similarity between p and q. The formula for binary cross-entropy is shown in Formula 2:
[0045]
[0046] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A small-sample SAR image target recognition method based on a Siamese neural network, characterized in that, it includes two symmetric branches, and each branch uses DenseNet as the backbone extraction network, and Inception is used to replace the conventional convolutional layer in the backbone network; each branch contains 3 DenseBlocks, each DenseBlock contains 4 Inception units, and each Inception unit outputs 12 channels of the feature maps it contains; the ratio of the output numbers of the feature map channels of different convolutional kernels is 1:3:8; If Hl represents the output of the l-th layer in the DenseBlock, then the output Hl+1 of the (l + 1)-th layer is expressed as formula (1): (1) fn() represents a convolutional layer with a convolutional kernel size of n, and [] represents the concatenation operation on the channels; the convolutional layer consists of four basic operations: processing normalization, bias-free convolution, ReLU activation function, and dropout; The Inception unit consists of multiple convolutional layers, and its first 1×1 convolution is introduced as a bottleneck layer, and this bottleneck layer generates 4 times the feature map dimension; Then, the feature map is processed using convolutional layers with different kernel sizes; The feature maps generated by the first DenseBlock and the second DenseBlock are added to the final fully connected layer through a 1×1 convolutional layer and global average pooling; The feature extraction module consists of 3 DenseBlocks; Transition layers are connected after the first two DenseBlocks respectively, and the grid size and the number of channels of the feature map are reduced to half of the original size through the transition layers between adjacent DenseBlocks; The transition layer contains a 1×1 convolution and average pooling with a stride of 2.
2. The small-sample SAR image target recognition method based on a Siamese neural network according to claim 1, characterized in that, during training, the binary cross-entropy loss function is used for model training to measure the similarity between two input samples; The binary cross-entropy is a Loss function, p represents the distribution of the true labels, q represents the predicted label distribution of the trained model, the cross-entropy loss function measures the similarity between p and q, and the binary cross-entropy is shown in formula (2): (2)。
Citation Information
Patent Citations
Local descriptor learning method based on RGBD data
CN108171249A
SAR image change detection method based on multi-scale differential feature attention mechanism
CN114926746A