Aircraft Target Recognition Method in SAR Images Based on Fusion of Semantic and Texture Features
By fusing high-level semantic features and texture features extracted by deep neural networks in SAR image aircraft target recognition, information fusion features are generated, and the problem of insufficient recognition accuracy and robustness in the prior art is solved, and higher recognition accuracy and effectiveness are achieved.
Patent Information
- Application Number
- CN202211056668.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-31
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-08-31
AI Technical Summary
The existing SAR image aircraft target recognition method relies on abstract features extracted by deep neural networks, resulting in insufficient recognition accuracy and robustness, especially under different imaging conditions.
A SAR image aircraft target recognition method based on the fusion of semantic and texture features is designed. By extracting high-level semantic features and texture features and fusing them to generate information fusion features, to improve the generalization ability and recognition accuracy of the model.
By fusing semantics and texture features, the accuracy and effectiveness of aircraft target recognition in SAR image are improved, the generalization ability and recognition accuracy of the model are enhanced, and the recognition bottlenecks are solved when deep neural networks are used alone.
Smart Images

Figure CN115471763B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of radar target recognition, and in particular to a SAR image aircraft target recognition method based on the fusion of semantic and texture features. Background Art
[0002] Synthetic Aperture Radar (SAR) technology is a pulse radar technology that uses mobile radars mounted on satellites or aircraft to obtain high-precision radar target images of geographic areas. It has all-weather and all-day working capabilities and a certain degree of penetration. In view of these advantages, it is widely used in mineral exploration, marine environment monitoring, military defense and other fields. In particular, the research on aircraft target recognition is of great significance both in the military field and in the civilian field. Therefore, the research on aircraft recognition in SAR images has attracted widespread attention from scholars at home and abroad.
[0003] Traditional aircraft target recognition methods mainly use some manually set features to distinguish different types of aircraft. These features include geometric shape, texture features, scattering features, scale-invariant features, grayscale features, etc. Classifiers use the differences between various features to determine the category of input samples. Widely used classifiers include support vector machines, K-nearest neighbor algorithms, and naive Bayes algorithms. Most SAR target methods based on traditional manual features require complex mathematical theories and repeated manual model modification. In addition, this method has weak model generalization ability, is time-consuming, labor-intensive, and inefficient.
[0004] With more spaceborne SAR systems in orbit, more and more SAR data are available, and thanks to the boom in artificial intelligence big data research, especially the rapid development of deep convolutional neural network technology, it has become possible to conduct research on large-scale SAR data aircraft target recognition. Deep learning can obtain better results by assigning features that previously required manual design to complex network structures. The target features automatically extracted by deep networks can provide better representation capabilities than traditional manual features. Therefore, models based on deep neural networks almost completely dominate the field of image target recognition. However, due to the special imaging mechanism of SAR, the imaging results of the same target under different imaging conditions often vary greatly, which is very unfavorable for neural networks that strongly rely on training data. Therefore, only using abstract features extracted by deep neural networks to classify SAR aircraft targets limits the recognition accuracy and robustness of the recognition method.
[0005] The applicant discovered in the study that texture is an important visual clue, a feature that is ubiquitous and difficult to describe in images, and the texture of SAR target images varies with the wavelength, resolution and angle of incidence of the radar system, and also with the composition of the target and the arrangement of background features. In other words, the difference between different types of aircraft targets is often not in the grayscale size, but in their texture differences. Therefore, if the texture features can be combined with the features obtained by the deep learning network, the generalization ability of the model can be improved, thereby improving the accuracy of classification. Therefore, how to design a SAR image aircraft target recognition method that can combine the target texture features with the high-level semantic features obtained by the deep learning network is a technical problem that needs to be solved urgently. Summary of the invention
[0006] In view of the deficiencies of the above-mentioned prior art, the technical problem to be solved by the present invention is: how to design a SAR image aircraft target recognition method based on the fusion of semantic and texture features, so as to combine the target texture features with the high-level semantic features obtained by the deep learning network, thereby improving the generalization ability and precision of the model, thereby improving the accuracy and effectiveness of SAR image aircraft target recognition.
[0007] In order to solve the above technical problems, the present invention adopts the following technical solutions:
[0008] The SAR image aircraft target recognition method based on semantic and texture feature fusion includes:
[0009] S1: Acquire the SAR image to be identified;
[0010] S2: Input the SAR image to be identified into the trained target recognition model and output the corresponding target recognition prediction value;
[0011] When training the target recognition model, firstly, a training set containing several SAR images is input into the target recognition model; secondly, the high-level semantic features of the SAR image are extracted through a deep neural network; at the same time, the texture features of the SAR image are extracted and a texture feature matrix is constructed; then, the high-level semantic features and the texture feature matrix are fused to generate information fusion features; finally, prediction is performed based on the information fusion features to generate target recognition prediction values, and model training is completed based on the target recognition prediction values;
[0012] S3: Realize target recognition of the SAR image to be identified based on the target recognition prediction value output by the target recognition model.
[0013] Preferably, in step S2, ResNet34 is used as the backbone network of the deep neural network for extracting high-level semantic features.
[0014] Preferably, high-level semantic features are extracted through the following steps:
[0015] S201: Input the SAR image into the convolution layer and output the feature map F1;
[0016] S202: Input the feature map F1 into the attention module, and output the feature map F2 with added attention;
[0017] S203: Input the feature map F2 with additional attention into the stacked residual layer for residual learning, deepen the depth of the network, and obtain the feature map F3 with high-level semantic information as a high-level semantic feature.
[0018] Preferably, in step S202, the attention module includes a channel attention module and a spatial attention module;
[0019] The feature map F1 is used as the input of the channel attention module: First, the feature map F1 is subjected to global maximum pooling and global average pooling to obtain the feature map F 1,1 and F 1,2 ; Secondly, the feature map F 1,1 and F 1,2 They are sent to the neural network respectively, and the output feature map F 1,1 ′ and F 1,2 ′ performs a single element-based addition operation, and then obtains the channel feature map F1′ through a sigmoid activation operation; finally, the channel feature map F1′ is multiplied by the feature map F1 to obtain the channel attention feature F1 CAM ;
[0020] The channel attention feature F1 CAM As the input of the spatial attention module: first, the channel attention feature F1 CAM After channel-based global maximum pooling and global average pooling, the feature map F is obtained. 1,1 CAM and F 1,2 CAM ; Secondly, the feature map F 1,1 CAM and F 1,2 CAM Channel splicing is performed to obtain a channel splicing feature map; the channel splicing feature map is then reduced to one channel through a convolution operation, and then the spatial attention feature F1 is generated through a Sigmoid activation function operation. SAM ; Finally, the spatial attention feature F1 SAM And channel attention feature F1 CAM Do multiplication and get the feature map F2 with additional attention.
[0021] Preferably, in step S203, the stacked residual layer consists of four connected residual layers.
[0022] Preferably, in step S2, a texture feature matrix is constructed by the following steps:
[0023] S211: define a circular neighborhood window with a radius of R on the SAR image, assume that there are P sampling points in the circular neighborhood window, calculate the coordinate value of the pixel position corresponding to each sampling point and determine its pixel gray value;
[0024] S212: Taking the pixel gray value of the pixel position corresponding to the center point in the circular neighborhood window as the threshold, then comparing the pixel gray values of the P sampling points with the threshold: if it is greater than the threshold, the pixel position corresponding to the sampling point is marked as 1; otherwise, it is marked as 0;
[0025] S213: Obtain P binary numbers in the circular neighborhood window through step S212, and then convert the P binary numbers into decimal numbers as the texture feature values of the pixel position corresponding to the center point of the circular neighborhood window;
[0026] S214: Slide the circular neighborhood window on the SAR image and repeat steps S211 to S213 to calculate the texture feature value of each pixel position on the SAR image, thereby forming a texture feature matrix of the SAR image.
[0027] Preferably, in step S211, the coordinate value of the sampling point is calculated by the following formula:
[0028]
[0029] Where: (x p ,y p ),p∈P represents the coordinate value of the pth sampling point; (x c ,y c ) represents the coordinate value of the center point of the circular neighborhood window;
[0030] If (x p ,y p ) is not at an integer position, the pixel value of the pixel position corresponding to the sampling point is determined by bilinear interpolation, and the formula is as follows:
[0031]
[0032] Where: f(x1,y1), f(x1,y2), f(x2,y1), and f(x2,y2) represent the pixel grayscale values of the four nearest pixels corresponding to the sampling point in the SAR image, that is, x1-x2=1, y1-y2=1, u=x-x1, and v=y-y1 represent the horizontal and vertical distances from the sampling point to the rectangle formed by the four nearest pixels; i(x p ,y p ) indicates that the coordinate is (xp ,y p )’s sampling point corresponds to the pixel gray value of the pixel position.
[0033] Preferably, in step S213, the texture feature value is calculated by the following formula:
[0034]
[0035] Where: LBP(x c ,y c ) indicates that the coordinate is (x c ,y c ) is the texture feature value of the pixel position corresponding to the center point of the circular neighborhood window; i c Indicates the center pixel (x c ,y c )’s gray value; i p It represents the pixel gray value of the pixel position corresponding to the pth sampling point in the circular neighborhood window; s represents the sign function.
[0036] Preferably, in step S2, the high-level semantic feature and texture feature matrices are converted into vectors respectively; then the vectors corresponding to the high-level semantic features and texture features are concatenated to obtain a new feature vector; finally, the new feature vector is input into the fully connected layer for information fusion, and then the corresponding information fusion features are output.
[0037] Preferably, in step S2, the information fusion feature is input into the Softmax function, and the corresponding target recognition prediction value is output; then the model training is completed according to the preset number of training iterations and gradient descent batch size combined with the cross entropy loss function.
[0038] The SAR image aircraft target recognition method based on semantic and texture feature fusion in the present invention has the following beneficial effects:
[0039] The present invention extracts high-level semantic features and texture features of SAR images respectively and fuses them to obtain information fusion features, and then predicts based on the information fusion features to complete model training. On the one hand, the high-level semantic features are highly associated with the targets in the SAR images, contain rich target information, and are conducive to improving the correct recognition rate of the targets, but the target positions are relatively rough. On the other hand, the texture features can provide discriminative target information, and have the requirements of grayscale and rotation invariance, and have the advantage of accurate target positions, but the feature semantic information contained is relatively small. Therefore, the high-level semantic features are fused with the texture features, which can not only provide richer discriminative target information for the target recognition model under the premise of ensuring the relevance of the aircraft targets, but also provide accurate target positions, thereby effectively improving the recognition accuracy of SAR aircraft targets, and solving the bottleneck caused by the existing SAR target classification relying only on deep neural networks. That is, the present invention can combine the target texture features with the high-level semantic features obtained by the deep learning network, thereby improving the generalization ability and precision of the model, thereby improving the accuracy and effectiveness of SAR image aircraft target recognition.
[0040] The present invention uses ResNet34 as the backbone network to extract high-level semantic features of aircraft targets, and combines the CBAM attention mechanism to guide the network to focus on the information of the target itself, rather than over-focusing on background clutter and speckle noise, which can ensure the effectiveness and accuracy of high-level semantic feature extraction, that is, it can extract high-level semantic features with a higher correlation with aircraft targets in SAR images, thereby further improving the accuracy of aircraft target recognition in SAR images.
[0041] The present invention extracts texture features of SAR images through a circular LBP operator, which can meet the needs of textures of different sizes and frequencies, can adapt to texture features of different scales and meet the requirements of grayscale and rotation invariance, that is, it can ensure the effectiveness and accuracy of texture feature extraction, so that texture features can better provide discriminative target information, thereby further improving the accuracy of SAR image aircraft target recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to make the purpose, technical solution and advantages of the invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings, in which:
[0043] Figure 1 It is the logic block diagram of the SAR image aircraft target recognition method based on the fusion of semantic and texture features;
[0044] Figure 2 This is a working principle diagram of the target recognition model;
[0045] Figure 3 This is the network structure diagram of the target recognition model;
[0046] Figure 4 There are five types of SAR aircraft targets and their corresponding optical images;
[0047] Figure 5 This is the network structure diagram of the CBAM attention module;
[0048] Figure 6 Visual graph of texture features extracted for each type of aircraft target. DETAILED DESCRIPTION
[0049] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. The components of the embodiments of the present invention generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but only represents selected embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work belong to the scope of protection of the present invention.
[0050] It should be noted that similar numbers and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in the subsequent drawings. In the description of the present invention, it should be noted that the orientation or position relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inside", "outside", etc. is based on the orientation or position relationship shown in the drawings, or the orientation or position relationship in which the invention product is usually placed when used, which is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation of the present invention. In addition, the terms "first", "second", "third", etc. are only used to distinguish the description, and cannot be understood as indicating or implying relative importance. In addition, the terms "horizontal", "vertical", etc. do not mean that the components are absolutely horizontal or suspended, but can be slightly tilted. For example, "horizontal" only means that its direction is more horizontal than "vertical", and does not mean that the structure must be completely horizontal, but can be slightly tilted. In the description of the present invention, it is also necessary to explain that, unless otherwise clearly specified and limited, the terms "set", "install", "connect", and "connect" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two elements. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0051] The following is a further detailed description through specific implementation methods:
[0052] Example:
[0053] This embodiment discloses a SAR image aircraft target recognition method based on the fusion of semantic and texture features.
[0054] like Figure 1 As shown, the SAR image aircraft target recognition method based on semantic and texture feature fusion includes:
[0055] S1: Acquire the SAR image to be identified;
[0056] S2: Input the SAR image to be identified into the trained target recognition model and output the corresponding target recognition prediction value;
[0057] Combination Figure 2 and Figure 3As shown in the figure, when training the target recognition model: first, a training set containing several SAR images is input into the target recognition model; secondly, high-level semantic features of the SAR image are extracted through a deep neural network; at the same time, texture features of the SAR image are extracted, and a texture feature matrix is constructed; then, high-level semantic features and texture feature matrix are fused to generate information fusion features; finally, prediction is performed based on the information fusion features to generate target recognition prediction values, and model training is completed based on the target recognition prediction values;
[0058] S3: Realize target recognition of the SAR image to be identified based on the target recognition prediction value output by the target recognition model.
[0059] In this embodiment, the target recognition prediction value refers to the output probability predicted by the target recognition model, which is a vector. The position corresponding to the element with the largest target recognition prediction value is the category of the sample.
[0060] Take the aircraft target as an example. Figure 4 As shown in the figure, there are five categories of aircraft samples (such as civil aircraft, fighters, bombers, transport aircraft, and tankers). Specifically, taking an aircraft sample as an example, the target recognition prediction value finally output by the network is assumed to be a probability vector of length 5 [0.04, 0, 0.92, 0.02, 0.02], and all values add up to 1. The element value at each position corresponds to the probability that the sample belongs to that category, that is, 0.92 corresponds to the category: bomber.
[0061] The present invention extracts high-level semantic features and texture features of SAR images respectively and fuses them to obtain information fusion features, and then predicts based on the information fusion features to complete model training. On the one hand, the high-level semantic features are highly associated with the targets in the SAR images, contain rich target information, and are conducive to improving the correct recognition rate of the targets, but the target positions are relatively rough. On the other hand, the texture features can provide discriminative target information, and have the requirements of grayscale and rotation invariance, and have the advantage of accurate target positions, but the feature semantic information contained is relatively small. Therefore, the high-level semantic features are fused with the texture features, which can not only provide richer discriminative target information for the target recognition model under the premise of ensuring the relevance of the aircraft targets, but also provide accurate target positions, thereby effectively improving the recognition accuracy of SAR aircraft targets, and solving the bottleneck caused by the existing SAR target classification relying only on deep neural networks. That is, the present invention can combine the target texture features with the high-level semantic features obtained by the deep learning network, thereby improving the generalization ability and precision of the model, thereby improving the accuracy and effectiveness of SAR image aircraft target recognition.
[0062] In the specific implementation process, ResNet34 is used as the backbone network of the deep neural network for extracting high-level semantic features. The network structure and parameters of ResNet34 are shown in Table 1.
[0063] Table 1 Network structure and parameters of ResNet34
[0064]
[0065] Combination Figure 3 As shown in Figure 2, high-level semantic features are extracted through the following steps:
[0066] S201: Input the SAR image into a convolution layer with a convolution kernel size of 7×7, a step size of 2, and a channel number of 64, and output a feature map F1;
[0067] S202: Input the feature map F1 into the CBAM (Convolutional Block Attention Module) attention module, and output the feature map F2 with additional attention;
[0068] S203: Input the feature map F2 with additional attention into the stacked residual layer for residual learning, deepen the depth of the network, and obtain the feature map F3 with high-level semantic information as a high-level semantic feature.
[0069] In this embodiment, the stacked residual layer consists of four connected residual layers.
[0070] like Figure 5 As shown in (a), the attention module includes a channel attention module (CAM) and a spatial attention module (SAM);
[0071] like Figure 5 As shown in (b), the feature map F1 is used as the input of the channel attention module: First, the feature map F1 is subjected to global maximum pooling and global average pooling to obtain the feature map F 1,1 and F 1,2 ; Secondly, the feature map F 1,1 and F 1,2 They are sent to the neural network respectively, and the output feature map F 1,1 ′ and F 1,2 ′ performs a single element-based addition operation, and then obtains the channel feature map F1′ through a sigmoid activation operation; finally, the channel feature map F1′ is multiplied by the feature map F1 to obtain the channel attention feature F1 CAM ;
[0072] like Figure 5 As shown in (c), the channel attention feature F1 CAM As the input of the spatial attention module: first, the channel attention feature F1 CAMAfter channel-based global maximum pooling and global average pooling, the feature map F is obtained. 1,1 CAM and F 1,2 CAM ; Secondly, the feature map F 1,1 CAM and F 1,2 CAM Channel splicing is performed to obtain a channel splicing feature map; the channel splicing feature map is then reduced to one channel through a 7×7 convolution operation, and then the spatial attention feature F1 is generated through a Sigmoid activation function operation. SAM ; Finally, the spatial attention feature F1 SAM And channel attention feature F1 CAM Do multiplication and get the feature map F2 with additional attention.
[0073] The present invention uses ResNet34 as the backbone network to extract high-level semantic features of aircraft targets, and combines the CBAM attention mechanism to guide the network to focus on the information of the target itself, rather than over-focusing on background clutter and speckle noise, which can ensure the effectiveness and accuracy of high-level semantic feature extraction, that is, it can extract high-level semantic features with a higher correlation with aircraft targets in SAR images, thereby further improving the accuracy of aircraft target recognition in SAR images.
[0074] In the specific implementation process, the texture features are extracted through the circular LBP (Local Binary Pattern, LBP) operator.
[0075] The texture feature matrix is constructed by the following steps:
[0076] S211: define a circular neighborhood window with a radius of R on the SAR image, assume that there are P sampling points in the circular neighborhood window, calculate the coordinate value of the pixel position corresponding to each sampling point and determine its pixel gray value;
[0077] In this embodiment, since the used data set has serious noise interference, the original SAR image is preprocessed first, and pixel values above 70% of the normalized amplitude value are taken to reduce obvious noise interference.
[0078] The coordinate values of the sampling points are calculated using the following formula:
[0079]
[0080] Where: (x p ,y p ),p∈P represents the coordinate value of the pth sampling point; (x c ,y c ) represents the coordinate value of the center point of the circular neighborhood window;
[0081] If (x p ,y p ) is not at an integer position, the pixel value of the pixel position corresponding to the sampling point is determined by bilinear interpolation, and the formula is as follows:
[0082]
[0083] Where: f(x1,y1), f(x1,y2), f(x2,y1), and f(x2,y2) represent the pixel grayscale values of the four nearest pixels corresponding to the sampling point in the SAR image, that is, x1-x2=1, y1-y2=1, u=x-x1, and v=y-y1 represent the horizontal and vertical distances from the sampling point to the rectangle formed by the four nearest pixels; i(x p ,y p ) indicates that the coordinate is (x p ,y p )’s sampling point corresponds to the pixel gray value of the pixel position.
[0084] S212: Taking the pixel gray value of the pixel position corresponding to the center point in the circular neighborhood window as the threshold, then comparing the pixel gray values of the P sampling points with the threshold: if it is greater than the threshold, the pixel position corresponding to the sampling point is marked as 1; otherwise, it is marked as 0;
[0085] S213: Obtain P binary numbers in the circular neighborhood window through step S212, and then convert the P binary numbers into decimal numbers as the texture feature values of the pixel position corresponding to the center point of the circular neighborhood window;
[0086] The texture feature value is calculated by the following formula:
[0087]
[0088] Where: LBP(x c ,y c ) indicates that the coordinate is (x c ,y c ) is the texture feature value of the pixel position corresponding to the center point of the circular neighborhood window; i c Indicates the center pixel (x c ,y c )’s gray value; i p It represents the pixel gray value of the pixel position corresponding to the pth sampling point in the circular neighborhood window; s represents the sign function.
[0089] S214: Slide the circular neighborhood window on the SAR image and repeat steps S211 to S213 to calculate the texture feature value of each pixel position on the SAR image, thereby forming a texture feature matrix of the SAR image.
[0090] The present invention extracts texture features of SAR images through a circular LBP operator, which can meet the needs of textures of different sizes and frequencies, can adapt to texture features of different scales and meet the requirements of grayscale and rotation invariance, that is, it can ensure the effectiveness and accuracy of texture feature extraction, so that texture features can better provide discriminative target information, thereby further improving the accuracy of SAR image aircraft target recognition.
[0091] During the specific implementation process, the high-level semantic feature and texture feature matrices are converted into vectors respectively; then the vectors corresponding to the high-level semantic features and texture features are concatenated to obtain a new feature vector; finally, the new feature vector is input into the fully connected layer for information fusion, and then the corresponding information fusion features are output.
[0092] During the specific implementation process, the information fusion features are input into the Softmax function to output the target recognition prediction value; then the model training is completed according to the preset network parameters such as the number of training iterations and gradient descent batch size combined with the cross entropy loss function.
[0093] In order to better illustrate the advantages of the technical solution of the present invention, the following experiments are disclosed in this embodiment.
[0094] 1. Evaluation Metrics
[0095] 1) Overall Accuracy (OA): OA refers to the ratio of correctly classified samples to the total number of samples. Generally speaking, the higher the OA, the higher the classification accuracy, as shown in the following formula:
[0096]
[0097] Among them, TP represents true positive, FP represents false positive, TN represents true negative, and FN represents false negative.
[0098] 2) Precision: It indicates the proportion of instances predicted to be positive that are actually positive, as shown in the following formula:
[0099]
[0100] 3) Recall: It indicates the ratio of instances predicted as positive to all actual positive examples, as shown in the following formula:
[0101]
[0102] 4) F1-score: expressed as the weighted harmonic mean of precision and recall, namely:
[0103]
[0104] 5) AUC value: First, the ROC (Receiver Operating Characteristic Curve) curve is a common performance indicator for measuring recognition tasks. The closer the ROC curve is to the upper left corner, the better the discrimination effect. However, when the distances between different ROC curves are close, it is difficult to directly judge the pros and cons of recognition performance. AUC is also a performance indicator for measuring recognition tasks. It is defined as the area enclosed by the coordinate axis under the ROC curve, which can intuitively reflect the performance of the recognition task.
[0105] 2. Experimental Preparation
[0106] The aircraft dataset used in this experiment comes from satellite SAR large scene images. There are five types of aircraft target images, and each type of SAR aircraft image and corresponding optical image are as follows: Figure 4 As shown in Table 2, the dataset is divided into training set and test set in a ratio of 7:3.
[0107] Table 2 Training set and test set
[0108]
[0109] 3. Experimental setup
[0110] After debugging various parameters of ResNet34 for many times, the parameter combination that optimizes network performance is selected. The pre-trained ResNet34 network is used to fine-tune the target recognition model (hereinafter referred to as the multi-feature fusion network model), the training iteration round (Epoch) is set to 30, and the adaptive momentum estimation algorithm (Adaptive Moment Estimation, Adam) based on small batch training data is selected to train the multi-feature fusion network model. The learning rate parameter (LearningRate) and the gradient descent batch size (Batch Size) are selected as 0.0001 and 16 respectively, and the loss function uses the cross entropy loss function (Cross Entropy Loss).
[0111] 4. Extract texture features
[0112] When the circular LBP algorithm is used to extract the texture features of the SAR aircraft target, the circular window radius is 3, which contains 8 sampling points. The extracted visual texture feature map and the corresponding SAR image are shown in Figure 2. Figure 6As shown. It can be seen that the circular LBP algorithm realizes the texture feature extraction of the complete aircraft target, and shows a strong suppression ability to noise by performing statistical calculations on the information of the pixel area. However, due to the interference of some discrete clutter points or speckle noise, there are still some feature points of non-target textures in the texture feature map. In addition, the texture features of different aircraft targets also differ in the wings and fuselage. Feature fusion with the high-level semantic features extracted by the ResNet34 network is conducive to assisting the deep neural network to capture more discriminative target information and obtain higher recognition accuracy.
[0113] 5. Performance description
[0114] In order to verify the performance of SAR aircraft target recognition based on multi-feature fusion in the present invention, target recognition was performed on ResNet-34, the model after introducing the attention mechanism, and the multi-feature fusion model in the aircraft data set, and the results are shown in Table 3. It can be seen that the recognition rate is 95.33% when only ResNet-34 is used. After further introducing the CBAM attention mechanism on the basis of ResNet-34, the recognition rate of aircraft targets is increased to 95.63%. The recognition rate after fusing texture features is 96.35%, which is about 0.72 percentage points higher than the model introducing the attention mechanism and about 1 percentage point higher than ResNet-34. The recognition results illustrate the effectiveness of the SAR aircraft target recognition method based on multi-feature fusion in the present invention, and the present invention significantly improves the recognition accuracy of SAR aircraft targets.
[0115] Table 3 Performance comparison
[0116]
[0117] 6. Model comparison
[0118] Six typical convolutional neural network models are selected to conduct recognition experiments on the aircraft dataset, and the results are shown in Table 4.
[0119] Combined with Table 4, it can be seen that the multi-feature fusion model of the present invention is significantly better than all the comparison models. Among the six comparison models, EfficientNet achieved the highest recognition accuracy of 94.28%, but it is still about 2 percentage points lower than the model of the present invention. Table 5 shows the AUC measurement results of different network models. It can be seen that the multi-feature fusion model of the present invention achieved the highest AUC value of 0.9952, showing better recognition performance than the existing optimal method.
[0120] Table 4 Performance comparison of six comparison models
[0121]
[0122] Table 5 AUC performance comparison of six comparison models
[0123]
[0124] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit the technical solution. Those skilled in the art should understand that those modifications or equivalent substitutions of the technical solution of the present invention that do not depart from the purpose and scope of the technical solution should be included in the scope of the claims of the present invention.
Claims
1. A SAR image aircraft target recognition method based on the fusion of semantic and texture features, characterized in that: include: S1: Acquire the SAR image to be identified; S2: Input the SAR image to be identified into the trained target recognition model and output the corresponding target recognition prediction value; When training the target recognition model, firstly, a training set containing several SAR images is input into the target recognition model; secondly, the high-level semantic features of the SAR image are extracted through a deep neural network; at the same time, the texture features of the SAR image are extracted and a texture feature matrix is constructed; then, the high-level semantic features and the texture feature matrix are fused to generate information fusion features; finally, prediction is performed based on the information fusion features to generate target recognition prediction values, and model training is completed based on the target recognition prediction values; In step S2, ResNet34 is used as the backbone network of the deep neural network for extracting high-level semantic features; The high-level semantic features are extracted through the following steps: S201: Input the SAR image into the convolution layer and output the feature map F1; S202: Input the feature map F1 into the attention module, and output the feature map F2 with added attention; S203: Input the feature map F2 with additional attention into the stacked residual layer for residual learning, deepen the depth of the network, and obtain the feature map F3 with high-level semantic information as a high-level semantic feature; The texture feature matrix is constructed by the following steps: S211: define a circular neighborhood window with a radius of R on the SAR image, assume that there are P sampling points in the circular neighborhood window, calculate the coordinate value of the pixel position corresponding to each sampling point and determine its pixel gray value; S212: Taking the pixel gray value of the pixel position corresponding to the center point in the circular neighborhood window as the threshold, then comparing the pixel gray values of the P sampling points with the threshold: if it is greater than the threshold, the pixel position corresponding to the sampling point is marked as 1; otherwise, it is marked as 0; S213: Obtain P binary numbers in the circular neighborhood window through step S212, and then convert the P binary numbers into decimal numbers as the texture feature values of the pixel position corresponding to the center point of the circular neighborhood window; S214: sliding the circular neighborhood window on the SAR image, and repeating steps S211 to S213 to calculate the texture feature value of each pixel position on the SAR image, thereby forming a texture feature matrix of the SAR image; S3: Realize target recognition of the SAR image to be identified based on the target recognition prediction value output by the target recognition model.
2. The SAR image aircraft target recognition method based on semantic and texture feature fusion as claimed in claim 1, characterized in that: In step S202, the attention module includes a channel attention module and a spatial attention module; The feature map F1 is used as the input of the channel attention module: First, the feature map F1 is subjected to global maximum pooling and global average pooling to obtain the feature map F 1,1 and F 1,2 ; Secondly, the feature map F 1,1 and F 1,2 They are sent to the neural network respectively, and the output feature map F 1,1 ′ and F 1,2 ′ performs a single element-based addition operation, and then obtains the channel feature map F1′ through a sigmoid activation operation; finally, the channel feature map F1′ is multiplied by the feature map F1 to obtain the channel attention feature F1 CAM ; The channel attention feature F1 CAM As the input of the spatial attention module: first, the channel attention feature F1 CAM After channel-based global maximum pooling and global average pooling, the feature map F is obtained. 1,1 CAM and F 1,2 CAM ; Secondly, the feature map F 1,1 CAM and F 1,2 CAM Channel splicing is performed to obtain a channel splicing feature map; the channel splicing feature map is then reduced to one channel through a convolution operation, and then the spatial attention feature F1 is generated through a Sigmoid activation function operation. SAM ; Finally, the spatial attention feature F1 SAM And channel attention feature F1 CAM Do multiplication and get the feature map F2 with additional attention.
3. The SAR image aircraft target recognition method based on semantic and texture feature fusion as claimed in claim 1, characterized in that: In step S203, the stacked residual layer consists of four connected residual layers.
4. The SAR image aircraft target recognition method based on semantic and texture feature fusion as claimed in claim 1, characterized in that: In step S211, the coordinate value of the sampling point is calculated by the following formula: Where: (x p ,y p ),p∈P represents the coordinate value of the pth sampling point; (x c ,y c ) represents the coordinate value of the center point of the circular neighborhood window; If (x p ,y p ) is not at an integer position, the pixel value of the pixel position corresponding to the sampling point is determined by bilinear interpolation, and the formula is as follows: Where: f(x1,y1), f(x1,y2), f(x2,y1), and f(x2,y2) represent the pixel grayscale values of the four nearest pixels corresponding to the sampling point in the SAR image, that is, x1-x2=1, y1-y2=1, u=x-x1, and v=y-y1 represent the horizontal and vertical distances from the sampling point to the rectangle formed by the four nearest pixels; i(x p ,y p ) indicates that the coordinate is (x p ,y p )’s sampling point corresponds to the pixel gray value of the pixel position.
5. The SAR image aircraft target recognition method based on semantic and texture feature fusion as claimed in claim 4 is characterized in that: In step S213, the texture feature value is calculated by the following formula: Where: LBP(x c ,y c ) indicates that the coordinate is (x c ,y c ) is the texture feature value of the pixel position corresponding to the center point of the circular neighborhood window; i c Indicates the center pixel (x c ,y c )’s gray value; i p It represents the pixel gray value of the pixel position corresponding to the pth sampling point in the circular neighborhood window; s represents the sign function.
6. The SAR image aircraft target recognition method based on semantic and texture feature fusion as claimed in claim 1, characterized in that: In step S2, the high-level semantic feature and texture feature matrices are converted into vectors respectively; then the vectors corresponding to the high-level semantic features and texture features are concatenated to obtain a new feature vector; finally, the new feature vector is input into the fully connected layer for information fusion, and then the corresponding information fusion features are output.
7. The SAR image aircraft target recognition method based on semantic and texture feature fusion as claimed in claim 6, characterized in that: In step S2, the information fusion feature is input into the Softmax function, and the corresponding target recognition prediction value is output; then the model training is completed according to the preset number of training iterations and gradient descent batch size combined with the cross entropy loss function.