An image-based aircraft target recognition method

By constructing an SDRSC target recognition model in air aircraft target recognition, and using the RSC module to optimize feature extraction and loss functions, the problem of low accuracy of target recognition in the prior art in air aircraft is solved, and higher recognition accuracy is achieved.

CN115082769BActive Publication Date: 2025-05-27HARBIN INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210750617.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-28
Publication Date
2025-05-27
Estimated Expiration
2042-06-28

AI Technical Summary

Technical Problem

The existing deep learning methods have low recognition accuracy when identifying air aircraft targets and cannot obtain accurate recognition results, mainly due to the scene complexity and feature surface correlation of air aircraft targets.

Method used

An image-based aircraft target recognition method is proposed. By constructing an SDRSC target recognition model, ResNet-50 is used as the backbone network, and RSC modules are embedded between its average pooling layer and linear classification layer. The RSC module calculates the degree of feature contribution, silences high-contribution features, forces the network to learn more features, and optimizes the loss function to suppress gradient hunger.

Benefits of technology

By optimizing feature extraction and loss functions, the accuracy of airplane target recognition is improved, the problem of low recognition accuracy in the prior art is solved, and higher recognition accuracy is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115082769B_ABST
    Figure CN115082769B_ABST
Patent Text Reader

Abstract

A method for recognizing an aircraft target based on an image belongs to the field of target recognition. The present invention solves the problem of low recognition accuracy when using existing deep learning methods to recognize aircraft targets in the air. Aiming at the high complexity of the data set, the present invention optimizes the feature extraction in the model training process by constructing an RSC module, and forces the network to learn more features to complete the target category prediction by muting high-gradient features. Aiming at the problem of surface correlation of the data set, a regularization method is proposed to optimize the loss function, thereby suppressing the gradient starvation phenomenon and improving the recognition accuracy. The method of the present invention can be applied to the field of target recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of target recognition, and particularly relates to an image-based aircraft target recognition method. Background Art

[0002] Although deep learning has achieved extremely brilliant achievements in the fields of image classification, target recognition, etc., for the recognition of airborne aircraft targets in multi-scenarios including military scenarios and civilian scenarios, due to the two main characteristics of airborne aircraft targets. One is that the scene is complex, and there are large differences in the postures, backgrounds, and observation perspectives of the same aircraft target; the other is that some features have surface correlations. For example, the same type of aircraft may have different paint jobs but the same fuselage structure, and different types of aircraft may have slight differences in the fuselage structure but similar paint jobs, and the paint job information may be assigned more learning weights. Therefore, this is a feature of surface correlation. Therefore, due to the particularity of airborne aircraft targets, the accuracy obtained by using existing deep learning methods to recognize airborne aircraft targets is low, and accurate recognition results of airborne aircraft targets cannot be obtained. Summary of the Invention

[0003] The purpose of the present invention is to solve the problem of low recognition accuracy obtained when using existing deep learning methods to recognize airborne aircraft targets, and to propose an image-based aircraft target recognition method.

[0004] The technical solution adopted by the present invention to solve the above technical problems is:

[0005] An image-based aircraft target recognition method, the method specifically includes the following steps:

[0006] Step 1: Construct a data set of airborne aircraft target images, and label the aircraft target images in the data set;

[0007] Step 2: Build an SDRSC target recognition model based on ResNet-50, and use the labeled data set to train the constructed SDRSC target recognition model;

[0008] Step 3: Collect the airborne aircraft target image to be recognized, and then input the airborne aircraft target image to be recognized into the SDRSC target recognition model trained in Step 2, and output the target recognition result through the SDRSC target recognition model.

[0009] Further, the data set constructed in Step 1 includes multi-type and multi-model aircraft target images in military scenarios and civil airliner images.

[0010] Further, the SDRSC target recognition model is specifically:

[0011] Select ResNet-50 as the backbone network, and embed the RSC module between the average pooling layer and the linear classification layer of the backbone network to obtain the SDRSC target recognition model.

[0012] Furthermore, the working process of the RSC module is as follows:

[0013] Step 1: Calculate the contribution degree of different-dimensional features output by the average pooling layer to the classification prediction;

[0014] Step 2: Initialize the hyperparameter p, and silence the dimensional features ranked in the top p% of the contribution degree. That is, construct a vector m with the same total dimension as the features output by the average pooling layer. The i-th element m(i) in the vector m is:

[0015]

[0016] where g z(i) is the contribution degree of the i-th dimensional feature output by the average pooling layer, and q p is the minimum contribution degree among the dimensional features ranked in the top p% of the contribution degree;

[0017] Step 3: Perform the Hadamard product of the features output by the average pooling layer and the vector m to obtain the updated features

[0018]

[0019] where ⊙ represents the Hadamard product;

[0020] Step 4: Calculate the output result of the updated features after passing through the softmax function:

[0021]

[0022] where is the output of the part used for class prediction in the backbone network;

[0023] According to the calculated use the gradient descent formula to update the model parameters, where y is the sample label, is the cross-entropy loss function, is the parameter of the feature with the highest weight.

[0024] Furthermore, the value of the hyperparameter p is 5.

[0025] Furthermore, the expression of the contribution degree g z(i) is as follows:

[0026]

[0027] Among them, z(i) represents the feature of the i-th dimension output by the average pooling layer, is the output based on the feature of the i-th dimension of the part used for class prediction in the backbone network, is the parameter of the part used for class prediction in the backbone network.

[0028] Furthermore, the loss function L adopted by the SDRSC target recognition model is:

[0029]

[0030] Among them, Y is the set of input sample labels, is the class prediction result of the SDRSC target recognition model for the input sample, λ ∈ [0, ∞) is the weight decay coefficient, and ∥∥ is the norm.

[0031] The beneficial effects of the present invention are:

[0032] Aiming at the characteristics of high complexity of the dataset, the present invention optimizes the feature extraction in the model training process by constructing an RSC module, and by measures of muting features with high contribution degrees, forces the network to learn more features to complete the target class prediction. Aiming at the problem of surface correlation of the dataset, the present invention optimizes the loss function by proposing a regularization method to suppress the gradient starvation phenomenon, thereby improving the recognition accuracy. Description of the Drawings

[0033] Figure 1 is the training flowchart of the SDRSC target recognition model;

[0034] Figure 2 is the typical sample diagram in the constructed dataset;

[0035] Figure 3 is the working flowchart of the RSC module;

[0036] Figure 4 is the loss function diagram of the SDRSC network;

[0037] Figure 5 is the comparison diagram of the recognition accuracy between the SDRSC network and ResNet-50;

[0038] Figure 6 is the comparison diagram of the iteration process between the SDRSC network of the present invention and the existing five algorithms of SelfRag, CDANN, DRO, MMD and IGA;

[0039] Figure 7It is a comparison chart of the recognition accuracy between the SDRSC network of the present invention and the existing five algorithms of SelfRag, CDANN, DRO, MMD, and IGA. Detailed implementation manners

[0040] Detailed implementation manner 1: A method for aircraft target recognition based on images described in this implementation manner specifically includes the following steps:

[0041] Step 1: Construct a dataset of aerial aircraft target images and label the aircraft target images in the dataset.

[0042] After labeling the aircraft type and model of the aircraft target image, it is used as the training set for the subsequent target recognition model.

[0043] Step 2: Construct an SDRSC target recognition model based on ResNet-50, and use the labeled dataset to train the constructed SDRSC target recognition model.

[0044] Step 3: Collect the aerial aircraft target image to be recognized, and then input the aerial aircraft target image to be recognized into the SDRSC target recognition model trained in Step 2, and output the target recognition result through the SDRSC target recognition model.

[0045] Detailed implementation manner 2: The difference between this implementation manner and Detailed implementation manner 1 is that the dataset constructed in Step 1 includes multi-type and multi-model aircraft target images in military scenarios and civil airliner images.

[0046] The civil airliner images in this implementation manner come from the FGVC-Aircraft dataset. Typical samples in the dataset are as Figure 2 shown. The sample images are taken in different backgrounds such as high altitude, low altitude, sea surface, and ground, and the aircraft is in different states such as takeoff, landing, cruising, and maneuvering; the sample images include various different angles such as front view, upward view, and side view.

[0047] Other steps and parameters are the same as those in Detailed implementation manner 1.

[0048] Detailed implementation manner 3: Combined with Figure 1 to illustrate this implementation manner. The difference between this implementation manner and Detailed implementation manner 1 or 2 is that the SDRSC target recognition model is specifically:

[0049] Select ResNet-50 as the backbone network, and embed the RSC module between the average pooling layer and the linear classification layer of the backbone network to obtain the SDRSC target recognition model.

[0050] The training process of the SDRSC target recognition model is as follows: Input an aircraft picture, which first enters the convolutional layer for feature extraction, then enters the normalization layer to optimize the output of neurons so that the output results follow a Gaussian distribution. Subsequently, it enters the max-pooling layer for dimensionality reduction. The residual network between the max-pooling layer and the average-pooling layer ensures that the network can be deeper without causing gradient explosion and degradation problems. The picture after dimensionality reduction is output to the RSC module by the average-pooling layer, aiming to extract more comprehensive features. After the network learns more diverse features, the picture is sent to a linear classifier for classification. Finally, the loss is calculated to suppress gradient starvation and achieve a better recognition effect.

[0051] Other steps and parameters are the same as those in the first or second specific implementation manner.

[0052] Specific implementation manner four: The difference between this implementation manner and one of the first to third specific implementation manners is that the working process of the RSC module is as follows:

[0053] Step 1: Calculate the contribution degree of different-dimensional features output by the average-pooling layer to classification prediction;

[0054] Step 2: Initialize the hyperparameter p, and perform silence processing on the dimensional features ranked in the top p% of the contribution degree, that is, construct a vector m with the same total dimension as the features output by the average-pooling layer. The i-th element m(i) in the vector m is:

[0055]

[0056] where g z(i) is the contribution degree of the i-th dimensional feature output by the average-pooling layer, and q p is the minimum contribution degree among the dimensional features ranked in the top p% of the contribution degree;

[0057] That is, if the corresponding element in g is one of the elements with the highest contribution degree within the discarded range, the element of the m vector is set to 0, otherwise it is set to 1;

[0058] Step 3: Perform the Hadamard product of the features output by the average-pooling layer and the vector m to silence the features with the greatest contribution to classification prediction within the p range, and obtain the updated features

[0059]

[0060] where ⊙ represents the Hadamard product;

[0061] Step 4: Calculate the output result of the updated features after passing through the softmax function:

[0062]

[0063] Among them, is the output of the part used for class prediction in the backbone network;

[0064] According to the calculated Adopt the gradient descent formula to update the model parameters Among them, y is the sample label, is the cross-entropy loss function, is the parameter of the feature with the highest weight.

[0065] In this embodiment, the RSC module is used to extract and optimize the features output by the average pooling layer. The extraction and optimization process only includes operations such as pooling and Hadamard product, and no additional parameters need to be learned except the weights of the original network, so as to learn more features, prevent overfitting caused by certain typical features, and optimize the feature extraction ability. The algorithm pseudocode is shown in Table 1:

[0066] Table 1

[0067]

[0068] Denote the feature tensor of the last convolutional layer of the backbone network as Z, and denote its gradient tensor as G. The gradient tensor can be calculated by backpropagation from the predicted class scores. The sizes of both Z and G are [7×7×512]. Set the hyperparameter p as the percentage of the number of features to be discarded in the total number of features (set to 5 here). Before the feature map is sent to the fully connected layer, dimensionality reduction is required. Here, global pooling between channels is performed on it, and this process will generate a weight matrix wi of size [7×7]. Use the weight matrix wi and the parameter p to mute several features in the feature z that contribute the most to the classification prediction, and obtain the updated feature tensor Z. At this time, the size of the updated feature tensor Z is [7×7], and its features are 7×7 = 49. Finally, the entire network is updated through backpropagation. The work flow is as Figure 3 shown.

[0069] Most neural networks will preferentially learn the most interesting features and suppress other features. The RSC module improves such a feature extraction method: mute the features related to the highest contribution degree, so that the network is forced to predict the class label by learning other features. By preventing the fully connected layer from making predictions with the most representative features (such as the most frequent colors, edges or shapes in the training data), the model is forced to use the remaining information to predict the class label. Compared with the conventional model, this model will use more features to complete the target recognition task.

[0070] Other steps and parameters are the same as those in any one of the first to third specific embodiments.

[0071] Embodiment 5: The difference between this embodiment and any one of Embodiments 1 to 4 is that the value of the hyperparameter p is 5.

[0072] Other steps and parameters are the same as those in any one of Embodiments 1 to 4.

[0073] Embodiment 6: The difference between this embodiment and any one of Embodiments 1 to 5 is that the contribution degree g z(i) has the following expression:

[0074]

[0075] where z(i) represents the feature of the i-th dimension output by the average pooling layer, is the output based on the feature of the i-th dimension by the part in the backbone network for class prediction, is the parameter of the part in the backbone network for class prediction.

[0076] Other steps and parameters are the same as those in any one of Embodiments 1 to 5.

[0077] Embodiment 7: The difference between this embodiment and any one of Embodiments 1 to 6 is that the loss function L adopted by the SDRSC target recognition model is:

[0078]

[0079] where Y is the set of input sample labels, is the class prediction result of the SDRSC target recognition model for the input sample, λ ∈ [0, ∞) is the weight decay coefficient, and ∥∥ is the norm.

[0080] The loss function in this embodiment completes the isolation of strong features and weak features through adding a regularization method and simple punishment of the output weights, thereby alleviating the gradient starvation phenomenon.

[0081] Other steps and parameters are the same as those in any one of Embodiments 1 to 6.

[0082] Experimental Results and Analysis

[0083] 1. SDRSC Network Performance Experiment

[0084] The dataset constructed in the present invention contains 8350 training set pictures and 8350 test set pictures each. According to experience, the training parameters are adjusted: batch_size is set to 32, the minimum batch size is set to 256, the initial weight is 0.001, the momentum size is set to 0.9, and the number of iterations is set to 30 times. This parameter setting is the same as the parameter setting of ResNet-50 in the comparative experiment of the backbone network. Using the SDRSC network for recognition, its loss function is as Figure 4As shown, record its accuracy rate and compare it with the recognition accuracy rate of ResNet-50 as Figure 5 shown.

[0085] It can be seen that SDRSC has a faster convergence speed, and its recognition accuracy rate always remains leading compared with ResNet-50. Finally, the recognition accuracy rate of SDRSC is 82.91%, while the recognition accuracy rate of ResNet-50 is 75.67%. It can be seen that the recognition accuracy of SDRSC has increased by 7.24% compared with the previous ResNet-50, and the network performance has been greatly improved.

[0086] 2. Ablation Experiment

[0087] The SDRSC network has greatly improved the final recognition accuracy rate compared with the original network ResNet-50 by introducing two optimized parts. In order to prove that each optimized part can effectively contribute to the improvement of the recognition accuracy rate of the data set, an ablation experiment was conducted in this section. The original network, the network with each optimized part added separately, and the network with all optimized parts added were run on the data set. Four rounds of experiments were carried out and the recognition accuracy rate of each round of experiments was recorded as the evaluation index, and the number of iterations was set to 30 times. The final results are shown in Table 2:

[0088] Table 2 Comparison of Classification Effects of Different Algorithms

[0089]

[0090] It can be observed from Table 2 that the recognition effects of each optimized part of the SDRSC network on the data set have been improved to a certain extent. Specifically, the ResNet-50 network without adding any optimized parts already has strong competitiveness on the data set with a recognition accuracy rate of 75.67%. In the second round of experiments, the loss function optimization was added to the original network ResNet-50, and the recognition accuracy rate of the network reached 79.65%, which was an increase of 3.98% compared with the accuracy rate of the original network ResNet-50, which is sufficient to prove that by punishing the output weights, the gradient starvation phenomenon can be effectively alleviated, and thus the recognition performance has been greatly improved. In the third round of experiments, the RSC module was added to the original network ResNet-50, and the recognition accuracy rate of the network reached 79.08%, which was an increase of 3.41% compared with when the module was not added, confirming that the feature extraction strategy of suppressing significant attention has better recognition for targets in complex scenes. In the fourth round of experiments, the above two optimized parts were added at the same time, and the recognition accuracy rate reached 82.91%. Compared with the two separate optimized parts, the recognition accuracy rate increased by 3.26% and 3.83% respectively, indicating that the combination of the two optimized parts has the greatest improvement on the network performance. The experiment proves that each optimized part of the SDRSC network proposed by the present invention can improve the network performance.

[0091] The SDRSC network of the present invention is compared with the existing five algorithms of SelfRag, CDANN, DRO, MMD, and IGA. The number of iterations is set to 30 times, and the iterative process on the dataset of the present invention is as follows Figure 6 shown, and the comparison of the final recognition accuracies is as follows Figure 7 shown.

[0092] (1) SelfRag: This algorithm is a regularization method for domain generalization. It only uses the self-supervision of positive data pairs to contrast the regularization loss to reduce the problems caused by negative pair sampling.

[0093] (2) CDANN: Convolutional-deconvolutional alternating neural network. It uses 3D convolutional / deconvolutional layers. By alternately connecting convolutional and deconvolutional layers, it repeatedly refines the high-level and abstract feature information of the image, and introduces residual learning to simplify the network. The network structure design of CDANN aims to achieve simple and efficient task completion.

[0094] (3) DRO: Proposed by Alibaba Cloud's AI Lab, its essence is an end-to-end deep recurrent optimizer based on neural networks, making it possible to solve some optimization problems that cannot calculate gradients. The recognition results of this algorithm on outdoor KITTI and indoor Scannet datasets have exceeded the results of all previous algorithms.

[0095] (4) MMD: The maximum-minimum distance algorithm is a heuristic clustering algorithm in pattern recognition. It is based on the Euclidean distance and takes samples that are as far apart as possible as the clustering centers. It can quickly obtain the specified k clustering centers according to the number of clusters in the dataset, and obtain samples with a relatively large distance as the initial clustering centers.

[0096] (5) IGA: The improved genetic algorithm is a bionic algorithm. It is a computer algorithm proposed by humans imitating the survival and evolution phenomena of various biological species in nature. Its biggest feature is that it can perform global search optimization and is not easily trapped in local extrema.

[0097] It can be seen from Figure 7 that the recognition accuracies of all optimization algorithms have been significantly improved compared with the ResNet-50 network, and the convergence speed is also faster; the differences in recognition accuracies and convergence speeds of the other five algorithms are relatively small, while the recognition accuracy of the algorithm of the present invention is stably better than that of other algorithms. From Figure 7It can be seen that from the perspective of the final recognition accuracy, the accuracies of the other five algorithms did not reach 80%, and the maximum accuracy gap among them was only 1.33%. The algorithm of the present invention achieved the best result with an accuracy of 82.91%, which was 2.96% higher than the accuracy of SelfReg, the best performing among the five algorithms. This accuracy gap exceeded twice the maximum accuracy gap of other algorithms. Therefore, the algorithm of the present invention has better discrimination ability compared with other mainstream algorithms.

[0098] The above examples of the present invention are only to illustrate in detail the calculation model and calculation process of the present invention, rather than to limit the implementation manner of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is impossible to list all the implementation manners here. Any obvious changes or modifications derived from the technical solution of the present invention still fall within the protection scope of the present invention.

Claims

1. An image-based aircraft target recognition method, characterized in that, the method specifically includes the following steps: Step 1: Construct a dataset of aerial aircraft target images and annotate the aircraft target images in the dataset; Step 2: Construct an SDRSC target recognition model based on ResNet-50 and use the annotated dataset to train the constructed SDRSC target recognition model; The SDRSC target recognition model is specifically: Select ResNet-50 as the backbone network, and embed the RSC module between the average pooling layer and the linear classification layer of the backbone network to obtain the SDRSC target recognition model; The working process of the RSC module is: Step 1: Calculate the contribution degree of different-dimensional features output by the average pooling layer to the classification prediction; Step 2: Initialize the hyperparameter p, and perform silence processing on the dimensional features ranked in the top p% of the contribution degree, that is, construct a vector m with the same total dimension as the output features of the average pooling layer. The i-th element m(i) in the vector m is: Among them, g z(i) is the contribution degree of the i-th dimensional feature output by the average pooling layer, and q p is the minimum contribution degree among the dimensional features whose contribution degrees rank in the top p%; Step 3: Perform Hadamard product on the features output by the average pooling layer and the vector m to obtain the updated features where, ⊙ represents the Hadamard product; Step 4, calculate the updated features The output result after passing through the softmax function: Among them, is the output of the part used for class prediction in the backbone network; According to the calculated Use the gradient descent formula to update the model parameters, where y is the sample label, is the cross-entropy loss function, is the parameter of the feature with the highest weight; Step 3: Collect the aerial aircraft target image to be recognized, and then input the aerial aircraft target image to be recognized into the SDRSC target recognition model trained in Step 2, and output the target recognition result through the SDRSC target recognition model.

2. The image-based aircraft target recognition method according to claim 1, characterized in that, the dataset constructed in Step 1 includes multi-type and multi-model aircraft target images in military scenarios and civil airliner images.

3. The image-based aircraft target recognition method according to claim 2, characterized in that, the value of the hyperparameter p is 5.

4. The image-based aircraft target recognition method according to claim 3, characterized in that, The contribution degree g z(i) has the following expression: Among them, z(i) represents the feature of the i-th dimension output by the average pooling layer, which is the output of the part for class prediction in the backbone network based on the feature of the i-th dimension, and is the parameter of the part for class prediction in the backbone network.

5. The image-based aircraft target recognition method according to claim 4, characterized in that, the loss function L adopted by the SDRSC target recognition model is: where Y is the set of input sample labels, is the class prediction result of the SDRSC target recognition model for the input sample, λ ∈ [0, ∞) is the weight decay coefficient, and ∥ ∥ is the norm.

Citation Information

Patent Citations

  • Convolutional neural network feature extraction method based on principal component analysis

    CN107844795A

  • Joint training model method and device and computer readable storage medium

    CN113268727A