Physical information guided SAR (Synthetic Aperture Radar) aircraft target detection method
Through the physical information guidance method, the Gaussian hybrid model and the target scattering key point prediction network are used, combined with the deep learning backbone network and the target detection network, the problem of similar targets and backgrounds and discrete structures in the SAR aircraft target detection is solved, and efficient and accurate target detection is achieved.
Patent Information
- Application Number
- CN202411987055.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-31
AI Technical Summary
SAR aircraft target detection method faces two main difficulties in practical application: one is that the target is similar to the surrounding background, resulting in inaccurate detection; the other is that the target structure is discrete, resulting in network error segmentation and reducing detection performance.
Using a physical information-guided method, the aircraft target slice data set is constructed, scattering key points are extracted, and the Gaussian mixed model is used to generate a probability density image as the truth heat map. Then, a target scattering key point prediction network and feature reweighting network are built, combined with deep learning backbone network and target detection network, and a joint model is formed for training to improve detection efficiency and accuracy.
It realizes complete and accurate detection of SAR aircraft targets, effectively suppresses background interference, and improves the accuracy of detection results and the completeness of targets.
Smart Images

Figure CN119942323A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a SAR aircraft target detection method guided by physical information. Background Art
[0002] Synthetic Aperture Radar (SAR) technology is one of the common remote sensing technologies. It has the advantages of all-day, all-weather, and almost unaffected by climatic conditions. It is one of the important means of earth observation. In the face of future combat scenarios, the mapping and support information required by the perception system needs to be updated in real time and quickly in the pre-war stage to adapt to the complex and real-time changes in the future battlefield environment. Synthetic aperture radar reconnaissance means are not affected by factors such as weather and lighting. The intelligent target recognition technology of SAR images has broad application prospects in modern warfare. For example, it can conduct all-day and all-weather military target reconnaissance in cloudy and rainy areas, collect intelligence information on military targets of interest, and guide combat decisions and tactical deployment.
[0003] SAR aircraft target detection for high-resolution remote sensing images is widely used in traffic management, urban planning, precision strikes on combat targets and other scenarios, and has high application value. In recent years, with the continuous improvement of the data acquisition capabilities of high-resolution SAR remote sensing platforms, SAR target detection and recognition technology based on deep learning has developed rapidly, but the current SAR aircraft target detection method still faces the following two difficulties in actual application scenarios:
[0004] 1. SAR aircraft targets are similar to the surrounding background
[0005] First, most aircraft targets in SAR images are located in airports, terminals, etc. The appearance geometry distribution of the surrounding background has a large inter-class similarity with the target, which makes it difficult for conventional target detection neural networks to distinguish between real aircraft targets and surrounding background features. The strong scattering key points of buildings distributed around the aircraft targets can easily be confused with aircraft parts, resulting in inaccurate target positioning.
[0006] 2. Discrete structure of aircraft targets in SAR images
[0007] In addition, since the imaging mechanism of SAR images is quite different from that of optical images, the aircraft targets in SAR images are a collection of a series of discrete scattering key points or scattering clusters. The discretization of the target structure makes it easy for the network to mistakenly divide a complete aircraft target into multiple ones when extracting features, which will lead to a decrease in the detection performance of the network. Summary of the invention
[0008] The purpose of the present invention is to provide a physical information guided SAR aircraft target detection method to further improve the efficiency and accuracy of SAR aircraft target detection.
[0009] In order to achieve the above tasks, the present invention adopts the following technical solutions:
[0010] A physical information guided SAR aircraft target detection method, comprising:
[0011] Step 1, construct an aircraft target slice data set; extract scattering key points of the aircraft target slices in the aircraft target slice data set; based on the Gaussian mixture model, use the scattering key points to generate a probability density image corresponding to the aircraft target slice as a true value heat map;
[0012] Step 2, construct a target scattering key point prediction network, and use the aircraft target slice data set to train the target scattering key point prediction network. During the training process, the network loss is calculated using the true value heat map corresponding to the aircraft target slice and the predicted heat map obtained by the target scattering key point prediction network.
[0013] Step 3, obtain a SAR aircraft detection image data set and construct a feature reweighted network; form a joint model with the feature reweighted network and the trained target scattering key point prediction network, and connect the joint model to the deep learning backbone network and the target detection network, so as to construct a physical information guided SAR aircraft target detection network; use the SAR aircraft detection image data set to train the SAR aircraft target detection network, and save the trained SAR aircraft target detection network model for identifying SAR aircraft target detection images of unknown target categories; wherein the feature reweighted network includes a compression module and an iterative feature enhancement module;
[0014] The SAR aircraft detection image is respectively input into the deep learning backbone network and the target scattering key point prediction network for feature extraction to obtain the corresponding depth feature map and prediction heat map. The depth feature map and the prediction heat map are compressed by the compression module respectively, and then enter the iterative feature enhancement module for feature enhancement. The enhanced features are then fused with the depth feature map to obtain the detection features and input into the target detection network to obtain the classification and detection results of the aircraft target.
[0015] Furthermore, the extraction of scattering key points of the aircraft target slices in the aircraft target slice data set first generates a multi-layer scale space by convolving the Gaussian kernel function with the aircraft target slices, and searches for candidate corner points in each layer of the scale space;
[0016] After obtaining the candidate corner points, the LOG operator is used to screen the candidate corner points in different scale spaces, and the iterative method is used to check whether the LOG operation value of the candidate corner point in each scale space is the extreme point in all scale spaces. If so, the candidate corner point is considered to be the final corner point; finally, the set of all corner points retained by the aircraft target slice is the scattering key point of the aircraft target slice.
[0017] Furthermore, the method of generating a probability density image corresponding to the aircraft target slice as a true value heat map based on the Gaussian mixture model using scattering key points includes:
[0018] Use K-means clustering method to generate the initial sub-distribution parameters of the mixed Gaussian model, including the mean, covariance matrix and weight coefficient of the cluster sub-distribution;
[0019] Calculate the posterior probability γ that each scattering key point i belongs to the kth sub-Gaussian distribution ik , the calculation formula is as follows:
[0020]
[0021] Where K represents the sub-distribution N(x i |u k ,Cov k ), u k represents the mean of the kth sub-distribution, and Cov k represents the covariance matrix corresponding to the kth sub-distribution, α k is the weight coefficient corresponding to the kth sub-distribution in the Gaussian mixture model, which is determined by the number of samples divided into the sub-distribution. j represents the i-th sub-distribution, i represents the i-th scattering key point, and x i Represents the coordinates of the i-th scattering key point;
[0022] Update the parameters in the posterior probability formula. The update formula for each parameter is as follows:
[0023]
[0024] Where N represents the number of scattering key points, and the superscript T represents the transposition operation;
[0025] Iterate the update of the parameters until the parameter u k , Cov k and α k The updated values of are all less than the corresponding preset values, or the preset iteration stop round is reached;
[0026] Then according to the parameters u of the K sub-distributions k , Cov k and α kThe probability density function p(X|u,Cov) for constructing the Gaussian mixture model is:
[0027]
[0028] Where X represents the scattering key point, μ and Cov are the mean vector and covariance matrix of the Gaussian mixture model;
[0029] By using the probability density function, a probability density image corresponding to each aircraft target slice, that is, a true value heat map, is obtained.
[0030] Furthermore, the target scattering key point prediction network is constructed, and the target scattering key point prediction network is trained using the aircraft target slice data set, including:
[0031] First, the aircraft target slices are preprocessed to adjust the size, and the preprocessed aircraft target slices and the corresponding true value heat map are input into the target scattering key point prediction network; the target scattering key point prediction network includes a backbone part and a key point feature learning part, in which:
[0032] The preprocessed aircraft target slices first pass through the backbone part, which first uses two convolutional layers to reduce the image resolution by a quarter, and then uses four BottleNeck layers for preliminary feature extraction. The extracted feature map is input into the key point feature learning part;
[0033] The key point feature learning part includes four sub-networks. The first to fourth sub-networks contain 10, 8, 5, and 2 convolutional layers connected in sequence, respectively. The resolution of the feature map in the second sub-network is half of that in the first sub-network, but the number of channels is doubled.
[0034] The feature map A extracted from the main part is input into the first sub-network, and passes through each convolution layer in the sub-network in sequence to extract the feature map, and outputs the feature map A1; the feature map A extracted from the main part is input into the second sub-network after downsampling, and passes through each convolution layer in the sub-network in sequence to extract the feature map, and outputs the feature map A2; the feature map output by the second convolution layer of the second sub-network is input into the third sub-network after downsampling, and passes through each convolution layer in the sub-network in sequence to extract the feature map, and outputs the feature map A3; the feature map output by the second convolution layer of the third sub-network is input into the fourth sub-network after downsampling, and passes through each convolution layer in the sub-network in sequence to extract the feature map, and outputs the feature map A4;
[0035] Among them, the feature map output by the jth convolutional layer of the i-th layer network is recorded as a ij , then the following relationship exists:
[0036] a 12After downsampling and 22 After addition, it enters the third convolutional layer of the second sub-network; and a 21 After upsampling and combining with a 13 After adding, it enters the fourth convolutional layer of the first sub-network;
[0037] a 24 、a 31 They are upsampled and then combined with a 16 Add, and then enter the seventh convolutional layer of the first sub-network; a 15 After downsampling, a 31 The result after upsampling is similar to a 25 Added, and then enter the sixth convolutional layer of the second sub-network; a 15 、a 24 The results after downsampling are respectively 32 Add them together and then enter the third convolutional layer of the third sub-network;
[0038] a 27 、a 34 、a 41 After upsampling, 19 Add, and then enter the tenth convolutional layer of the first sub-network; a 18 After downsampling, a 34 After upsampling, a 41 The result after upsampling is similar to a 28 Add them together and the result is A2; a 18 After downsampling, a 27 After downsampling, a 41 The result after upsampling is similar to a 35 Add them together and the result is A3; a 18 After downsampling, a 27 After downsampling, a 34 The result after downsampling is similar to a 42 Add them together and the result is A4;
[0039] The final predicted heat map is the feature map of A2, A3, and A4 after upsampling and A1 added together.
[0040] Furthermore, the network loss is calculated by using the true value heat map corresponding to the aircraft target slice and the predicted heat map obtained by the aircraft target slice predicted by the target scattering key point prediction network, including:
[0041] Construct a loss function based on the predicted heat map and the true value heat map to calculate the network loss during training; the loss function is as follows:
[0042]
[0043] Where L heatmap represents the network loss, N represents the total number of pixels in the true value heat map, and y wh , They respectively represent the true value heat map corresponding to the aircraft target slice input to the network and the value of the pixel with coordinates (w, h) in the predicted heat map obtained by the network prediction. W and H represent the total number of pixels in the length and width directions.
[0044] Furthermore, the deep learning backbone network adopts ResNet18 or VGG16; the target detection network adopts YOLOv5.
[0045] Furthermore, the compression module includes a mean pooling layer and a maximum pooling layer. The depth feature map extracted by the deep learning backbone network of the SAR aircraft detection image and the predicted heat map extracted by the target scattering key point prediction network are processed by the mean pooling layer and the maximum pooling layer respectively, and the two pooled images processed by the mean pooling layer and the maximum pooling layer are weighted fused according to pre-set weights. The fusion process is to add each pixel element of the pooled image to obtain the corresponding backbone compression feature map and the key point compression feature map, and input them into the iterative feature enhancement module.
[0046] Furthermore, the processing process of the iterative feature enhancement module is:
[0047] The backbone compressed feature map is projected to the key value matrix, and the key point compressed feature map is projected to the query matrix. The query matrix of the key point compressed feature map and the key value matrix of the backbone compressed feature map are used to perform cross attention calculation to obtain the feature Z representing the correlation between the two. Then, the feature Z is reprojected through a linear layer, and the projected result is compared with the backbone compressed feature map according to the weight φ. Connect and fuse the complementary information from different features to obtain the complementary feature Z a ; complementary feature Z a The result after the feedforward network processing is then combined with the complementary feature Z a Connect according to the weights η and ε respectively to obtain the enhanced feature Z b ; Enhance feature Z b The deep feature map extracted by the deep learning backbone network is added element by element to obtain the final detection feature Z c , Z c Input into the target detection network to obtain the detection result of the aircraft target.
[0048] A terminal device comprises a processor, a memory and a computer program stored in the memory; when the processor executes the computer program, the SAR aircraft target detection method guided by physical information is implemented.
[0049] A computer-readable storage medium stores a computer program; when the computer program is executed by a processor, the physical information-guided SAR aircraft target detection method is implemented.
[0050] Compared with the prior art, the present invention has the following technical features:
[0051] 1. In view of the fact that the strong scattering key points of SAR aircraft targets are usually distributed in key parts such as the nose and wings, and contain a large amount of structural feature information, the present invention first extracts the scattering key points of the target, and in order to further represent the extracted scattering key points as key cluster perceptible features of the SAR aircraft target structure, a Gaussian mixture model is used to cluster the scattering key points; the aircraft target slice is first passed through a backbone network to reduce the image resolution by one quarter to obtain basic features, and then the basic features are extracted through a subsequent multi-scale feature pyramid network HRNet, and the predicted heat map of the aircraft target slice is output; and the loss function of the predicted heat map and the true value heat map is calculated for evaluation.
[0052] 2. For the prediction heat map of the target scattering key point prediction network, the present invention designs a feature reweighted network based on the cross-attention mechanism guided by physical information; first, the two input features are compressed by information dimension pooling through the spatial feature shrinkage module, and the compressed features are interactively fused through the iterative cross-modal feature enhancement module; finally, the feedforward network is used to further refine the feature representation to improve the robustness and accuracy of the model.
[0053] Based on the above design, this method can, on the one hand, completely and accurately detect SAR aircraft targets, and on the other hand, can effectively suppress background interference similar to the target in the SAR image. It has the characteristics of good target detection completeness and high detection result accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 It is a flow chart of the SAR aircraft target detection algorithm guided by physical information of the present invention;
[0055] Figure 2 This is a flow chart of the SAR aircraft target slice data set training of the present invention;
[0056] Figure 3 This is a schematic diagram of the structure of the target scattering key point prediction network of the present invention;
[0057] Figure 4 It is a schematic diagram of the structure of the up- and down-sampling network fusion method of the present invention.
[0058] Figure 5 This is a schematic diagram of the structure of the feature re-weighted network of the present invention. DETAILED DESCRIPTION
[0059] In view of the problems existing in the prior art, the present invention provides a SAR aircraft target detection method guided by physical information, adopts physical knowledge such as target scattering cluster model to uniformly describe and quantitatively express the characteristics of SAR aircraft targets; based on the physical characteristics of SAR targets, constructs a high-resolution feature pyramid-style scattering key point prediction network self-supervised learning framework guided by physical characteristics; constructs a SAR aircraft detection and recognition hybrid neural network model coupled with physical knowledge, etc., completes the research on the intelligent interpretation technology of interpretable SAR targets based on physical characteristics, completes the scattering key point prediction network algorithm experiment with SAR aircraft target slices constituting training data, and the SAR aircraft target detection and recognition experiment with SAR aircraft detection images constituting training data, and forms a complete physical characteristic combined SAR aircraft intelligent detection and recognition algorithm software with SAR target physical characteristic extraction algorithm, SAR target self-supervised feature learning algorithm, physical characteristic combined intelligent recognition algorithm, SAR aircraft detection and recognition algorithm and other unit modules. It can achieve robust detection of aircraft targets in complex environments, improve SAR aircraft target detection and recognition performance, and improve algorithm transparency and credibility.
[0060] The present invention provides a physical information guided SAR aircraft target detection method, comprising the following steps:
[0061] Step 1, construct an aircraft target slice dataset; extract scattering key points of the aircraft target slices in the aircraft target slice dataset; based on the Gaussian mixture model (GMM), use the scattering key points to generate a probability density image corresponding to the aircraft target slice as a true value heat map.
[0062] Firstly, we select multiple clear aircraft target slices from the public SAR aircraft dataset. Each aircraft target slice contains a complete aircraft target, thus constructing an aircraft target slice dataset.
[0063] In this embodiment, 280, 131, 443 and 170 slices of aircraft target are screened from SADD, TerraSAR-X, Multiangle SAR Dataset and SAR-CAD dataset respectively, and finally 1024 slices of aircraft target are obtained.
[0064] Secondly, extract the scattering key points of the aircraft target slice, see Appendix Figure 2 , as follows:
[0065] (1) Extract the corner points of the aircraft target slice.
[0066] First, the Gaussian kernel function is convolved with the aircraft target slice to generate the scale space, and the candidate corner points are found in each layer of the scale space: This method generates the scale space by convolving the Gaussian kernel function with different scale parameters with the aircraft target slice, and the expression is:
[0067]
[0068] Where L(x,y,σ) is the scale space, Represents the convolution operation, I(x,y) is the input aircraft target slice, G(x,y,σ) is the Gaussian kernel function x and y are the horizontal and vertical coordinates of the pixel position of the aircraft target slice, and σ is the scale parameter.
[0069] Inspired by the autocorrelation function of signal processing, the eigenvalue of the autocorrelation matrix M is the first-order curvature of the autocorrelation function. If the two first-order curvature values of a pixel are both large, it is a corner point. The autocorrelation matrix M is extended to the scale space and expressed as follows:
[0070]
[0071] In the formula, g(σ I ) is a Gaussian function, Q x ,Q y Respectively represent the scale space in the x and y directions and σ = σ D The differential is calculated by multiplying the Gaussian convolution kernel of D is the differential scale, Used to offset the scaling of the Gaussian convolution kernel by the differential scale; σ I It is the integral scale, which is used to adjust the scale of the Gaussian convolution kernel. The larger the value, the larger the corresponding scale.
[0072] Define the corner response function:
[0073] R=det(M)-αtrace(M) 2
[0074] Wherein, R is the corner point response value, det is the determinant of the matrix, trace is the trace of the matrix, α is a constant, which usually takes a value in the range of [0.04, 0.06]. In this embodiment, the value of α is 0.04.
[0075] The corner point response value is calculated for each pixel in a single scale space. When the corner point response value exceeds a preset threshold, the pixel is judged as a candidate corner point. In this way, the candidate corner points in a single scale space are obtained. However, such candidate corner points are not scale invariant, so it is necessary to establish multiple scale spaces and search for candidate corner points in different spaces.
[0076] By predefining a set of σ={σ1,σ2,σ3,σ4,...,σ n},σ n is the nth scale parameter, and a set of scale spaces under Gaussian kernel functions of different scales for the slice of the aircraft target is obtained. In this embodiment, n is set to 3. The candidate corner points in each scale space are extracted according to the above method.
[0077] (2) Since there are a lot of redundancies or detection errors in the candidate corner points obtained in multiple scale spaces, the LOG operator (Laplacian of Gaussian) is introduced to screen the candidate corner points in different scale spaces. The iterative method is used to check whether the LOG operation value of the candidate corner point in each scale space is the extreme point in all scale spaces. If not, it is discarded. If so, the candidate corner point is considered to be the final corner point.
[0078] (3) The set of all corner points retained by the final aircraft target slice is the scattering key points of the aircraft target slice.
[0079] In step 1.4, the scattering key points extracted from each aircraft target slice are clustered using a Gaussian mixture model to obtain a probability density image of the clustering of the scattering key points of the aircraft target slice as a true value heat map.
[0080] (1) Parameter initialization settings:
[0081] The K-means clustering method is used to generate the initial sub-distribution parameters of the mixed Gaussian model, including the mean, covariance matrix and weight coefficient of the cluster sub-distribution. In this embodiment, the number of cluster centers is 9.
[0082] For the scattering key points of the aircraft target slice, the mean of each sub-distribution is the coordinate mean of the scattering key points in the class, the covariance matrix is calculated by the covariance of the scattering key points, and the weight coefficient is determined by the number of scattering key points in each sub-distribution.
[0083] (2) Calculate the posterior probability:
[0084] Calculate the posterior probability γ that each scattering key point i belongs to the kth sub-Gaussian distribution ik , the calculation formula is as follows;
[0085]
[0086] Where K represents the sub-distribution N(x i |u k ,Cov k ), u k represents the mean of the kth sub-distribution, and Cov krepresents the covariance matrix corresponding to the kth sub-distribution, α k is the weight coefficient corresponding to the kth sub-distribution in the Gaussian mixture model, which is determined by the number of samples divided into the sub-distribution. j represents the i-th sub-distribution, i represents the i-th scattering key point, and x i Represents the coordinates of the i-th scattering keypoint.
[0087] (3) Update the parameters in the posterior probability formula. The update formula for each parameter is as follows:
[0088]
[0089] Where N represents the number of scattering key points, and the superscript T represents the transpose operation.
[0090] (4) Repeated iteration
[0091] Repeat steps (2) and (3) until the parameter u k , Cov k and α k The updated values of are all less than the corresponding preset values, or the preset iteration stop rounds are reached; in this embodiment, the preset value is 0.01, and the iteration rounds are 50. The mean, covariance matrix and weight coefficient of the updated sub-distribution are obtained.
[0092] Through the above method, we can estimate the Gaussian sub-distribution to which each scattering key point belongs, as well as the mean u of the sub-distribution corresponding to the Gaussian sub-distribution. k , covariance matrix Cov k and weight coefficient α k , and then according to the parameters u of the K sub-distributions k , Cov k and α k The probability density function p(X|u,Cov) for constructing the Gaussian mixture model is:
[0093]
[0094] Among them, X represents the scattering key points, μ and Cov are the mean vector and covariance matrix of the Gaussian mixture model.
[0095] Using the probability density function, the probability density image corresponding to each aircraft target slice is obtained, that is, the true value heat map, such as Figure 2 As shown, the ground truth heatmap size is one-fourth of the aircraft target slice size.
[0096] Step 2: construct a target scattering key point prediction network and use the aircraft target slice data set to train the target scattering key point prediction network. During the training process, the network loss is calculated using the true value heat map corresponding to the aircraft target slice and the predicted heat map obtained by the target scattering key point prediction network.
[0097] Step 2.1, when using the aircraft target slice data set for network training, the aircraft target slice is first preprocessed to adjust the size; during the preprocessing, the original aspect ratio of the image is kept unchanged, and the remaining part is padded with 0; in this embodiment, the adjusted size is (256, 192).
[0098] In step 2.2, the preprocessed aircraft target slices and the corresponding true value heat map are input into the target scattering key point prediction network to learn the scattering key point feature distribution of the aircraft target slices.
[0099] The target scattering key point prediction network provided by the present invention is as follows Figure 3 As shown, it includes the backbone part and the key point feature learning part, where:
[0100] (1) Main part
[0101] The preprocessed aircraft target slices first pass through the backbone part, which first uses two 3*3 convolutional layers to reduce the image resolution by a quarter to match the resolution size of the real heat map. Then four BottleNeck layers are used for preliminary feature extraction, and the extracted feature map is input into the key point feature learning part. In this embodiment, the size of the feature map output by the backbone part and the real heat map are both (32, 24) to facilitate subsequent loss calculation.
[0102] (2) Key point feature learning part
[0103] The key point feature learning part includes four sub-networks. The first to fourth sub-networks contain 10, 8, 5, and 2 convolutional layers connected in sequence respectively. The resolution of the feature map in the second sub-network is half of that in the first sub-network, but the number of channels is doubled. The specific design is as follows:
[0104] The feature map A extracted from the main part is input into the first sub-network, and passes through each convolution layer in the sub-network in sequence to extract the feature map, and outputs the feature map A1; the feature map A extracted from the main part is input into the second sub-network after downsampling, and passes through each convolution layer in the sub-network in sequence to extract the feature map, and outputs the feature map A2; the feature map output by the second convolution layer of the second sub-network is input into the third sub-network after downsampling, and passes through each convolution layer in the sub-network in sequence to extract the feature map, and outputs the feature map A3; the feature map output by the second convolution layer of the third sub-network is input into the fourth sub-network after downsampling, and passes through each convolution layer in the sub-network in sequence to extract the feature map, and outputs the feature map A4;
[0105] Among them, the feature map output by the jth convolutional layer of the i-th layer network is recorded as a ij , then the following relationship exists:
[0106] a 12 After downsampling and 22 After addition, it enters the third convolutional layer of the second sub-network; and a 21 After upsampling and combining with a 13 After adding, it enters the fourth convolutional layer of the first sub-network;
[0107] a 24 、a 31 They are upsampled and then combined with a 16 Add, and then enter the seventh convolutional layer of the first sub-network; a 15 After downsampling, a 31 The result after upsampling is similar to a 25 Added, and then enter the sixth convolutional layer of the second sub-network; a 15 、a 24 The results after downsampling are respectively 32 Add them together and then enter the third convolutional layer of the third sub-network;
[0108] a 27 、a 34 、a 41 After upsampling, 19 Add, and then enter the tenth convolutional layer of the first sub-network; a 18 After downsampling, a 34 After upsampling, a 41 The result after upsampling is similar to a 28 Add them together and the result is A2; a 18 After downsampling, a 27 After downsampling, a 41 The result after upsampling is similar to a 35 Add them together and the result is A3; a 18 After downsampling, a27 After downsampling, a 34 The result after downsampling is similar to a 42 Add them together and the result is A4.
[0109] The final predicted heat map is the feature map of A2, A3, and A4 after upsampling and A1 added together.
[0110] In the multi-layer sub-network structure designed in this scheme, the resolution of each sub-network is half of that of the previous layer, but the number of channels is doubled; sub-networks with different resolutions are connected in parallel and fused at different convolutional layers; the specific fusion method is as follows: Figure 4 As shown in the figure, when the low-resolution feature map is fused into the high-resolution feature map, an upsampling module is added to enlarge the feature map for every half difference in resolution, and it is added element by element to the high-resolution feature map; similarly, when the feature map is fused from high to low, a downsampling module is added for every half difference in resolution. This module is implemented by a 3*3 convolutional layer with a step size of 2, and the reduced feature map is added element by element to the feature map to be fused; finally, the feature maps output by the last three sub-networks are upsampled and added to the feature map of the first sub-network to obtain the final predicted heat map.
[0111] Step 2.3, construct a loss function based on the predicted heat map and the true value heat map to calculate the network loss during training; the loss function is as follows:
[0112]
[0113] Where L heatmap represents the network loss, N represents the total number of pixels in the true value heat map, and y wh , They represent the true heat map corresponding to the slice of the aircraft target input to the network and the value of the pixel with coordinates (w, h) in the predicted heat map obtained by the network prediction, respectively. W and H represent the total number of pixels in the length and width directions. This loss function aims to evaluate and constrain the learning effect of the network on key point features by calculating the mean square difference of each pixel between the true heat map and the predicted heat map.
[0114] After network training, when the network loss converges, a trained target scattering key point prediction network is obtained, which can achieve good prediction of the key point features in the aircraft target slice, and the prediction result is the corresponding prediction heat map.
[0115] Step 3, obtain a SAR aircraft detection image data set and construct a feature reweighted network; form a joint model with the feature reweighted network and the trained target scattering key point prediction network, and connect the joint model to the deep learning backbone network and the target detection network, so as to construct a physical information guided SAR aircraft target detection network; use the SAR aircraft detection image data set to train the SAR aircraft target detection network, and save the trained SAR aircraft target detection network model for identifying SAR aircraft target detection images of unknown target categories; wherein the feature reweighted network includes a compression module and an iterative feature enhancement module;
[0116] The SAR aircraft detection images in the SAR aircraft detection image dataset are respectively input into the backbone network of the target detection network and the target scattering key point prediction network for feature extraction to obtain the corresponding depth feature map and prediction heat map. The depth feature map and the prediction heat map are compressed by the compression module respectively, and then enter the iterative feature enhancement module for feature enhancement. The enhanced features are then input into the target detection network to obtain the classification and detection results of the aircraft target.
[0117] In this step, the SAR aircraft detection images in the SAR aircraft detection image dataset are real SAR images taken for aircraft targets, which are different from the aircraft target slices in step 1; the SAR aircraft detection images may contain one or more aircraft targets. Although the target scattering key point prediction network is trained for aircraft target slices, its network parameters also have a good key point detection effect for the complete SAR aircraft detection image. In this scheme, the SAR aircraft detection image dataset is selected from the public dataset SAR-Aircraft1.0, and the ratio of the training set and the test set is consistent with the original dataset.
[0118] (1) Deep Learning Backbone Network
[0119] The deep learning backbone network in this solution can use an existing network structure, such as ResNet18 or VGG16, to process the input SAR aircraft detection image to obtain a deep feature map. In this example, the deep learning backbone network is ResNet18.
[0120] (2) Compression module
[0121] In order to reduce the computational complexity of subsequent modules and minimize the loss of key point feature information, the spatial feature compression module compresses the deep feature map extracted by the backbone network and the predicted heat map extracted by the target scattering key point prediction network, and combines average pooling and maximum pooling to adaptively aggregate feature information; average pooling retains background information, and maximum pooling retains texture features, and the final compressed feature map is obtained by weighted summation; through the spatial feature compression module, the spatial dimension of the feature map can be significantly reduced, thereby reducing the computational complexity of subsequent modules.
[0122] Specifically, Figure 5 As shown in FIG. 1 , the compression module includes a mean pooling layer and a maximum pooling layer. The deep feature map extracted by the deep learning backbone network of the SAR aircraft detection image and the predicted heat map extracted by the target scattering key point prediction network are processed by the mean pooling layer and the maximum pooling layer respectively. The two pooled images processed by the mean pooling layer and the maximum pooling layer are weighted fused according to the preset weight λ. The fusion process is to add each pixel element of the pooled image to obtain the corresponding backbone compression feature map and key point compression feature map, and input them into the iterative feature enhancement module. In this example, the value of λ is 0.5.
[0123] (3) Iterative feature enhancement module
[0124] The iterative feature enhancement module continuously strengthens the cross-modal feature information in an iterative manner, thereby improving the discriminability of feature representation.
[0125] The processing process of the iterative feature enhancement module is:
[0126] The backbone compressed feature map is projected to the key value (K, V) matrix, while the key point compressed feature map is projected to the query Q matrix. The query matrix Q of the key point compressed feature map and the key value matrix (K, V) of the backbone compressed feature map are used to perform cross attention calculation to obtain the feature Z representing the correlation between the two. Then, the feature Z is reprojected through a linear layer, and the projected result is compared with the backbone compressed feature map according to the weight φ. Connect and fuse the complementary information from different features to obtain the complementary feature Z a , thereby improving the robustness and accuracy of the model; complementary feature Z a The result after the feedforward network FNN processing is then combined with the complementary feature Z a Connect according to the weights η and ε respectively to obtain the enhanced feature Z b , thereby improving the robustness and accuracy of the model; enhancing feature Z b The deep feature map extracted by the deep learning backbone network is added element by element to obtain the final detection feature Z c .
[0127] Among them, φ, η, ε are both learnable parameters in the training process. The initialization values of η and ε are both 0.5.
[0128] (4) Object Detection Network
[0129] The target detection network in this scheme is used to detect the target in the feature map. It can use existing networks, such as YOLOv5, etc., to convert the final detection feature Z c Input into the target detection network to obtain the classification and detection results of the aircraft target. The target detection network in this solution uses YOLOv5.
[0130] Example:
[0131] In one embodiment of the present invention, the Pytorch framework is used for simulation on a CPU of Intel(R) i9-10900X 3.7GHz CPU, 256G memory, 4 Nvidia GTX3090 GPUs, and Ubuntu 18.04 operating system; during the experiment, the ratio of the training set to the test set of the SAR aircraft detection image is set according to SAR-Aircraft1.0. The test results are shown in Table 1; from the test results, it can be seen that the detection method proposed by the present invention has significantly improved the recognition rate of different targets compared with the traditional network.
[0132] Table 1 Verification results of SAR aircraft detection method guided by physical information
[0133]
[0134] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A physical information guided SAR aircraft target detection method, characterized in that: include: Step 1, construct an aircraft target slice dataset; Extracting scattering key points of aircraft target slices in the aircraft target slice data set; Based on the Gaussian mixture model, the probability density image corresponding to the aircraft target slice is generated using the scattering key points as the true value heat map; Step 2, construct a target scattering key point prediction network, and use the aircraft target slice data set to train the target scattering key point prediction network. During the training process, the network loss is calculated using the true value heat map corresponding to the aircraft target slice and the predicted heat map obtained by the target scattering key point prediction network. Step 3, obtain the SAR aircraft detection image dataset and construct a feature reweighted network; The feature reweighting network and the trained target scattering key point prediction network are combined into a joint model, and the joint model is connected to the deep learning backbone network and the target detection network to construct a SAR aircraft target detection network guided by physical information; the SAR aircraft target detection network is trained using the SAR aircraft detection image dataset, and the trained SAR aircraft target detection network model is saved for identifying SAR aircraft target detection images of unknown target categories; wherein the feature reweighting network includes a compression module and an iterative feature enhancement module; The SAR aircraft detection image is respectively input into the deep learning backbone network and the target scattering key point prediction network for feature extraction to obtain the corresponding depth feature map and prediction heat map. The depth feature map and the prediction heat map are compressed by the compression module respectively, and then enter the iterative feature enhancement module for feature enhancement. The enhanced features are then fused with the depth feature map to obtain the detection features and input into the target detection network to obtain the classification and detection results of the aircraft target.
2. The physical information guided SAR aircraft target detection method according to claim 1, characterized in that: The scattering key points of the aircraft target slices in the aircraft target slice data set are extracted by first convolving the Gaussian kernel function with the aircraft target slices to generate a multi-layer scale space, and searching for candidate corner points in each layer of the scale space; After obtaining the candidate corner points, the LOG operator is used to screen the candidate corner points in different scale spaces, and the iterative method is used to check whether the LOG operation value of the candidate corner point in each scale space is the extreme point in all scale spaces. If so, the candidate corner point is considered to be the final corner point; finally, the set of all corner points retained by the aircraft target slice is the scattering key point of the aircraft target slice.
3. The physical information guided SAR aircraft target detection method according to claim 1, characterized in that: The method of generating a probability density image corresponding to an aircraft target slice as a true value heat map based on a Gaussian mixture model using scattering key points includes: Use K-means clustering method to generate the initial sub-distribution parameters of the mixed Gaussian model, including the mean, covariance matrix and weight coefficient of the cluster sub-distribution; Calculate the posterior probability γ that each scattering key point i belongs to the kth sub-Gaussian distribution ik , the calculation formula is as follows: Where K represents the sub-distribution N(x i |u k ,Cov k ), u k represents the mean of the kth sub-distribution, and Cov k represents the covariance matrix corresponding to the kth sub-distribution, α k is the weight coefficient corresponding to the kth sub-distribution in the Gaussian mixture model, which is determined by the number of samples divided into the sub-distribution. j represents the i-th sub-distribution, i represents the i-th scattering key point, and x i Represents the coordinates of the i-th scattering key point; Update the parameters in the posterior probability formula. The update formula for each parameter is as follows: Where N represents the number of scattering key points, and the superscript T represents the transposition operation; Iterate the update of the parameters until the parameter u k , Cov k and α k The updated values of are all less than the corresponding preset values, or the preset iteration stop round is reached; Then according to the parameters u of the K sub-distributions k , Cov k and α k The probability density function p(X|u,Cov) for constructing the Gaussian mixture model is: Where X represents the scattering key point, μ and Cov are the mean vector and covariance matrix of the Gaussian mixture model; By using the probability density function, a probability density image corresponding to each aircraft target slice, that is, a true value heat map, is obtained.
4. The physical information guided SAR aircraft target detection method according to claim 1, characterized in that: The target scattering key point prediction network is constructed, and the target scattering key point prediction network is trained using the aircraft target slice data set, including: First, the aircraft target slices are preprocessed to adjust the size, and the preprocessed aircraft target slices and the corresponding true value heat map are input into the target scattering key point prediction network; the target scattering key point prediction network includes a backbone part and a key point feature learning part, in which: The preprocessed aircraft target slices first pass through the backbone part, which first uses two convolutional layers to reduce the image resolution by a quarter, and then uses four BottleNeck layers for preliminary feature extraction. The extracted feature map is input into the key point feature learning part; The key point feature learning part includes four sub-networks. The first to fourth sub-networks contain 10, 8, 5, and 2 convolutional layers connected in sequence, respectively. The resolution of the feature map in the second sub-network is half of that in the first sub-network, but the number of channels is doubled. The feature map A extracted from the main part is input into the first sub-network, and passes through each convolution layer in the sub-network in sequence to extract the feature map, and outputs the feature map A1; the feature map A extracted from the main part is input into the second sub-network after downsampling, and passes through each convolution layer in the sub-network in sequence to extract the feature map, and outputs the feature map A2; the feature map output by the second convolution layer of the second sub-network is input into the third sub-network after downsampling, and passes through each convolution layer in the sub-network in sequence to extract the feature map, and outputs the feature map A3; the feature map output by the second convolution layer of the third sub-network is input into the fourth sub-network after downsampling, and passes through each convolution layer in the sub-network in sequence to extract the feature map, and outputs the feature map A4; Among them, the feature map output by the jth convolutional layer of the i-th layer network is recorded as a ij , then the following relationship exists: a 12 After downsampling and 22 After addition, it enters the third convolutional layer of the second sub-network; and a 21 After upsampling and combining with a 13 After adding, it enters the fourth convolutional layer of the first sub-network; a 24 、a 31 They are upsampled and then combined with a 16 Add, and then enter the seventh convolutional layer of the first sub-network; a 15 After downsampling, a 31 The result after upsampling is similar to a 25 Added, and then enter the sixth convolutional layer of the second sub-network; a 15 、a 24 The results after downsampling are respectively 32 Add them together and then enter the third convolutional layer of the third sub-network; a 27 、a 34 、a 41 After upsampling, 19 Add, and then enter the tenth convolutional layer of the first sub-network; a 18 After downsampling, a 34 After upsampling, a 41 The result after upsampling is similar to a 28 Add them together and the result is A2; a 18 After downsampling, a 27 After downsampling, a 41 The result after upsampling is similar to a 35 Add them together and the result is A3; a 18 After downsampling, a 27 After downsampling, a 34 The result after downsampling is similar to a 42 Add them together and the result is A4; The final predicted heat map is the feature map of A2, A3, and A4 after upsampling and A1 added together.
5. The physical information guided SAR aircraft target detection method according to claim 1, characterized in that: The network loss is calculated by using the true value heat map corresponding to the aircraft target slice and the predicted heat map obtained by the aircraft target slice through the target scattering key point prediction network, including: Construct a loss function based on the predicted heat map and the true value heat map to calculate the network loss during training; the loss function is as follows: Where L heatmap represents the network loss, N represents the total number of pixels in the true value heat map, and y wh , They respectively represent the true value heat map corresponding to the aircraft target slice input to the network and the value of the pixel with coordinates (w, h) in the predicted heat map obtained by the network prediction. W and H represent the total number of pixels in the length and width directions.
6. The physical information guided SAR aircraft target detection method according to claim 1, characterized in that: The deep learning backbone network uses ResNet18 or VGG16; the target detection network uses YOLOv5.
7. The physical information guided SAR aircraft target detection method according to claim 1, characterized in that: The compression module includes a mean pooling layer and a maximum pooling layer. The depth feature map extracted by the deep learning backbone network of the SAR aircraft detection image and the predicted heat map extracted by the target scattering key point prediction network are processed by the mean pooling layer and the maximum pooling layer respectively, and the two pooled images processed by the mean pooling layer and the maximum pooling layer are weighted fused according to the preset weights. The fusion process is to add each pixel element of the pooled image to obtain the corresponding backbone compression feature map and the key point compression feature map, and input them into the iterative feature enhancement module.
8. The physical information guided SAR aircraft target detection method according to claim 1, characterized in that: The processing process of the iterative feature enhancement module is: The backbone compressed feature maps are projected to the key value matrix, and the key point compressed feature map is projected to the query matrix. The query matrix of the key point compressed feature map and the key value matrix of the backbone compressed feature map are used to perform cross attention calculation to obtain the feature Z representing the correlation between the two. Then The feature Z is reprojected through a linear layer, and the projected result is passed through the backbone compressed feature map according to the weight φ. Connect and fuse the complementary information from different features to obtain the complementary feature Z a ; complementary feature Z a The result after the feedforward network processing is then combined with the complementary feature Z a Connect according to the weights η and ε respectively to obtain the enhanced feature Z b ; Enhance feature Z b The deep feature map extracted by the deep learning backbone network is added element by element to obtain the final detection feature Z c , Z c Input into the target detection network to obtain the detection result of the aircraft target.
9. A terminal device comprising a processor, a memory and a computer program stored in the memory; characterized in that: When the processor executes the computer program, the SAR aircraft target detection method guided by physical information according to any one of claims 1 to 8 is implemented.
10. A computer-readable storage medium, wherein a computer program is stored in the medium; characterized in that: When the computer program is executed by a processor, the SAR aircraft target detection method guided by physical information according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Lightweight SAR image ship target inclined frame detection method and system
CN114445721A
Multi-scale SAR (Synthetic Aperture Radar) image target detection method, device, equipment and medium
CN115272859A