Small sample-oriented SAR image vehicle target detection and identification method

The SAR image samples are generated through the region growth algorithm and the vehicle object detection model of hollow convolution and dynamic convolution layer is solved, and data scarcity and similarity recognition problems in SAR image vehicle object detection recognition are achieved, high-precision vehicle object detection and improved recognition effect.

CN120355902APending Publication Date: 2025-07-22XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510492263.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

SAR image vehicle target detection and recognition faces the problems of scarce data sets, complex background distribution and high difficulty in precise identification. Especially in complex terrestrial environments, vehicle targets are disturbed by imaging angles, surface coverings and surrounding artificial facilities, resulting in weak echo signals and blurred image features, increasing the risk of missed detection and false alarms. In addition, the high similarity of vehicles of the same type but different models in SAR images puts strict requirements on the recognition algorithm.

Method used

The region growth algorithm is used to generate SAR image samples, and combined with the hollow convolution layer and the full-dimensional dynamic convolution layer sampled in parallel with multiple sampling rates, a vehicle object detection model is built. Through feature fusion pyramid network and decoupled detection head, multi-scale information and detailed target structure characteristics are extracted, and the network's ability to distinguish similar targets is enhanced.

Benefits of technology

The high-precision detection and recognition of vehicle targets in SAR images has been successfully achieved, the performance of the network model has been improved, the ability to distinguish similar targets has been significantly improved, the recognition effect has been optimized, and more robust and accurate detection performance has been provided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355902A_ABST
    Figure CN120355902A_ABST
Patent Text Reader

Abstract

The invention discloses a small sample-oriented SAR image vehicle target detection and identification method. The method comprises the steps of obtaining a to-be-identified SAR image; the to-be-recognized SAR image is input into the trained vehicle target detection model, a detection result of the to-be-recognized SAR image is output, and the detection result is used for representing whether the to-be-recognized SAR image contains vehicles or not and representing the type and the position of each contained vehicle; wherein the trained vehicle target detection model is obtained through training by adopting a training set, and the SAR image sample in the training set is obtained through synthesizing real SAR images by adopting a region growing algorithm; the trained vehicle target detection model comprises a dilated convolutional layer for parallel sampling at multiple sampling rates, and a classification branch of a decoupling detection head of the trained vehicle target detection model comprises a full-dimensional dynamic convolutional layer. The method can improve the recognition precision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of deep learning, and in particular relates to a method for detecting and recognizing vehicle targets in SAR images with small samples. Background Art

[0002] SAR technology, with its unique imaging capabilities that are not restricted by lighting conditions and strong penetration, can provide detailed information about ground objects in any environment around the clock. It has become the core supporting technology in many key fields such as environmental reconnaissance, ocean monitoring, and geological exploration. Especially in marine scenes, thanks to the promotion of multiple public SAR ship target detection and recognition data sets, the accurate detection of ship targets has made significant progress. However, the core of human social activities is still focused on land. All kinds of information on land, especially the accurate detection and recognition of vehicle targets, have shown extremely high research value and urgent practical needs, whether in civil traffic management, urban planning, or in environmental reconnaissance, border monitoring and other fields. These needs require not only the ability to quickly and accurately identify vehicles, but also the ability to cope with complex and changing land environments.

[0003] Deep learning technology has shown great potential in the field of target detection and recognition due to its powerful image feature extraction capabilities. Compared with traditional methods, deep learning technology can deeply explore the deep-level, high-dimensional features of images and reveal more complex relationships in the data. It not only has significant advantages in feature extraction and can efficiently and accurately extract key information, but also the algorithm process is more concise and automatic, reducing manual intervention, showing excellent generalization performance and higher detection accuracy. In different scenarios, deep learning technology can maintain stable recognition effects.

[0004] However, although deep learning technology has great potential in the field of vehicle detection and recognition in SAR images, its development is restricted by the extremely scarce data sets. Traditional SAR target detection and recognition methods, such as the strategies based on constant false alarm rate and saliency analysis, mainly rely on the difference in scattering intensity between the target and the background. When facing SAR images with complex backgrounds, low target intensity, and insignificant features, the performance of these traditional methods often drops significantly. The rapid development of deep learning technology has provided new ideas and solutions for SAR target detection. The target detection algorithm based on deep learning can automatically learn and extract deep features in the image by constructing a deep neural network architecture, realizing the accurate detection and recognition of the target. However, a high-precision neural network model usually requires the support of a large amount of training data. Without sufficient data, it is difficult for the network to fully learn the features of the target, thus affecting its performance. In addition, in a complex terrestrial environment, vehicle targets are often interfered by imaging angles, surface coverings, and surrounding man-made facilities, resulting in weak echo signals and blurred image features. This not only increases the risk of missed detection and false alarms but also poses a huge challenge to accurate recognition. Especially when facing vehicles of the same type but different models, their high similarity in SAR images places strict requirements on the recognition algorithm.

[0005] Therefore, the main problems faced in the current field of vehicle target detection in SAR images include: scarce data sets, complex background distributions, and high difficulty in accurate recognition. Summary of the Invention

[0006] To solve the above problems existing in the prior art, the present invention provides a method for vehicle target detection and recognition in SAR images for small samples.

[0007] The technical problems to be solved by the present invention are realized through the following technical solutions:

[0008] The present invention provides a method for vehicle target detection and recognition in SAR images for small samples, including:

[0009] Obtain the SAR image to be recognized;

[0010] Input the SAR image to be recognized into the trained vehicle target detection model, and output the detection result of the SAR image to be recognized, where the detection result is used to indicate whether the SAR image to be recognized contains a vehicle, and the category and position of each vehicle included.

[0011] Among them, the trained vehicle target detection model is obtained by training with a training set, and the SAR image samples in the training set are synthesized from real SAR images using the region growing algorithm; the trained vehicle target detection model contains atrous convolution layers with multiple sampling rates in parallel sampling, and moreover, the classification branch of the decoupled detection head of the trained vehicle target detection model contains a full-dimensional dynamic convolution layer.

[0012] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0013] The present invention uses the region growing algorithm to generate SAR image samples, overcomes the inherent defects of limited applicability and insufficient stability of methods such as transfer learning, unsupervised learning, and semi-supervised learning, successfully realizes the effective expansion of small sample data, and solves the problem of the scarcity of SAR vehicle target detection and recognition data sets. The vehicle target detection model of the present invention contains atrous convolution layers with multiple sampling rates in parallel sampling, avoids the problem of information loss that may be caused by a complex pyramid structure, and also efficiently extracts multi-scale information, significantly improving the performance of the network model and making it particularly suitable for the detection and recognition tasks of vehicle targets. In addition, the present invention incorporates a full-dimensional dynamic convolution into the classification branch of the decoupled detection head in the vehicle target detection model, enhancing the attention and sensitivity of the network model to target features, significantly improving the ability to distinguish similar targets, and thus optimizing the recognition effect. The present invention successfully realizes the high-precision detection and recognition of vehicle targets in SAR images, demonstrates more robust and accurate detection performance, and provides strong technical support for the research and application in related fields.

[0014] The following will further elaborate on the present invention in detail in conjunction with the drawings and specific embodiments. Description of the Drawings

[0015] Figure 1 is a schematic flowchart of a method for detecting and recognizing vehicle targets in SAR images for small samples provided by an embodiment of the present invention;

[0016] Figure 2 is a comparison schematic diagram of sample generation results;

[0017] Figure 3 is a schematic flowchart of generating samples using the region growing algorithm provided by an embodiment of the present invention;

[0018] Figure 4 is a schematic diagram of the complete vehicle target region mask generated by an embodiment of the present invention;

[0019] Figure 5 is a schematic diagram of the synthetic image generated by an embodiment of the present invention;

[0020] Figure 6 It is a schematic diagram of a framework of the vehicle target detection model provided by an embodiment of the present invention;

[0021] Figure 7 It is a comparative schematic diagram of different feature pyramid structures;

[0022] Figure 8 It is a schematic diagram of a structure of the ASPP network;

[0023] Figure 9 It is a schematic diagram of a structure of the feature fusion pyramid network provided by an embodiment of the present invention;

[0024] Figure 10 It is a schematic diagram of a structure of the improved decoupled detection head provided by an embodiment of the present invention;

[0025] Figure 11 It is a schematic diagram of SAR images of ten types of vehicle targets and corresponding optical images;

[0026] Figure 12 They are two rural scene images;

[0027] Figure 13 They are partial synthetic images and corresponding background images;

[0028] Figure 14 It is a schematic diagram of detection results. Specific embodiments

[0029] The following further describes the present invention in detail with reference to specific embodiments, but the embodiments of the present invention are not limited thereto.

[0030] The present invention deeply analyzes the severe challenges faced in the current field of SAR vehicle target detection and recognition, including the problem that network training is prone to falling into the overfitting dilemma due to the scarcity of experimental samples, and the problems of poor detection and recognition effects caused by complex image backgrounds, similar target sizes, and highly similar target structures. To address these problems, the present invention proposes a SAR image vehicle target detection and recognition technology for small samples. First, using the region growing algorithm, on the basis of existing SAR images, a SAR vehicle target detection and recognition dataset that meets the experimental requirements is ingeniously synthesized, effectively alleviating the problem of sample scarcity; then, dilated convolutions with multiple sampling rates in parallel sampling are innovatively incorporated into the pyramid network to deeply extract multi-scale features, not only achieving the effective fusion of context information but also avoiding the generation of a large amount of redundant information, thereby improving the feature extraction ability of the network; finally, fully-dimensional dynamic convolution is introduced into the classification branch of the decoupled detection head, and by assigning different weights to the convolution kernels, more detailed target structure features are extracted, significantly enhancing the network's ability to distinguish similar targets and achieving high-precision detection and recognition of vehicle targets.

[0031] Figure 1 It is a schematic flowchart of a method for detecting and recognizing vehicle targets in SAR images for small samples provided by an embodiment of the present invention. As Figure 1 shown, the method includes:

[0032] S101. Obtain the SAR image to be recognized.

[0033] S102. Input the SAR image to be recognized into the trained vehicle target detection model, and output the detection result of the SAR image to be recognized. The detection result is used to indicate whether the SAR image to be recognized contains a vehicle, and the category and position of each vehicle included; wherein, the trained vehicle target detection model is obtained by training with a training set, and the SAR image samples in the training set are synthesized from the original SAR image by using the region growing algorithm; the trained vehicle target detection model includes dilated convolutional layers with multiple sampling rates in parallel sampling, and moreover, the classification branch of the decoupled detection head of the trained vehicle target detection model includes a full-dimensional dynamic convolutional layer.

[0034] The traditional sample generation method is to directly embed the vehicle target slices into the background image. This method cannot simulate the SAR image in the real situation, as Figure 2 (a) shown in. To make the difference between the generated image and the real image smaller, the present invention proposes a method for generating a synthetic image by the region growing algorithm. By finely extracting the vehicle target region, it is ensured that the clutter around the vehicle target is consistent with the clutter distribution in the background image, and a generated image closer to the real image as Figure 2 (b) shown in is obtained. As Figure 3 shown, the method for generating a synthetic image by the region growing algorithm proposed by the present invention is as follows:

[0035] S1. Obtain the original SAR images of at least one vehicle target to obtain at least one vehicle target slice image, and obtain multiple different SAR background images;

[0036] S2. For each SAR background image, embed all of the at least one vehicle target slice images into the SAR background image or embed some of the at least one vehicle target slice images into the SAR background image to obtain a spliced image.

[0037] S3. Use the region growing algorithm to generate a vehicle region mask for each vehicle target and a vehicle shadow region mask for each vehicle target according to the spliced image.

[0038] Specifically, filter a stitched image to obtain a filtered image; use the region growing algorithm to generate a vehicle region mask for each vehicle target based on a filtered image; use the region growing algorithm to generate a vehicle shadow region mask for each vehicle target based on a filtered image.

[0039] The vehicle target slice images used in the present invention are derived from real SAR images. Due to the imaging characteristics of SAR images, there are significant differences in the gray values of the vehicle body region, the vehicle shadow region, and the background clutter region. Therefore, through Gaussian filtering, the noise in the image is effectively smoothed, making the boundaries between regions clearer. This preprocessing step helps the region growing algorithm generate finer target region blocks. The Gaussian filtering function is as follows: In the formula, σ represents the variance, which determines the smoothness of the processed image. The larger σ is, the better the smoothing effect; the smaller σ is, the worse the smoothing effect. Using the Gaussian filtering function as a convolution kernel to perform convolution operation with the SAR image, a smoothed image can be obtained, and its formula is as follows: In the formula, f(x, y) represents the input SAR image, and I(x, y) represents the smoothed image after filtering.

[0040] The core principle of the region growing algorithm is to achieve image segmentation by merging pixel points with similar properties into a region. The entire algorithm mainly includes the following four steps:

[0041] 1) Seed point selection: Select one or more initial seed pixel points from the target image, and these points represent the starting points of the target regions to be segmented;

[0042] 2) Similarity judgment and region growth: Define a similarity criterion, then check the pixel points in the eight-neighborhood around the seed pixel points, and merge those neighborhood pixel points that are similar enough to the seed pixel points into the current segmentation region according to the similarity criterion, and use the newly merged pixel points as new seed points;

[0043] 3) Iterative update and termination: Repeat step 2), continuously merge new similar pixel points into the segmentation region, and update the set of seed pixel points until no new pixel points meet the merging conditions, that is, the iteration terminates;

[0044] 4) Generate mask and result screening: Generate a mask image according to the final segmentation region, and then cover the mask image on the original scene image for segmentation to obtain the final vehicle target region extraction result.

[0045] S4. Merge the vehicle region mask of each vehicle target with the vehicle shadow region mask to obtain a complete vehicle target region mask for each vehicle target.

[0046] S5. Extract the vehicle target regions that do not contain background information from the complete vehicle target region masks of each vehicle target, and invert the pixel values of the mask images after the vehicle target regions are extracted to obtain the background regions;

[0047] S6. Synthesize the vehicle target regions that do not contain background information with the background regions to obtain a SAR image sample.

[0048] In the SAR image, the vehicle body part shows a relatively high gray value due to the high scattering intensity of the metal material, and the vehicle shadow part shows a relatively low gray value due to the small amount of radar signals received. Therefore, it is necessary to extract the vehicle body and the shadow separately. In the present invention, pixel points with relatively high gray values (i.e., pixel points with gray values greater than or equal to the preset gray threshold) are first selected as the initial seed points, and based on this, the mask of the vehicle body region is gradually generated; then, pixel points with relatively low gray values (i.e., pixel points with gray values less than the preset gray threshold) are selected as the new seed points to generate the mask of the vehicle shadow region; finally, the masks are superimposed to obtain the complete vehicle target overall mask, realizing the accurate extraction of the vehicle target region. Exemplarily, the mask extraction results of the vehicle target are as Figure 4 shown Figure 4 In (b) of [], the pixel points with high gray values are used as the seed points to obtain the mask of the vehicle body region (i.e., the vehicle region masks of the three vehicle targets outlined by the three blue frames), Figure 4 In (c) of [], the pixel points with low gray values are used as the seed points to obtain the mask of the vehicle shadow region (i.e., the vehicle shadow region masks of the three vehicle targets outlined by the three blue frames). For Figure 4 the mask in (b) of [] and Figure 4 the mask in (c) of [] are stitched together to obtain the complete vehicle target region mask in (d) of [], Figure 4 i.e., the complete vehicle target region masks of the three vehicle targets outlined by the three blue frames. Comparing with the original image shown in (a) of [], it can be seen that the contour information of each vehicle is accurately segmented, and the clutter regions are discarded as much as possible. Figure 4 In (b) of [], in addition to the mask of the vehicle body region, i.e., the region within the blue square, there is also a part of the mask of the tree crown region. This is because the tree crown has a high scattering intensity and a large area, and will also generate a complete region in the region growing algorithm. However, as shown in the background region in (a) of [], the vehicle target region described in (b) of [], and the synthesized image shown in (c) of [], this part of the tree crown region comes from the background and acts on the background, so it does not affect the image generation result. Figure 4 In (b) of [], in addition to the mask of the vehicle body region, i.e., the region within the blue square, there is also a part of the mask of the tree crown region. This is because the tree crown has a high scattering intensity and a large area, and will also generate a complete region in the region growing algorithm. However, as shown in the background region in (a) of [], the vehicle target region described in (b) of [], and the synthesized image shown in (c) of [], this part of the tree crown region comes from the background and acts on the background, so it does not affect the image generation result. Figure 5 the background region shown in (a) of [], Figure 5 the vehicle target region described in (b) of [], and Figure 5 the synthesized image shown in (c) of [], it can be seen that this part of the tree crown region comes from the background and acts on the background, so it does not affect the image generation result.

[0049] In the present invention, the vehicle target detection model sequentially includes: a backbone feature extraction network, a feature fusion pyramid network, and a decoupled detection head module. For the above S102, specifically, after inputting the SAR image to be recognized into the trained vehicle target detection model, the backbone feature extraction network extracts N initial features of the SAR image to be recognized and inputs the N initial features into the feature fusion pyramid network; wherein, the feature fusion pyramid network includes atrous convolution layers with multiple parallel sampling rates; N is a positive integer; the feature fusion pyramid network performs feature fusion on the N initial features to obtain N fused features and inputs the N fused features into the decoupled detection head module; the decoupled detection head module performs classification and localization according to the N fused features to obtain the detection result. For example, N is 3. Exemplarily, Figure 6 is a schematic diagram of the framework of the vehicle target detection model, wherein, C3, C4, and C5 are three input layers of the feature fusion pyramid network, P3, P4, and P5 are three output layers of the feature fusion pyramid network, C3, C4, and C5 are used to receive each initial feature output by the backbone feature extraction network, and after the feature fusion pyramid network performs feature fusion on the N initial features, each fused feature is output through P3, P4, and P5 to the decoupled detection head module for target classification and position regression processing. Exemplarily, the backbone feature extraction network is ResNet50.

[0050] In the present invention, the feature fusion pyramid network is a network obtained by adding an ASPP network to the Feature Pyramid Network (FPN), wherein the ASPP network includes multiple parallel atrous convolution layers with different atrous rates. The Feature Pyramid Network (FPN) adjusts the size of the feature map through upsampling and downsampling operations on the feature map, achieving a perfect fusion of low-resolution and high-resolution feature maps. This design enables the model to extract and integrate multi-scale feature information and demonstrates stronger generalization ability when facing targets with different shapes and sizes. Figure 7The improvements from the FPN to the BIFPN pyramid shown above are all achieved by adding top-down or bottom-up feature propagation paths to strengthen feature communication between different levels. Such improvement measures can well enhance the detection and recognition effects of multi-scale targets. However, when detecting and recognizing land vehicle targets, the sizes of vehicle targets are similar and the differences are very small in SAR images. Therefore, if PAFPN or BIFPN is used as the neck network, the complex network structure will bring a large amount of redundant information, obscuring the effective feature information. At the same time, although a large number of downsamplings can increase the receptive field, they will also reduce the spatial resolution, resulting in information loss. Therefore, in the feature fusion pyramid network proposed in the present invention, the most original FPN network is selected as the feature fusion network. Dilated convolution is a special convolution that can increase the receptive field of the convolution kernel while keeping the computational amount unchanged. The number of holes filled between the convolution kernels can be controlled by the dilation rate, changing the area of the input feature map covered by the convolution kernel and expanding the receptive field of the convolution kernel. In the ASPP structure shown in Figure 8 , parallel sampling of the input feature map with dilated convolutions at different dilation rates (i.e., three dilated convolution layers with dilation rates of 3, 6, and 12 respectively and convolution kernel sizes of 3×3) can fuse deep semantic information and shallow detail information, improving the network's detection and recognition capabilities.

[0051] The main design idea of the feature fusion pyramid network proposed in the present invention is based on the FPN structure. By adding an ASPP module, parallel convolutions with multiple different sampling rates are used to extract context information, enriching the feature representation. At the same time, the problem of information loss introduced by downsampling is avoided. In addition, due to the sparse sampling method of dilated convolution, the obtained information is all from long distances, lacking the correlation of local information. Therefore, a residual structure using ordinary convolution to capture local features is added to each layer of features. The feature fusion pyramid network proposed in the present invention includes: N input layers, N feature channel adjustment layers, N intermediate layers, N two-dimensional convolution layers, N upsampling layers, N ASPP networks, N addition processing units, and N output layers; where the output of the nth input layer is connected to the input of the nth feature channel adjustment layer, the output of the nth feature channel adjustment layer is simultaneously connected to the input of the nth intermediate layer and the input of the nth two-dimensional convolution layer, the output of the nth intermediate layer is simultaneously connected to the input of the nth ASPP network and the input of the nth upsampling layer, the output of the nth upsampling layer is connected to the input of the (n + 1)th intermediate layer, the outputs of the nth ASPP network and the nth two-dimensional convolution layer are both connected to the input of the nth addition processing unit, and the output of the nth addition processing unit is connected to the input of the nth output layer; n is a positive integer, and the value of n ranges from 1 to N. Exemplarily, Figure 9 is a schematic structural diagram of the feature fusion pyramid network, as shown in Figure 9As shown, each feature channel adjustment layer is a 1*1 convolutional layer. C3, C4, and C5 are the input layers, and the outputs of C3, C4, and C5 are each one of the three outputs from the backbone feature extraction network. N3, N4, and N5 are the intermediate feature layers. As Figure 9 shown, first, the number of feature channels of the output features of the input layer is adjusted to 256 through a 1*1 convolutional layer, then the high-level feature resolution of the intermediate layer is upsampled to be consistent with the adjacent low-level features through the upsampling layer, and finally, these two features are added to obtain the features output by the intermediate feature layer. The calculation formula is: N i = upsample(N i+1 ) + Conv 1×1 (c i ), i = 3, 4, N i = Conv 1×1 (c i ), i = 5. P3, P4, and P5 are the output layers. Part of the features output by each output layer come from the results of pooling the intermediate layer through the ASPP network. Among them, the dilation rates of the three dilated convolutional layers in the ASPP network are 3, 6, and 12 respectively, and the kernel sizes are all 3×3. The other part comes from the features output by the input layer after passing through a 1*1 convolutional layer and then a 3*3 two-dimensional convolution. The calculation formulas for the features output by each output layer of P3, P4, and P5 are: P i = Conv 3×3,r=3,6,12 (N i ) + Conv 3×3 (Conv 1×1 (c i ), i = 3, 4, 5.

[0052] The decoupled head assigns the classification and localization tasks to different network branches, enabling each branch to focus on learning the weights required for its respective task, and then generates prediction results for classification, bounding box regression, and confidence estimation through convolutional layers, fully connected layers, and activation functions. In SAR images, vehicles of the same category but different models will exhibit extremely similar structural characteristics, which makes it difficult for the convolutional layer of the classification branch to capture effective features that can distinguish these subtle differences. This high degree of similarity brings great interference to the accurate classification of vehicle targets, thereby affecting the recognition effect of the entire network. Therefore, how to enhance the network's ability to distinguish different model vehicle targets by optimizing the structure of the classification branch is the key work to improve the network's recognition effect. In the present invention, the decoupled detection head module contains N decoupled detection heads, and each output of the feature fusion pyramid network is connected to a decoupled detection head. For example, when the structure of the feature fusion pyramid network is as Figure 9When, a decoupling detection head is respectively connected behind P3, P4, and P5. Each decoupling detection head includes a classification branch and a localization branch, and the convolutional layers in the classification branch of each decoupling detection head are all full-dimensional dynamic convolutional layers. Exemplarily, Figure 10 is a schematic structural diagram of an improved decoupling detection head proposed by the present invention. Figure 10 The yellow rectangular block in it represents the input layer. Two branches are connected behind the input layer. Among them, the four sequentially connected blue rectangular blocks in the classification branch represent four full-dimensional dynamic convolutional layers (ODConv) with a stride of 1, a padding of 1, and a convolutional kernel of 3×3. A green rectangular block represents a fully connected layer, and the output feature size is h×w×c; the four sequentially connected red rectangular blocks in the localization branch represent four two-dimensional convolutional layers with a stride of 1, a padding of 1, and a convolutional kernel of 3×3. Two green rectangular blocks represent two fully connected layers, and the output feature sizes are h×w×4 and h×w×1 respectively. The ordinary convolutional layers in the classification branch of the existing detection heads cannot adaptively adjust the weights according to the target feature distribution to capture more contributive effective features, which results in that in complex tasks containing multiple similar targets, ordinary convolution cannot achieve good recognition effects. Compared with ordinary convolution, dynamic convolution can dynamically adjust the weights of each convolutional kernel by calculating the attention weights of the input features in the spatial dimension, input channel dimension, output channel dimension, and convolutional kernel dimension, and capture the key features of the target more accurately, so as to more flexibly adapt to different input samples. Therefore, after replacing the ordinary convolution in the classification branch with dynamic convolution, the dynamic convolution can extract effective features of vehicle targets from SAR images, covering rich target structure features and context information, which helps to enhance the distinguishability of various types of vehicles and alleviate the problem that the structures of different models of vehicles are highly similar, thereby improving the detection and recognition performance of the network for vehicles.

[0053] The present invention has the following advantages:

[0054] 1. Aiming at the lack of SAR vehicle target detection and recognition datasets, a sample augmentation method based on the region growing algorithm is designed, which can synthesize a sufficient amount of datasets that meet the experimental requirements by using a small number of existing pictures;

[0055] 2. Aiming at the problem that in the vehicle target detection task, the target sizes are similar and the complex pyramid structure will cause feature loss, a top-down unidirectional feature pyramid is used as the feature fusion network, and dilated convolutions with multiple sampling rates for parallel sampling are incorporated to mine multi-scale features, avoiding the generation of a large amount of redundant information while realizing the fusion of context information;

[0056] 3. To address the problem of high difficulty in identifying similar vehicle targets, a full-dimensional dynamic convolution is added to the classification branch of the decoupled detection head. By assigning different weights to the convolution kernels, more detailed target structure features are extracted, enhancing the network's ability to distinguish similar targets and improving the target recognition accuracy.

[0057] The effectiveness of the method proposed in the present invention is illustrated through experiments below.

[0058] (1) Experimental data and parameters

[0059] The data used in the training and testing processes of the network proposed in the present invention is sourced from the publicly available MSTAR dataset. This dataset contains sliced images of ten types of vehicle targets and large-scene clutter images. The size of the sliced images of the 10 types of vehicle targets is generally 128×128 pixels, including images of various vehicle targets at different azimuth angles. Each image contains only one stationary vehicle target. Figure 11 (a) - (j) in it show the SAR images and corresponding optical images of the ten types of vehicle targets. There are 100 large-scene clutter images in the MSTAR dataset, including natural scenes such as lakes, farmlands, grasslands, forests, and shrubs, and artificial buildings such as houses, roads, and bridges. Figure 12 (a) and (b) in it show two rural scene images. (a) is a forest scene, and (b) is a farmland scene.

[0060] The sample generation method based on the region growing algorithm proposed in the present invention is used to synthesize the sliced images of vehicle targets and large-scene clutter images in MSTAR to obtain a dataset that meets the task application scenario. To reduce the network burden, the background image is cropped into an image with a size of 512×512 pixels, as shown in Figure 13 (a) in it. Compared with the optical image, the target in the SAR image is just a pile of bright spots and cannot be recognized by the naked eye. In addition, vehicle targets of different models under the same category also have similar scattering characteristics, further increasing the difficulty of detection and recognition. Therefore, images of four types of vehicle targets, namely 2S1, BMP2, BTR60, and BTR70, are selected as experimental images and embedded in the background image to synthesize 700 training and testing images, which altogether contain 1800 vehicle targets, fully simulating the real scene to test the detection and recognition effects of different categories and different models of targets under the same category, ensuring the authenticity and effectiveness of the experiment. Some of the synthesized images are shown in Figure 13 (b) in it.

[0061] In the experimental process of the present invention, the number of model training rounds is set to 300, batch_size is set to 4, the SGD optimizer is used to update the parameters, the initial learning rate of the optimizer is set to 0.01, the momentum factor is set to 0.9, and the weight decay factor is set to 0.0001. The experiment is run on a computer equipped with an NVIDIA T40c GPU.

[0062] (2) Measurement Metrics

[0063] The average precision (AP) is used as the evaluation criterion to quantitatively evaluate the detection performance of the network proposed in this invention. The calculation formula is as follows:

[0064]

[0065] Among them, P represents precision and R represents recall. AP50 is the AP score when the IoU threshold is selected as 0.5. The precision and recall are defined as follows:

[0066]

[0067]

[0068] Among them, TP is the number of vehicles actually detected, FP is the number of backgrounds detected as vehicles, and FN represents the number of real vehicles detected as backgrounds. The actually detected vehicles are defined as the targets whose IoU value between their bounding boxes and the ground truth annotations is greater than 0.5.

[0069] (3) Comparison of Detection Performance

[0070] To verify the effectiveness and advancement of the vehicle target detection method proposed in this invention, some classic object detection network frameworks are selected for method verification as comparative experiments.

[0071] Table 1 Performance Comparison of the Method Proposed in this Invention with Other Methods

[0072]

[0073] It can be seen from Table 1 that the average AP50 value of the method proposed in this invention is 97.82%, which is 3.32 percentage points higher than that of the FCOS network as the baseline model. The detection and recognition effect has been significantly improved, fully verifying the effectiveness of the algorithm proposed in this invention. At the same time, the average AP50 of the proposed algorithm is 5.52 percentage points higher than that of Faster R-CNN, 5.12 percentage points higher than that of RetinaNet, 4.37 percentage points higher than that of SSD, 3.92 percentage points higher than that of YOLOv3, and 1.07 percentage points higher than that of YOLOv5. This further shows that the method proposed in this invention has obvious advantages in detection and recognition performance.

[0074] (4) Ablation Experiment

[0075] An ablation experiment was conducted to analyze the impact of each module in the algorithm proposed in the present invention on the detection performance, with the FCOS network as the baseline model. Accordingly, atrous spatial pyramid and improved decoupled head were introduced into the FCOS network to verify the module functions. Table 2 shows the comparison results of the network performance after adding each module. "√" indicates the adoption of this module, and "×" indicates the non-adoption of this module.

[0076] Table 2 Comparison of Module Performance

[0077]

[0078] It can be seen from Table 2 that compared with the original FCOS network, after adding the atrous spatial pyramid network, the average AP50 increased by 2.02 percentage points. After adding the improved decoupled head, the average AP50 increased by another 1.30 percentage points. In summary, after the combination of the atrous spatial pyramid and the improved decoupled head, the average AP50 of the network can be increased from 94.50% of the FCOS network to 97.82%, an increase of 3.32 percentage points, which fully demonstrates the effectiveness of the method proposed in the present invention.

[0079] In addition, BTR60 and BTR70 are targets of the same type, with similar appearance structures and high differentiation difficulty. It can be seen from Table 2 that after adding the improved decoupled head, the AP50 of both BTR60 and BTR70 has been significantly improved, and the difference between the two has become smaller, verifying that the improved decoupled head proposed in the present invention can improve the recognition accuracy of the network.

[0080] (5) Result Analysis

[0081] To further intuitively demonstrate the detection and recognition advantages of the algorithm proposed in the present invention compared with other algorithms, Figure 14 the detection results of the algorithm proposed in the present invention, the SSD network, the FCOS network, and the YOLOv5 network in different scenarios are given. In addition, the ground truth labels are provided as a reference. In the figure, the white detection box represents BMP2, the green detection box represents BTR70, the red detection box represents 2S1, and the blue detection box represents BTR60.

[0082] From Figure 14It can be seen that in a variety of land scenarios, due to the influence of conditions such as forests, roads, and shadows, the performance of networks such as SSD, FCOS, and YOLOv5 networks has been interfered, and problems such as false alarms, missed alarms, and misdetections have occurred to varying degrees. Among them, the SSD network is biased towards the BTR70 type during recognition and misdetects the target as BTR70 many times. The FCOS network has poor detection effects and often has false alarms and missed alarms. YOLOv5 is biased towards the BTR60 type during recognition and repeatedly misdetects the target as BTR60. These problems have all led to a decline in network performance. However, the algorithm proposed in the present invention correctly detects and recognizes all targets without problems such as false alarms, missed alarms, and misdetections, and has good detection and recognition effects.

[0083] In summary, the present invention proposes a sample augmentation method based on the region growing algorithm, which overcomes the inherent defects of limited applicability and insufficient stability of methods such as transfer learning, unsupervised learning, and semi-supervised learning, and successfully realizes the effective augmentation of small sample data. The atrous spatial pyramid used in the neck network avoids the information loss problem that may be caused by a complex pyramid structure, and also efficiently extracts multi-scale information, significantly improving the performance of the network and making it particularly suitable for the detection and recognition tasks of vehicle targets. In addition, dynamic convolution is incorporated into the classification branch of the decoupled detection head, enhancing the network's attention and sensitivity to target features, significantly improving the ability to distinguish similar targets, and thus optimizing the recognition effect. The present invention has successfully realized the high-precision detection and recognition of vehicle targets in SAR images, demonstrating more robust and accurate detection performance, and providing strong technical support for the research and application in related fields.

[0084] It should be noted that the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more features. In the description of the present invention, "a plurality" means two or more unless otherwise specifically defined.

[0085] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.

[0086] In the description, the term "comprising" does not exclude other components or steps, and the article "a" or "an" does not exclude a plurality. Certain measures are recited in mutually different embodiments, but this does not mean that these measures cannot be combined to produce good results.

[0087] The above content is a further detailed description of the present invention in conjunction with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can be made, and all should be regarded as falling within the protection scope of the present invention.

Claims

1. A vehicle target detection and recognition method for SAR images facing small samples, characterized in that Including: Obtain the SAR image to be recognized; Input the SAR image to be recognized into the trained vehicle target detection model, and output the detection result of the SAR image to be recognized, where the detection result is used to indicate whether the SAR image to be recognized contains a vehicle, and the category and location of each vehicle included; Among them, the trained vehicle target detection model is obtained by training with a training set, and the SAR image samples in the training set are synthesized from real SAR images by using the region growing algorithm; the trained vehicle target detection model contains atrous convolution layers with multiple sampling rates in parallel, and moreover, the classification branch of the decoupled detection head of the trained vehicle target detection model contains a full-dimensional dynamic convolution layer.

2. The method according to claim 1, wherein The step of inputting the SAR image to be recognized into the trained vehicle target detection model and outputting the detection result of the SAR image to be recognized includes: Input the SAR image to be recognized into the trained vehicle target detection model, and the backbone feature extraction network in the vehicle target detection model extracts N initial features of the SAR image to be recognized, and inputs the N initial features into the feature fusion pyramid network in the vehicle target detection model; among them, the feature fusion pyramid network contains atrous convolution layers with multiple sampling rates in parallel; N is a positive integer; The feature fusion pyramid network performs feature fusion on the N initial features to obtain N fused features, and inputs the N fused features into the decoupled detection head module in the vehicle target detection model; The decoupled detection head module classifies and locates according to the N fused features to obtain the detection result.

3. The method according to claim 1, wherein The vehicle target detection model contains a feature fusion pyramid network for feature fusion, and the feature fusion pyramid network is a network obtained by adding an ASPP network to an FPN network, where the ASPP network contains multiple parallel atrous convolution layers with different dilation rates.

4. The method according to claim 3, characterized in that The feature fusion pyramid network includes: N input layers, N feature channel adjustment layers, N intermediate layers, N two-dimensional convolution layers, N upsampling layers, N ASPP networks, N addition processing units, and N output layers; Among them, the output of the nth input layer is connected to the input of the nth feature channel adjustment layer, the output of the nth feature channel adjustment layer is simultaneously connected to the input of the nth intermediate layer and the input of the nth two-dimensional convolution layer, the output of the nth intermediate layer is simultaneously connected to the input of the nth ASPP network and the input of the nth upsampling layer, the output of the nth upsampling layer is connected to the input of the (n + 1)th intermediate layer, the output of the nth ASPP network and the output of the nth two-dimensional convolution layer are both connected to the input of the nth addition processing unit, and the output of the nth addition processing unit is connected to the input of the nth output layer; n is a positive integer, and the value of n ranges from 1 to N.

5. The method according to claim 1 or 2, characterized in that, The vehicle target detection model includes a decoupled detection head module, and the decoupled detection head module includes N decoupled detection heads. The convolutional layers in the classification branches of each decoupled detection head are all full-dimensional dynamic convolutional layers.

6. The method according to claim 2, wherein N is 3.

7. The method according to claim 1, characterized in that, The method for synthesizing SAR image samples from real SAR images using the region growing algorithm includes: Obtaining real SAR images of at least one vehicle target to obtain at least one vehicle target slice image, and obtaining multiple different SAR background images; For each SAR background image, embedding all of the at least one vehicle target slice image into the SAR background image or embedding some of the at least one vehicle target slice images into the SAR background image to obtain a stitched image; Using the region growing algorithm to generate a vehicle region mask for each vehicle target and a vehicle shadow region mask for each vehicle target respectively according to the stitched image; Fusing the vehicle region mask of each vehicle target with the vehicle shadow region mask of each vehicle target to obtain a complete vehicle target region mask for each vehicle target; Extracting the vehicle target region that does not contain background information from the complete vehicle target region mask of each vehicle target, and inverting the pixel values of the mask image after extracting the vehicle target region to obtain the background region; Synthesizing the vehicle target region that does not contain background information with the background region to obtain a SAR image sample.

8. The method according to claim 7, wherein The step of using the region growing algorithm to generate a vehicle region mask for each vehicle target and a vehicle shadow region mask for each vehicle target respectively according to the stitched image includes: Filtering the stitched image to obtain a filtered image; Using the region growing algorithm to generate a vehicle region mask for each vehicle target according to the filtered image; Using the region growing algorithm to generate a vehicle shadow region mask for each vehicle target according to the filtered image.