Few-shot SAR Target Detection and Recognition Method Based on Meta-Learning and Metric Learning

By building a small sample SAR object detection and recognition network based on meta-learning and metric learning, the problems of insufficient detection performance of new object detection and low recognition accuracy in fine-grained tasks are solved, and efficient detection and accurate recognition of new object targets are achieved.

CN116664823BActive Publication Date: 2025-07-29XIDIAN UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310639895.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-31
Publication Date
2025-07-29
Estimated Expiration
2043-05-31

AI Technical Summary

Technical Problem

The prior art has insufficient performance in SAR object detection and recognition under small sample conditions, and has low recognition accuracy and poor robustness in fine-grained tasks.

Method used

A small sample SAR object detection and recognition network based on meta-learning and metric learning is built, and the quality of new candidate areas and the accuracy of fine-grained feature recognition is improved by designing the area recommendation module of class attention modulation and the fine-grained detection and recognition module based on multi-relational metrics.

Benefits of technology

It significantly improves the detection and recognition performance of new class targets and the recognition accuracy and robustness in fine-grained tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116664823B_ABST
    Figure CN116664823B_ABST
Patent Text Reader

Abstract

The present invention discloses a few-shot SAR target detection and recognition network based on meta-learning and metric learning, which mainly solves the problems of insufficient detection performance for new-class targets and low recognition accuracy and poor robustness in fine-grained tasks in the prior art. The implementation solution is as follows: annotate and partition the SAR measured data to generate the support sets and query sets of the base classes and new classes; construct a few-shot SAR target detection and recognition network composed of a data preprocessing module, a feature extraction module, a region proposal module, and a fine-grained detection and recognition module; based on the stochastic gradient descent algorithm, use the support sets and query sets of the base classes and new classes to train the network; input the SAR image to be detected and recognized into the trained few-shot SAR target detection and recognition network to obtain the detection and recognition results of the SAR target. The present invention significantly improves the detection and recognition performance of new-class targets and the recognition accuracy and robustness in fine-grained tasks, and can be used for environmental reconnaissance and situation awareness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of radar remote sensing, and further relates to a small-sample SAR target detection method, which can be used for environmental reconnaissance and situation awareness. Background Art

[0002] Synthetic Aperture Radar (SAR) is a microwave imaging radar, which can achieve all-weather, all-day, long-distance, and high-resolution imaging of ground, sea surface and other targets without being restricted by illumination and climate conditions. Therefore, it plays an important role in both military and civilian fields. Among them, the detection and recognition of key targets in SAR images are the difficulties and key issues in the field of Automatic Target Recognition (ATR) of radar.

[0003] With the application of deep learning-based target detection and recognition methods to SAR image interpretation, great progress has been made in SAR target detection and recognition. At present, most of the mainstream SAR target detection and recognition methods are based on general target detection algorithms, such as single-stage algorithms SSD, YOLO, and two-stage algorithms Faster RCNN, Casacade RCNN. These algorithms based on large-scale deep networks have good detection and recognition performance when there are sufficient training samples. However, due to the non-cooperativeness of targets and the limitation of observation conditions, it is difficult to collect and label specific category samples. At this time, the deep network model driven by big data will overfit due to severely insufficient training samples, resulting in a significant decline or even complete failure of the detection and recognition performance of the model. At present, small-sample learning methods including data augmentation, semi-supervised learning, and transfer learning have been proposed in the field of computer vision.

[0004] Small-sample target detection and recognition aims to design effective network structures and training strategies, and optimize the network model through a large number of base class training sets and a small number of new class training sets, so that it can better detect and recognize new class targets.

[0005] The patent document with the application number CN202211273499.4 discloses "a small-sample SAR image target detection method and system based on meta-learning". It is completed in three parts, that is, first, a backbone network with shared weights is used to extract support and query features; then, a Region Proposal Network (RPN) is used to generate candidate regions of the query image and extract the corresponding RoI features; then, cross-correlation operations are used to aggregate the query RoI features and support features, and input them into the backend of the target detector to complete class inference and bounding box regression. Although this method effectively improves the detection and recognition performance of the deep network model under small-sample conditions, it has the following problems: 1) Since it only relies on a general Region Proposal Network to generate candidate regions of new class targets, the quality is low, resulting in insufficient detection performance for new class targets and affecting subsequent recognition tasks. 2) Since it only relies on a single global feature similarity for class inference, the recognition accuracy is low and the robustness is poor in fine-grained tasks. Summary of the Invention

[0006] The object of the present invention is to propose a few-shot SAR target detection and recognition method based on meta-learning and metric learning for the deficiencies of the above-mentioned existing technologies, so as to improve the detection and recognition performance of new-class targets and enhance the accuracy and robustness of recognition in fine-grained tasks.

[0007] The technical idea of the present invention is to improve the detection performance of new-class targets and enhance the accuracy and robustness of recognition in fine-grained tasks by designing a few-shot SAR target detection and recognition network based on meta-learning and metric learning. The implementation steps are as follows:

[0008] (1) Construct a support set and a query set:

[0009] (1a) Collect a large number of SAR images containing base-class targets and a small number of SAR images containing new-class targets, and label the target positions and categories in each SAR image.

[0010] (1b) Use the collected SAR images of base-class targets as the query set of the base class, and randomly sample some SAR images from the collected SAR images of base-class targets to form the support set of the base class; use all the collected SAR images of new-class targets as the query set and support set of the new class respectively.

[0011] (2) Construct a few-shot SAR target detection and recognition network based on meta-learning and metric learning:

[0012] (2a) Establish a data preprocessing module consisting of a query image preprocessing process that sequentially performs multi-scale transformation, random flipping, and data normalization and a support image preprocessing process that sequentially performs target cropping, random flipping, and data normalization.

[0013] (2b) Establish a feature extraction module A consisting of a twin network with shared weights.

[0014] (2c) Establish a region proposal module G consisting of a cascade of a feature aggregation sub-module, an anchor box generation sub-module, a positive and negative sample assignment sub-module, a classification and regression sub-module, and a post-processing sub-module, and use the cross-entropy loss function and the SmoothL1 loss function as the classification loss and the bounding box regression loss

[0015] (2d) Establish a fine-grained detection and recognition module D consisting of a region of interest extraction sub-module, a local relationship sub-module, a global relationship sub-module, and a cross-relationship sub-module. Its output is the bounding box coordinates (x, y, w, h) of the inspected target and the similarity s of each category of the target, and these two parameters are respectively substituted into the cross-entropy loss function and the Smooth L1 loss function to calculate the classification loss value and the bounding box regression loss value

[0016] (2e) Cascade the modules established in (2a), (2b), (2c), and (2d) in sequence to form a small - sample SAR target detection and recognition network, and define the loss function of this network as:

[0017] (3) Conduct base - class training on the small - sample SAR target detection network:

[0018] (3a) Input the support set and query set of the base class into the small - sample SAR target detection and recognition network, and calculate the loss value of each iteration where i is the iteration number during base - class training, and update the network parameters through the stochastic gradient descent algorithm according to this loss value;

[0019] (3b) Repeat the process of (3a) until the network converges to obtain a preliminarily trained small - sample SAR target detection and recognition network;

[0020] (4) Conduct small - sample fine - tuning on the small - sample SAR target detection network:

[0021] (4a) Input partial support images of the base - class support set, query - set images, and all new - class support sets and query - set images into the preliminarily trained small - sample SAR target detection network, and calculate the loss value of each iteration where j represents the iteration number in the small - sample fine - tuning stage, and update the network parameters through the stochastic gradient descent algorithm using this loss value;

[0022] (4b) Repeat the process of (4a) until the network converges to obtain the finally trained small - sample SAR target detection and recognition network.

[0023] (5) Input the SAR image to be detected and recognized into the finally trained small - sample SAR target detection and recognition network obtained in step (4) to obtain the detection and recognition results.

[0024] Compared with the prior art, the present invention has the following advantages:

[0025] First, in constructing a small - sample SAR target detection and recognition network based on meta - learning and metric learning, since the present invention designs a region proposal module based on class - attention modulation, it can aggregate support features and query features through depth - wise separable convolution, improve the quality of the new - class candidate regions generated by this module, and thus improve the detection and recognition performance of new classes.

[0026] Second, in the construction of the few-shot SAR target detection and recognition network based on meta-learning and metric learning, since the fine-grained detection and recognition module based on multi-relationship metric is designed, the accuracy and robustness of the recognition of fine-grained features can be effectively improved by measuring the global similarity, local similarity, and cross-similarity of the support features and query features. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 It is a flowchart for implementing the present invention;

[0028] Figure 2 It is a model diagram of the few-shot SAR target detection and recognition network constructed in the present invention;

[0029] Figure 3 It is a simulation result diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0030] The following further elaborates on the examples and effects of the present invention with reference to the accompanying drawings.

[0031] Refer to Figure 1 , in this example, the few-shot SAR target detection and recognition method based on meta-learning and metric learning sequentially includes data collection and production, construction of a few-shot SAR target detection and recognition network, training of the few-shot SAR target detection and recognition network, and obtaining SAR target detection and recognition results. The specific implementation steps are as follows:

[0032] Step 1, data collection and production.

[0033] 1.1) Collect a large number of SAR images containing base-class targets and a small number of SAR images containing new-class targets, and label the target positions and categories in each SAR image;

[0034] 1.2) Compose the query set of base-class targets from the collected base-class target images, and randomly sample some base-class target images to form the support set of base-class targets;

[0035] 1.3) Compose the query set and support set of new-class targets from all the collected new-class target images.

[0036] In this example, the SAR images are from the spaceborne radar on the Gaofen-3 satellite; the sizes of the SAR images include 600×600, 1024×1024, and 2048×2048, and there are a total of seven types of aircraft targets, namely: A220, A320 - 321, A330, ARJ21, Boeing737 - 800, Boeing787, and other. ARJ21 and Boeing787 are regarded as new-class targets, and the rest are regarded as base-class targets.

[0037] Step 2, construct a few-shot SAR target detection and recognition network.

[0038] Reference Figure 2 , the implementation of this step is as follows:

[0039] 2.1) Establish a data preprocessing module consisting of a query image preprocessing process that sequentially performs multi-scale transformation, random flipping, and data normalization, and a support image preprocessing process that sequentially performs target cropping, random flipping, and data normalization. In the implementation example of the present invention, the scales of the multi-scale transformation include 440×440, 472×472, 504×504, 536×536, 568×568, and 600×600; the flipping probability of the random flipping is 0.5; the target cropping crops the target of the support image and scales it to a size of 320×320;

[0040] 2.2) Establish a feature extraction module A composed of a twin network with shared weights. This module is composed of 4 cascaded convolutional modules a1, a2, a3, and a4, where:

[0041] The second convolutional module a2 is composed of 3 residual blocks;

[0042] The third convolutional module a3 is composed of 4 cascaded residual blocks;

[0043] The fourth convolutional module a4 is composed of 6 cascaded residual blocks;

[0044] The query feature map output by the entire feature extraction module is denoted as where is the query feature map output by the i-th convolutional module a i , and X Q is the input query SAR image;

[0045] The support feature map output by the entire feature extraction module is denoted as where is the support feature map output by the i-th convolutional module a i , and X s is the input support SAR image.

[0046] In this example, each residual block in the four cascaded convolutional modules includes 3 convolutional modules c1, c2, and c3, where:

[0047] The first convolutional module c1 is composed of a 1×1 convolutional layer, a batch normalization layer, and a ReLU activation layer in cascade;

[0048] The second convolutional module c2 is composed of a 3×3 convolutional layer, a batch normalization layer, and a ReLU activation layer in cascade;

[0049] The third convolutional module c3 is composed of a 1×1 convolutional layer and a batch normalization layer in cascade;

[0050] The output feature map of each convolutional module in the residual block is denoted as where x is the input feature map of the residual block;

[0051] The output feature map of the residual block is denoted as z = relu(y s + x), where relu(·) represents the ReLU activation function.

[0052] 2.3) Establish a region proposal module G composed of a cascaded feature aggregation sub-module, anchor box generation sub-module, positive and negative sample assignment sub-module, classification and regression sub-module, and post-processing sub-module, where:

[0053] The feature aggregation sub-module uses depthwise separable convolution to complete the aggregation of the support feature map and the query feature map, and its output is:

[0054]

[0055] where, is the query feature obtained by the feature extraction module A, W and H respectively represent the width and height of the query feature map, is the support feature obtained by the feature extraction module A, K represents the side length of the support feature map, and C represents the depth of the query feature map and the support feature map;

[0056] The anchor box generation sub-module is used to generate rectangular anchor boxes with scales of 2 2 , 4 2 , 8 2 , 16 2 and 32 2 , and aspect ratios of 1:2, 1:1, and 2:1 at each point of the aggregated feature map P;

[0057] The positive and negative sample assignment sub-module is used to assign positive and negative samples to the anchor boxes, that is, the anchor boxes with an intersection over union greater than 0.7 with the labeled bounding box of the same category as the support feature are assigned as positive samples, the anchor boxes with an intersection over union less than 0.3 with the labeled bounding box of the same category as the support feature are assigned as negative samples, and the remaining anchor boxes are assigned as irrelevant samples;

[0058] The classification and regression sub-module is used to obtain candidate regions that may contain the target, and it is composed of a 3×3 convolutional layer g1 and two parallel 1×1 convolutional layers g cls and g reg cascaded, where g reg is used to obtain the bounding box offset of the candidate region relative to the anchor box, and g cls is used to predict the confidence that the candidate region contains the target;

[0059] The post - processing sub - module is used to filter redundant candidate regions and output the top N candidate regions with the highest confidence levels \(\{(p j , b j )|j = 1, 2, ..., N\}\), where \(p j \) is the confidence level of the \(j\) - th candidate region containing the target, and \(b j =(x j , y j , w j , h j )\) is the bounding box of the \(j\) - th candidate region. \((x j , y j )\) represents the center coordinates of the bounding box of the \(j\) - th candidate region, and \(w j \) and \(h j \) represent the width and height of the bounding box of the \(j\) - th candidate region respectively.

[0060] 2.4) Establish a fine - grained detection and recognition module D including a region of interest extraction sub - module, a local relationship sub - module, a global relationship sub - module, and a cross - relationship sub - module, where:

[0061] The region of interest extraction sub - module consists of an adaptive average pooling layer, which is used to average - pool the query feature map \(Y Q \) extracted by the feature extraction module A into a \(7\times7\) query RoI feature \) according to the candidate regions provided by the region proposal module G, and average - pool the support feature map \(Y S \) extracted by the feature extraction module A into a \(7\times7\) support RoI feature \), where C is the depth of the feature map;

[0062] The local relationship sub - module is used to measure the local feature similarity \(s l \) between the query RoI feature and the support RoI feature. This module consists of a cascaded 1×1 convolutional layer, a depth - separable convolutional layer, and a fully - connected layer with an output feature number of 1;

[0063] The global relationship sub - module is used to measure the global similarity \(s g \) between the query RoI feature and the support RoI feature. This module consists of a cascaded global average pooling layer, a feature concatenation layer, and a fully - connected layer with an output feature number of 1;

[0064] The cross - relationship sub - module is used to measure the cross - relationship similarity between the query RoI feature and the support RoI feature and obtain the predicted bounding box coordinates. This module consists of a cascaded feature concatenation layer, a 1×1 standard convolutional layer, a 3×3 standard convolutional layer, a 1×1 standard convolutional layer, a global average pooling layer, and a parallel fully - connected layer \(d cls \) with an output feature number of 1 and a fully - connected layer \(d reg \) with an output feature number of 4, where \(d clsUsed to predict the cross - relationship similarity s between the query RoI feature and the support RoI feature p , d reg Used to predict the bounding box coordinates (x, y, w, h), where (x, y) represents the center point coordinates of the bounding box, and w and h are the width and height of the bounding box respectively;

[0065] Parallelize the above - mentioned local relationship sub - module, global relationship sub - module and cross - relationship sub - module, and cascade the above - mentioned region of interest extraction sub - module with this parallel module to form the fine - grained detection and recognition module D.

[0066] In this example, the fine - grained recognition module D outputs the similarity s of each category of the detected target, and the formula is as follows:

[0067]

[0068] Among them, and are the local similarity, global similarity and cross - similarity of the j - th category of the target respectively, and are expressed as follows:

[0069]

[0070]

[0071]

[0072] Among them, represents the query RoI feature, represents the support Roi feature of the j - th category, C is the depth of the feature map, σ(·) represents the Sigmoid activation function, FC(·) represents a fully - connected layer with 1 output feature, DWConv(·, ·) represents a depth - separable convolution, Conv(·) represents a 1×1 standard convolution, Cat(·, ·) represents concatenating the feature maps by channel, GAP(·) represents global average pooling, and Conv3(·) represents a cascaded 1×1 standard convolution layer, 3×3 standard convolution layer and 1×1 standard convolution layer.

[0073] Step 3: Train the small - sample SAR target detection and recognition network.

[0074] 3.1) Train the base class of the small - sample SAR target detection and recognition network

[0075] 3.1.1) Input the support set and query set of the base class into the small - sample SAR target detection and recognition network, calculate the loss value of each iteration, and update the network parameters through the stochastic gradient descent algorithm according to this loss value;

[0076] 3.1.2) Calculate the loss of the small - sample SAR target detection and recognition network

[0077]

[0078] Among them: is the classification loss of the region proposal module G, where p and p gt are the object confidence of the candidate region and the ground truth label respectively;

[0079] is the bounding box regression loss of the region proposal module G, where p and p gt are the object confidence of the candidate region and the ground truth label respectively, and t i and are the coordinate encodings of the candidate region and the ground truth bounding box relative to the anchor box respectively, that is, when it is less than 1, the value of this expression is when it is greater than or equal to 1, the value of this expression is

[0080] is the classification loss of the fine-grained detection and recognition module D, where c q represents the category to which the query RoI feature belongs, and c represents the category to which the aggregated support RoI feature belongs, and represent the local similarity, global similarity, and cross similarity between the predicted editing box and the category c respectively;

[0081] is the bounding box regression loss of the fine-grained detection and recognition module D, where c q represents the category to which the query RoI feature belongs, and c represents the category to which the aggregated support RoI feature belongs, and t i and are the coordinate encodings of the predicted box and the ground truth bounding box relative to the candidate region respectively, that is, q is 1 when c

[0082] 3.1.3) Solve the loss in 3.1.2) The gradient of the network parameter θ0 before the training iteration of the small-sample SAR object detection and recognition network base class

[0083]

[0084] Among them and are the classification loss and the bounding box regression loss of the region proposal module G respectively, and They are the classification loss and the bounding box regression loss of the fine-grained detection and recognition module D respectively;

[0085] 3.1.4) According to the gradients obtained in 3.1.3) Update the parameters of the small sample SAR target detection and recognition network to obtain the network parameters θ' after iteration in the current base class training stage:

[0086]

[0087] where θ0 is the network parameter before iteration in the base class training stage of the small sample SAR target detection and recognition network, and lr base is the learning rate during base class training. In this example, lr base is taken as 0.005;

[0088] 3.2) Repeat step 3.1) until the network converges to obtain the final iterated network parameter θ base , and complete the base class training of the small sample SAR target detection and recognition network;

[0089] 3.3) Perform small sample fine-tuning on the small sample SAR target detection and recognition network that has completed base class training:

[0090] 3.3.1) Input partial support images of the base class support set, query set images, and all new class support set and query set images into the small sample SAR target detection and recognition network that has completed base class training, calculate the loss value for each iteration, and update the network parameters using this loss value through the stochastic gradient descent algorithm;

[0091] 3.3.2) Calculate the loss of the small sample SAR target detection and recognition network: where and are the classification loss and the bounding box regression loss of the region proposal module G respectively, and are the classification loss and the bounding box regression loss of the fine-grained detection and recognition module D respectively;

[0092] 3.3.3) Solve the loss in 3.3.2) for the gradient of the network parameter θ1 before iteration in the small sample fine-tuning stage of the small sample SAR target detection and recognition network

[0093]

[0094] where and are the classification loss and the bounding box regression loss of the region proposal module G respectively, and They are the classification loss and the bounding box regression loss of the fine-grained detection and recognition module D respectively;

[0095] 3.3.4) According to the gradients solved in 3.3.3) Update the parameters of the small-sample SAR target detection and recognition network to obtain the network parameters θ″ after iteration in the current small-sample fine-tuning stage:

[0096]

[0097] Among them, θ1 is the network parameter before iteration in the small-sample fine-tuning stage of the small-sample SAR target detection and recognition network, and lr ft is the learning rate in the small-sample fine-tuning stage. In this example, lr ft takes 0.0025.

[0098] 3.4) Repeat step 3.3) until the network converges to obtain the final iterated network parameter θ ft , and complete the small-sample fine-tuning of the small-sample SAR target detection and recognition network.

[0099] Step Four: Obtain the SAR target detection and recognition result.

[0100] Input the SAR image to be detected and recognized into the fine-tuned small-sample SAR target detection and recognition network to obtain the detection and recognition result.

[0101] The effect of the present invention can be further illustrated by the following simulation experiments:

[0102] I. Simulation experiment conditions

[0103] The software platform for the simulation experiment of the present invention is the Windows11 operating system and Pytorch 1.10.1, and the hardware configuration is: Core i7-11800H CPU and NVIDIA GeForce RTX 3080 Laptop GPU.

[0104] The simulation experiment of the present invention uses the measured data of GF-3 SAR. The scene type is an airport, the image resolution is 1m×1m, and there are seven types of aircraft targets, namely: A220, A320-321, A330, ARJ21, Boeing737-800, Boeing787, and other. ARJ21 and Boeing787 are used as new-class targets, and the rest are used as base-class targets.

[0105] The number of SAR images is 2000, the image sizes are 600×600, 1024×1024, and 2048×2048, the total number of targets is 6556, the number of training set images is 1400, and the number of test set images is 600.

[0106] II. Simulation Content and Result Analysis

[0107] Under the above simulation conditions, the two SAR images in the test set images were detected and recognized using the present invention and the existing "A Small-Sample Target Detection Method and System for SAR Images Based on Meta-Learning", and the detection and recognition results were visualized on the test set images. The results are as Figure 3 shown. Among them:

[0108] Figure 3 (a) shows the detection and recognition results of the above two SAR images by the prior art,

[0109] Figure 3 (b) shows the detection and recognition results of the above two SAR images by the present invention.

[0110] Figure 3 The solid rectangular boxes in the figure indicate the correctly detected and recognized targets, the dashed rectangular boxes indicate the incorrectly detected and recognized targets, and the circles indicate the missed targets.

[0111] Comparing Figure 3 (a) and 3(b), it can be seen that there are more false alarms and missed detections in the detection and recognition results obtained by the prior art, while there are fewer false alarms and missed detections in the detection and recognition results obtained by the present invention.

[0112] The base-class average precision bAP and the new-class average precision nAP of the test results of the present invention and the prior art were calculated respectively under the conditions of 5-shot for 5 new-class targets, 10-shot for 10 new-class targets, 20-shot for 20 new-class targets, and 30-shot for 30 new-class targets, as shown in Table 1:

[0113] Table 1 Comparison of the base-class average precision bAP and the new-class average precision nAP between the present invention and the prior art

[0114]

[0115] As can be seen from Table 1, under the experimental settings of 5-shot, 10-shot, 20-shot, and 30-shot, both the base-class average precision and the new-class average precision of the present invention are higher than those of the prior art.

Claims

1. A few-shot SAR target detection and recognition method based on meta-learning and metric learning, characterized in that, It includes the following steps: (1) Construct a support set and a query set: (1a) Collect a large number of SAR images containing base-class targets and a small number of SAR images containing new-class targets, and label the target positions and categories in each SAR image; (1b) Use the collected SAR images of base-class targets as the query set of the base class, and randomly sample some SAR images from the collected SAR images of base-class targets to form the support set of the base class; Use all the collected SAR images of new-class targets as the query set and the support set of the new class respectively; (2) Construct a few-shot SAR target detection and recognition network based on meta-learning and metric learning: (2a) Establish a data preprocessing module consisting of a query image preprocessing process that sequentially performs multi-scale transformation, random flipping, and data normalization and a support image preprocessing process that sequentially performs target cropping, random flipping, and data normalization; (2b) Establish a feature extraction module A composed of weight-sharing siamese networks; (2c) Establish a region proposal module G composed of a cascaded feature aggregation sub-module, an anchor box generation sub-module, a positive and negative sample assignment sub-module, a classification and regression sub-module, and a post-processing sub-module, and use the cross-entropy loss function and the Smooth L1 loss function as the classification loss and the bounding box regression loss (2d) Establish a fine-grained detection and recognition module D composed of a region of interest extraction sub-module, a local relationship sub-module, a global relationship sub-module, and an intersection relationship sub-module. Its output is the bounding box coordinates (x, y, w, h) of the target to be inspected and the similarity s of each category of the target, and these two parameters are respectively substituted into the cross-entropy loss function and the Smooth L1 loss function to calculate the classification loss value. and the bounding box regression loss value (2e) Cascade the modules established in (2a), (2b), (2c), and (2d) in sequence to form a small-sample SAR target detection and recognition network, and define the loss function of this network as: (3) Conduct base-class training on the few-shot SAR target detection network: (3a) Input the support set and query set of the base class into the few-shot SAR detection and recognition network, and calculate the loss value for each iteration. Where i is the number of iterations during the training of the base class, and update the network parameters according to this loss value through the stochastic gradient descent algorithm. (3b) Repeat the process of (3a) until the network converges to obtain a preliminarily trained few-shot SAR target detection and recognition network; (4) Conduct few-shot fine-tuning on the few-shot SAR target detection network: (4a) Input the partial support images of the base class support set, the query set images, and all new class support set and query set images into the pre-trained few-shot SAR target detection network, and calculate the loss value for each iteration. Where j represents the number of iterations in the few-shot fine-tuning stage, and use this loss value to update the network parameters through the stochastic gradient descent algorithm; (4b) Repeat the process of (4a) until the network converges to obtain the finally trained few-shot SAR target detection and recognition network; (5) Input the SAR image to be detected and recognized into the finally trained few-shot SAR target detection and recognition network obtained in step (4) to obtain the detection and recognition result.

2. The method according to claim 1, wherein In step (2b), a feature extraction module A composed of weight-sharing siamese networks is established, which includes a siamese network composed of two backbone networks with the same structure and weight sharing. Each backbone network includes four cascaded convolutional modules a1, a2, a3, and a4; The first convolutional module a1 is composed of a cascaded 7×7 standard convolutional layer, a batch normalization layer, a ReLU activation layer, and a max-pooling downsampling layer; The second convolutional module a2 is composed of 3 residual blocks; The third convolutional module a3 is composed of 4 cascaded residual blocks; The fourth convolutional module a4 is composed of 6 cascaded residual blocks; The query feature map output by the entire feature extraction module, denoted as where is the query feature map output by the i-th convolutional module a i and X Q is the input query SAR image; The support feature map output by the entire feature extraction module, denoted as where is the support feature map output by the i-th convolutional module a i and X S is the input support SAR image.

3. The method according to claim 2, wherein Each residual block in the four cascaded convolutional modules includes 3 convolutional modules c1, c2, and c3; The first convolutional module c1 is composed of a cascaded 1×1 convolutional layer, a batch normalization layer, and a ReLU activation layer; The second convolutional module c2 is composed of a cascaded 3×3 convolutional layer, a batch normalization layer, and a ReLU activation layer; The third convolutional module c3 is composed of a cascaded 1×1 convolutional layer and a batch normalization layer; The output feature map of each convolutional module in the residual block, denoted as where x is the input feature map of the residual block; The output feature map of the residual block is expressed as z = relu(y3 + x), where relu(·) represents the ReLU activation function.

4. The method according to claim 1, characterized in that In step (2c), the structures and functions of each sub-module in the region proposal module G are as follows: The feature aggregation sub-module uses depthwise separable convolution to complete the aggregation of the support feature map and the query feature map, and its output is: Among them, is the query feature obtained by the feature extraction module A. W and H respectively represent the width and height of the query feature map. is the support feature obtained by the feature extraction module A. K represents the side length of the support feature map, and C represents the depth of the query feature map and the support feature map. The anchor box generation sub-module is used to generate rectangular anchor boxes with scales of 2 2 , 4 2 , 8 2 , 16 2 and 32 2 , and aspect ratios of 1:2, 1:1, and 2:1 at each point of the aggregated feature map P; The positive and negative sample assignment sub-module is used to assign positive and negative samples to the anchor boxes, that is, the anchor boxes with an intersection over union (IoU) greater than 0.7 with the labeled bounding boxes of the same category as the support feature are assigned as positive samples, the anchor boxes with an IoU less than 0.3 with the labeled bounding boxes of the same category as the support feature are assigned as negative samples, and the remaining anchor boxes are assigned as irrelevant samples; The classification and regression sub-module is used to obtain candidate regions that may contain the target, which consists of a 3×3 convolutional layer g1 and two parallel 1×1 convolutional layers g cls and g reg cascaded, where g reg is used to obtain the bounding box offsets of the candidate regions relative to the anchor boxes, and g cls is used to predict the confidence that the candidate region contains the target; The post - processing sub - module is used to filter redundant candidate regions and output the N candidate regions with the highest confidence levels \(\{(p j , b j )|j = 1, 2,\cdots, N\}\), where \(p j \) is the confidence level of the \(j\) - th candidate region containing the target, and \(b j =(x j , y j , w j , h j )\) is the bounding box of the \(j\) - th candidate region, \((x j , y j )\) represents the center coordinates of the bounding box of the \(j\) - th candidate region, and \(w j \) and \(h j \) respectively represent the width and height of the bounding box of the \(j\) - th candidate region.

5. The method according to claim 1, wherein In step (2d), the structural parameters of each sub-module in the fine-grained detection and recognition module D are as follows: The region of interest extraction sub-module, which consists of an adaptive average pooling layer, is used to average pool the query feature map Y extracted by the feature extraction module A according to the candidate regions provided by the region proposal module G Q into 7×7 query RoI features and average pool the support feature map Y extracted by the feature extraction module A s into 7×7 support RoI features where C is the depth of the feature map; The local relationship sub-module is used to measure the local feature similarity s between the query RoI feature and the support RoI feature l , which is composed of a cascaded structure of a 1×1 convolutional layer, a depthwise separable convolutional layer, and a fully connected layer with an output feature number of 1; The global relationship sub-module is used to measure the global similarity s between the query RoI feature and the support RoI feature g , and this module is composed of a cascade of a global average pooling layer, a feature splicing layer, and a fully connected layer with an output feature number of 1; The cross-relationship sub-module is used to measure the cross-relationship similarity between the query RoI feature and the support RoI feature and obtain the predicted bounding box coordinates. This module consists of a feature splicing layer, a 1×1 standard convolutional layer, a 3×3 standard convolutional layer, a 1×1 standard convolutional layer, a global average pooling layer, and two parallel fully-connected layers d with 1 output feature and 4 output features respectively. cls The fully-connected layer d with 4 output features reg is cascaded. Among them, d cls is used to predict the cross-relationship similarity s between the query RoI feature and the support RoI feature p , and d reg is used to predict the bounding box coordinates (x, y, w, h), where (x, y) represents the center point coordinates of the bounding box, and w and h are the width and height of the bounding box respectively. The above-mentioned region of interest extraction sub-module is cascaded with the parallel local relationship sub-module, global relationship sub-module, and cross-relationship sub-module to form the fine-grained detection and recognition module D.

6. The method according to claim 5, wherein In step (2d), the fine-grained recognition module D outputs the similarity s of each category of the detected target, and the formula is as follows: wherein, and are the local similarity, global similarity, and cross similarity of the j-th category of the target, respectively, and are expressed as follows: Among them, represents querying the RoI features, represents the support RoI features of the j-th category, C is the depth of the feature map, σ(·) represents the Sigmoid activation function, FC(·) represents a fully connected layer with an output feature number of 1, DWConv(·, ·) represents depthwise separable convolution, Conv(·) represents a 1×1 standard convolution, Cat(·, ·) represents concatenating the feature maps by channels, GAP(·) represents global average pooling, and Conv3(·) represents a cascaded 1×1 standard convolution layer, 3×3 standard convolution layer, and 1×1 standard convolution layer.

7. The method according to claim 1, characterized in that, Calculating the classification loss value in step (2d) and the bounding box regression loss value The formulas are as follows: Among them, c q represents the category to which the query RoI feature belongs, and c represents the category to which the aggregated support RoI feature belongs. and respectively represent the local similarity, global similarity, and cross similarity between the predicted editing box and category c. t i and are the coordinate encodings of the predicted box and the true bounding box relative to the candidate region respectively. That is, it is 1 when c q is equal to c, and 0 when they are not equal. That is when it is less than 1, the value of this expression is when it is greater than or equal to 1, the value of this expression is 8. The method according to claim 1, wherein The classification loss in step (2c) and the bounding box regression loss are respectively expressed as follows: Among them, p and p gt are the target confidence and the true label of the candidate region, respectively, and t i and are the coordinate encodings of the candidate region and the true bounding box relative to the anchor box, respectively.

9. The method according to claim 1, characterized in that, In step (3a), the network parameters are updated using the stochastic gradient descent algorithm during the base class training stage, as follows: (3a1) Solve the gradient of the small sample SAR target detection and recognition network parameters, expressed as: Among them, is the loss in the training stage of the base class of the small-sample SAR target detection and recognition network, and are the classification loss and bounding box regression loss of the region proposal module G in the base class training stage, respectively, and are the classification loss and bounding box regression loss of the fine-grained detection and recognition module D in the base class training stage, respectively. θ0 is the network parameter before iteration in the base class training stage of the small-sample SAR target detection and recognition network; (3a2)According to the solved gradient Update the network parameters of the small-sample SAR target detection and recognition network to obtain the network parameters θ′ after iteration in the current base class training stage: where θ0 is the network parameter before iteration in the base class training stage of the small sample SAR target detection and recognition network, and lr base is the learning rate during base class training, and lr base is the learning rate during base class training.

10. The method according to claim 1, wherein In step (4a), the network parameters are updated using the stochastic gradient descent algorithm during the small sample fine-tuning stage, as follows: (4a1) Solve the gradient of the small sample SAR target detection and recognition network parameters, expressed as: Among them, is the loss in the small-sample fine-tuning stage of the small-sample SAR target detection and recognition network, and are the classification loss and bounding box regression loss of the region proposal module G in the small-sample fine-tuning stage respectively, and are the classification loss and bounding box regression loss of the fine-grained detection and recognition module D in the small-sample fine-tuning stage respectively, and θ1 is the network parameter before iteration in the small-sample fine-tuning stage of the small-sample SAR target detection and recognition network; (4a2)According to the solved gradient Update the network parameters of the small-sample SAR target detection and recognition network to obtain the network parameters θ″ after iteration in the current base class training stage: Among them, θ1 is the network parameter before iteration in the small-sample fine-tuning stage of the small-sample SAR target detection and recognition network, lr ft is the learning rate in the small-sample fine-tuning stage, lr ft is the learning rate in the small-sample fine-tuning stage.

Citation Information

Patent Citations

  • A Small-Sample Object Detection Method and System for SAR Images Based on Meta-Learning

    CN115578592B

  • Remote sensing image small sample target detection method based on prototype convolutional neural network

    CN112861720A

  • SAR (Synthetic Aperture Radar) target detection and identification method based on multistage enhancement network

    CN115909086A