Metalearning-based small sample SAR target detection and identification method and system
By building a local contrast aggregation network, using global cross attention and local cross attention mechanisms, the problems of insufficient feature extraction and insufficient local feature fusion in small sample SAR target detection are solved, and higher detection accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510351263.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-22
Smart Images

Figure CN120355937A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of target detection and recognition, and further relates to a small-sample SAR target detection and recognition method and system, which can be used for battlefield reconnaissance and situation awareness. Background Art
[0002] Synthetic Aperture Radar (SAR) has become an important means of earth observation due to its all-weather and all-day characteristics, and plays an important role in fields such as battlefield reconnaissance and situation awareness. As an important part of SAR image interpretation, SAR target detection has received increasing attention in recent years and has made remarkable progress.
[0003] In recent years, with the rapid development of deep learning theories and methods, SAR target detection has gradually shifted from traditional Constant False Alarm Rate (CFAR) methods to deep learning-based methods, and significant breakthroughs have been achieved in both detection speed and accuracy. However, deep learning-based target detection methods highly rely on a large amount of labeled data. Under the condition of scarce training samples, the model is prone to severe overfitting, resulting in a decline in detection performance.
[0004] In response, researchers have proposed various solutions to improve the target detection performance under small-sample conditions. These methods can be roughly divided into two categories: meta-learning-based methods, that is, designing the network structure through support samples and enhancing the detection performance of new category targets by aggregating support samples and query samples; transfer learning-based methods, improving the detection performance of target categories with only scarce samples through different fine-tuning methods. Since the small-sample target detection method based on meta-learning has the advantages of quickly adapting to new tasks and strong generalization ability, it has become a research hotspot in the field of SAR target detection.
[0005] The Meta R-CNN method proposed by Xiaopeng Yan et al. in 2019 at the IEEE / CVF International Conference on Computer Vision, titled "Meta R-CNN: Towards General Solver for Instance-Level Low-Shot Learning", is designed based on the meta-learning framework and mainly consists of three steps: First, the image is input into Faster R-CNN or Mask R-CNN to extract deep features, and the Region Proposal Network (RPN) is used to generate candidate regions that may contain the target. Then, the Prediction Head Reshaping Network (PRN) receives few-shot targets and their bounding boxes or mask images, calculates their class attention vectors, and then uses meta-learning to train the Prediction Head Reshaping Network. Finally, the Prediction Head Reshaping Network outputs the target detection box and the class of the target to achieve target detection and recognition. Although Meta R-CNN can effectively improve the adaptability of the model to new tasks in the few-shot scenario, it extracts target features insufficiently and lacks effective fusion of the local features of support samples and query samples, resulting in overlapping distributions of different-class targets in the feature space and poor separability.
[0006] The patent document with the application number CN202310639895.2 discloses "A Few-Shot SAR Target Detection and Recognition Method Based on Meta-Learning and Metric Learning". It is divided into three parts. First, a weight-sharing Siamese network is used to extract deep features. Then, a region proposal module is used to aggregate the support feature map and the query feature map. Finally, a fine-grained detection and recognition module is used to predict the position and class of the target. Although this method effectively improves the detection and recognition performance of the deep network model under few-shot conditions, it has the following problems: 1) The feature extraction module extracts features insufficiently, and the separability of different-class targets in the feature space is poor. 2) Only multiplication operations are performed when aggregating the support sample features and the query sample features, lacking effective fusion of the local features of support samples and query samples, and the detection accuracy and robustness are insufficient. 3) The loss function of the network lacks the contrast loss between the support sample features and the query sample features, resulting in poor separability between different-class targets. Summary of the Invention
[0007] The object of the present invention is to propose a few-shot SAR target detection and recognition method and system based on meta-learning in view of the above deficiencies of the prior art, so as to improve the separability of different-class targets in the feature space, effectively fuse the features of support samples and query samples, and enhance the accuracy and robustness of the detection algorithm.
[0008] The technical solutions to achieve the object of the present invention are as follows:
[0009] 1. A few-shot SAR target detection and recognition method based on meta-learning, characterized by comprising:
[0010] (1) Obtain multi-class target SAR images as the training set, and divide the base class support dataset, base class query dataset, new class support dataset, and new class query dataset;
[0011] (2) Construct a local contrast aggregation network including a candidate region feature extraction unit, a local feature aggregation unit, a meta-contrast loss unit, and a detection head unit:
[0012] The local feature aggregation unit is used to effectively aggregate the local features of the support image and the query image, and includes a global cross-attention module and a reweighted local cross-attention module;
[0013] The meta-contrast loss unit is used to calculate the contrast loss between the local features of the support image and the local features of the query image;
[0014] (3) Set the loss function of the local contrast aggregation network, and perform iterative training on it using the training dataset;
[0015] (4) Input the test data into the trained local contrast aggregation network, and output the detection and recognition results of the SAR target.
[0016] Furthermore, the construction of the local contrast aggregation network including a candidate region feature extraction unit, a local feature aggregation unit, a meta-contrast loss unit, and a detection head unit is realized as follows:
[0017] (2a) Select an existing candidate region feature extraction module to extract the candidate region feature F of the query image q and the candidate region feature F of the support image s ;
[0018] (2b) Establish a local feature aggregation unit composed of a global cross-attention module and a reweighted local cross-attention module connected:
[0019] The global cross-attention module inputs the candidate region feature F of the support sample q and the candidate region feature F of the query sample s , and calculates the global affinity matrix
[0020] The reweighted local cross-attention module inputs the candidate region feature F of the support sample q and the candidate region feature F of the query sample s , and is weighted by the global affinity matrix to obtain the aggregated local feature F;
[0021] (2c) Establish a meta-contrast loss unit for calculating the meta-contrast loss between the candidate region features of the support samples and the candidate region features of the query samples.
[0022] (2d) Select an existing detection head unit for obtaining the category and detection box of the target.
[0023] (2e) Connect the candidate region feature extraction unit with the meta-contrast loss unit, and then cascade the output end of the candidate region feature extraction unit with the local feature aggregation unit and the detection head unit in sequence to form a local contrast aggregation network.
[0024] 2. A few-shot SAR target detection and recognition system based on meta-learning, characterized by comprising:
[0025] Dataset generation module: used for obtaining multi-class target SAR images and dividing them into a base class dataset and a novel class dataset.
[0026] Image preprocessing module: used for performing different image preprocessing operations on the query image and the support image to enhance the generalization ability of the model.
[0027] Network construction module: used for constructing a local contrast aggregation network to extract and aggregate features.
[0028] Network training module: used for training the local contrast aggregation network.
[0029] Network inference module: used for performing feature extraction and matching on the input image through the trained local contrast aggregation network to achieve target detection and recognition.
[0030] Furthermore, the network construction module includes:
[0031] Candidate region feature extraction sub-module: used for extracting the global features of the image using a deep network and obtaining the candidate region features using a candidate region proposal network.
[0032] Local feature aggregation sub-module: used for effectively aggregating the candidate region features of the query samples and the candidate region features of the support samples on the basis of the candidate regions, ensuring that the query samples can focus on the correct support local features and effectively avoiding the interference of incorrect categories.
[0033] Meta-contrast loss sub-module: used for further optimizing the feature representation, constraining the projection of the target in the deep feature space, enhancing the separability between different category targets, and effectively improving the accuracy of few-shot detection and recognition.
[0034] Detection head sub-module: used for predicting the category and location of the target using the optimized candidate region features to achieve accurate target detection and recognition.
[0035] Compared with the prior art, the present invention has the following advantages:
[0036] Since the present invention designs a local feature aggregation unit, through the cross-attention mechanism, it can achieve effective aggregation between the features of the support sample candidate region and the query image candidate region, and at the same time effectively avoid the interference of incorrect categories.
[0037] Since the present invention designs a meta-contrast loss unit, by constraining the projection of the target in the deep feature space, it can enhance the separability between different category targets, thereby effectively improving the accuracy of few-shot detection and recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 is the implementation flowchart of the few-shot SAR target detection and recognition method based on meta-learning of the present invention;
[0039] Figure 2 is Figure 1 the structural diagram of the local contrast aggregation network in
[0040] Figure 3 is Figure 2 the structural diagram of the local feature aggregation module in
[0041] Figure 4 is the structural block diagram of the few-shot SAR target detection and recognition system based on meta-learning of the present invention;
[0042] Figure 5 is the comparison diagram of the simulation results of the present invention and the prior target detection and recognition methods. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be described more clearly and completely below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0044] It should be noted that the step numbers in the specification and claims of the present invention are only for clear description of the implementation solutions of the present invention for easy understanding, and their sequence numbers are not limited.
[0045] Embodiment 1: Few-shot SAR Target Detection and Recognition Method Based on Meta-Learning
[0046] Referring to Figure 1 , the implementation steps of this embodiment include the following:
[0047] Step 1, generate a data set.
[0048] 1.1) Obtain multiple SAR images containing multiple types of targets from different datasets, and mark the positions and types of the targets in each SAR image;
[0049] 1.2) Divide the SAR images and their labels into a base class dataset and a new class dataset according to the types;
[0050] 1.3) Use the images in the base class dataset as the query set of the base class, and sample the target regions of some images in the base class dataset to form the support set of the base class; Use all the images in the new class dataset as the query set and support set of the new class respectively.
[0051] In this embodiment, the SAR-Aircraft-1.0 dataset is used. The SAR data comes from the Gaofen-3 satellite, which includes 4 different sizes of 800×800, 1000×1000, 1200×1200, and 1500×1500, and it includes 7 types of SAR aircraft targets: A220, A320 / 321, A330, Boeing787, Other, ARJ21, and Boeing737;
[0052] Regard the 5 types of targets A220, A320 / 321, A330, Boeing787, and Other as the base class targets;
[0053] Regard the 2 types of targets ARJ21 and Boeing737 as the new class targets.
[0054] Step 2, preprocess the images in the dataset.
[0055] 2.1) Sequentially perform channel normalization, multi-scale transformation, random rotation, and random flipping on the images in the base class query set to obtain the preprocessed base class query images;
[0056] 2.2) Sequentially perform target sampling, channel normalization, random rotation, and random flipping on the images in the base class support set to obtain the preprocessed base class support images;
[0057] 2.3) Sequentially perform channel normalization, multi-scale transformation, random rotation, and random flipping on the images in the new class query set to obtain the preprocessed new class query images;
[0058] 2.4) Sequentially perform target sampling, channel normalization, random rotation, and random flipping on the images in the new class support set to obtain the preprocessed new class support images;
[0059] In this embodiment, the scales of the multi-scale transformation are 800×800, 1000×1000, 1120×1120, and 1200×1200, the probability of random flipping is 0.5, the probability of random rotation is 0.5, and the range is [-90°, 90°].
[0060] Step 3: Construct a local contrast aggregation network.
[0061] Refer to Figure 2 , and the implementation of this step includes the following:
[0062] 3.1) Set a candidate region feature extraction module composed of two branches, namely a query image branch and a support image branch:
[0063] The query image branch is used to obtain candidate regions using an existing region proposal network on the basis of extracting deep features, further extract candidate region features, and align the features using a candidate region alignment module to obtain the candidate region feature F q ;
[0064] The support image branch is used to directly extract candidate region features on the deep features using the ground truth as candidate regions, and align the features using a candidate region alignment module to obtain the candidate region feature F s ;
[0065] In this embodiment, the existing Faster R-CNN is selected as the backbone network to extract the deep features of the input image, and the query image branch and the support image branch share the backbone network parameters.
[0066] 3.2) Establish a local feature aggregation module for obtaining the aggregated local feature F:
[0067] Refer to Figure 3 , and the implementation of this step includes the following:
[0068] 3.2.1) Set a global cross-attention sub-module for calculating the global affinity matrix :
[0069] Perform global average pooling on the candidate region feature F q of the query image and the candidate region feature F s of the support image respectively to obtain the instance-level query semantic vector f q and the support semantic vector f s ;
[0070] Use the value space projection matrix W V to project the query semantic vector f q into the value space to obtain the value space feature W V fq , project the support semantic vector f K using the key space projection matrix W s onto the key space to obtain the feature W in the key space K f s ;
[0071] Obtain the global affinity matrix based on the spatial features of value and key
[0072]
[0073] where d k is the feature depth, W V and W K are the projection matrices projected onto the value space and the key space respectively, and softmax is the activation function;
[0074] 3.2.2) Set the reweighted local cross-attention sub-module for calculating the aggregated local feature F:
[0075] Flatten both the candidate region features F q of the query image and the candidate region features F s of the support image to obtain the query semantic vector and the support semantic vector
[0076] Use the value space projection matrix to project the support semantic vector onto the value space to obtain the value space feature Use the key space projection matrix to project the support semantic vector onto the key space to obtain the feature in the key space
[0077] Multiply the value space feature with the key space feature to obtain the local affinity matrix
[0078]
[0079] where T represents matrix transpose;
[0080] Use the global affinity matrix to weight the local affinity matrix to obtain the reweighted local affinity matrix
[0081]
[0082] Among them, Multiply is the element-wise multiplication of matrices, and softmax is the activation function;
[0083] The reweighted local affinity matrix and the value space features are multiplied by matrix, and added to the candidate region features F of the query image q to obtain the aggregated local features F:
[0084]
[0085] Among them, is the projection matrix projected into the value space.
[0086] 3.3) Set the meta-contrast loss calculation module:
[0087] Obtain all candidate region feature vectors of the query image and the prototype features of the category y i Calculate the contrast loss between the candidate region features of the support samples and the candidate region features of the query samples The formula is as follows:
[0088]
[0089]
[0090] Among them, i is the candidate region number, is the feature vector of the candidate region of the query image, is the prototype feature of its corresponding category y i i is the intersection over union of the candidate region and the ground truth bounding box, N proposal is the number of candidate regions, N ways is the number of categories, i is the normalized feature vector of the candidate region of the query image, is the normalized prototype feature of the category y i ways is the temperature coefficient;
[0091] In this embodiment, the temperature coefficient τ is set to 0.1.
[0092] 3.4) Set the detection head module for calculating the category and bounding box of the target:
[0093] Calculate the cosine similarity between the aggregated local features F and the prototypes of each category to obtain the confidence levels of different categories, and take the one with the highest confidence level as the predicted target category;
[0094] Input the aggregated local feature F into the fully connected layer to predict the offsets of the target bounding boxes and the candidate regions, and then calculate the predicted target bounding boxes.
[0095] 3.5) Connect one output end of the candidate region feature extraction module to the meta-contrast loss module, and cascade the other output end of the candidate region feature extraction module with the local feature aggregation module and the detection head module in sequence to form a local contrast aggregation network.
[0096] Step 4: Train the local contrast aggregation network.
[0097] 4.1) Conduct a base class training on the local contrast aggregation network:
[0098] 4.1.1) Set the learning rate lr of the local contrast aggregation network in the base class training phase base to be 0.001;
[0099] 4.1.2) Input the base class support images and the base class query images into the local contrast aggregation network, perform forward propagation, and calculate the loss function
[0100]
[0101] Among them, is the classification loss of the detection head module, is the localization loss of the detection head module, is the meta-contrast loss, f(λ) represents the weight of the meta-contrast loss. When calculating the loss of the candidate region feature extraction module, f(λ) = λ, and when calculating the loss of the detection head module, f(λ) = 0, set λ = 1;
[0102] In this embodiment, the classification loss uses the cross-entropy loss, expressed as:
[0103]
[0104] Among them, y i is the true class, is the class predicted by the network;
[0105] In this embodiment, the localization loss uses the Smooth L1 loss, expressed as:
[0106]
[0107] Among them, t i is the true target bounding box parameter, is the target bounding box parameter predicted by the network, and the smoothing parameter δ is set to 1;
[0108] 4.1.3) Use the loss function Calculate the gradient of the network parameter θ0
[0109]
[0110] 4.1.4) Use the gradient descent method for the gradient Perform backpropagation and update the network parameter θ0′:
[0111]
[0112] where lr base is the learning rate for base class training;
[0113] 4.2) Repeat step 4.1) until the network converges to obtain a well-trained local contrast aggregation network for the base class;
[0114] 4.3) Fine-tune the well-trained local contrast aggregation network for the base class with few-shot samples:
[0115] 4.3.1) In this embodiment, set the learning rate lr ft for the few-shot fine-tuning stage to 0.0005;
[0116] 4.3.2) Input some base class support images, some base class query images, all new class support images, and all new class query images into the well-trained local contrast aggregation network for the base class for forward propagation, calculate the loss function and use the loss function to calculate the gradient of the network parameter θ0
[0117] 4.3.3) Use the gradient descent method for the gradient to perform backpropagation and update the network parameter θ0″:
[0118]
[0119] where lr ft is the learning rate for fine-tuning;
[0120] 4.4) Repeat step 4.3) until the network converges to obtain a fully trained local contrast aggregation network.
[0121] Step 5, Input the test image into the fully trained local contrast aggregation network to obtain the target detection and recognition result.
[0122] Embodiment 2: Few-shot SAR target detection and recognition system based on meta-learning:
[0123] Refer to Figure 4, this embodiment includes: a dataset generation module 1, an image preprocessing module 2, a network construction module 3, a network training module 4, and a network inference module 5. Among them, the network construction module 3 includes: a candidate region feature extraction module 31, a local feature aggregation module 32, a meta-contrast loss module 33, and a detection head module 34. The working principle is as follows:
[0124] The dataset generation module 1 is used to obtain multi-class target SAR images and divide the base class query dataset, base class query support dataset, new class query dataset, and new class support dataset;
[0125] The image preprocessing module 2 is used to perform different image preprocessing operations on the images of the divided base class query dataset, base class query support dataset, new class query dataset, and new class support dataset to obtain preprocessed training images;
[0126] The network construction module 3 is used to construct a local contrast aggregation network, extract and aggregate features. Among them, the candidate region feature extraction module 31 is used to extract the global features of the image using a deep network and obtain candidate region features using a candidate region proposal network; the local feature aggregation sub-module 32 is used to effectively aggregate the candidate region features of the query sample and the candidate region features of the support sample on the basis of the candidate region, ensuring that the query sample can focus on the correct support local features and effectively avoiding the interference of incorrect categories; the meta-contrast loss sub-module 33 is used to further optimize the feature representation, constrain the projection of the target in the deep feature space, enhance the separability between different category targets, and effectively improve the accuracy of few-shot detection and recognition; the detection head sub-module 34 is used to predict the category and location of the target using the optimized candidate region features to achieve accurate target detection and recognition;
[0127] The network training module 4 is used to input the preprocessed training images into the local contrast aggregation network for training to obtain a fully trained local contrast aggregation network;
[0128] The network inference module 5 is used to perform feature extraction and matching on the input image through the fully trained local contrast aggregation network to achieve target detection and recognition.
[0129] It should be noted that: the above-mentioned functional modules can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a program instruction product. The program instruction product includes one or a set of program instructions. When the program instructions are loaded and executed on a computer, the above-mentioned process or function is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The program instructions can be stored in a computer-readable and writable storage medium, or transmitted from one computer-readable and writable storage medium to another computer-readable and writable storage medium.
[0130] In this embodiment, the direct coupling or communication connection between the modules shown or discussed with each other can be achieved through the indirect coupling or communication connection of some interfaces, devices or modules. Each functional module and sub-module in this embodiment can be dynamically in a processing component, or each module can exist physically alone, or two or more modules can be dynamically in a processing component. When the above dynamic components are implemented in the form of software functional modules and sold or used as independent products, they can also be stored in a computer-readable and writable storage medium. The storage medium can be a memory, a disk or an optical disc, etc.
[0131] The flowchart representation or method representation of the above embodiment can be understood as representing a module, segment or part of code including one or a group of executable instructions configured to implement a specific logical function or process. The scope of the preferred embodiments of the present disclosure includes additional implementations, where the functions may be executed in a substantially simultaneous manner according to the functions involved, rather than in the order shown or discussed.
[0132] The effects of the present invention can be further illustrated by the following simulation results
[0133] I. Simulation conditions
[0134] The software platform for the simulation experiment of the present invention is the Ubuntu 20.04 operating system, the experimental environment is Python 3.8, torch 1.11.0 and CUDA 11.3, and the hardware configuration is an Intel Core i9-14900K processor and an NVIDIA GeForce RTX 4090 graphics card.
[0135] The training data for the simulation experiment uses the SAR-Aircraft-1.0 dataset. The SAR data comes from the Gaofen-3 satellite, which includes 4 different sizes of 800×800, 1000×1000, 1200×1200, and 1500×1500. It includes 7 types of SAR aircraft targets: A220, A320 / 321, A330, Boeing787, Other, ARJ21, and Boeing737. The 5 types of targets A220, A320 / 321, A330, Boeing787, and Other are used as base class targets, and the 2 types of targets ARJ21 and Boeing737 are used as new class targets.
[0136] The test data for the simulation experiment uses two measured SAR images from the Gaofen-3 satellite.
[0137] II. Simulation content
[0138] Under the above simulation conditions, the proposed method and the existing detection and recognition method Meta R-CNN were respectively used to complete training on the training data, and then the test data was input into the trained networks of the proposed method and the existing detection and recognition method respectively to obtain the detection and recognition results. As Figure 5 shown, where the solid rectangular box represents a correct detection, the dashed rectangular box represents a wrong detection or a false alarm, and the circle represents a missed detection.
[0139] From Figure 5 it can be seen that there are more false alarms and missed detections in the existing technical solutions, and there are no false alarms and missed detections in the present invention, indicating that the detection and recognition performance of the present invention is significantly better than that of the existing detection methods.
Claims
1. A few-shot SAR target detection and recognition method based on meta-learning, characterized in that, Including: (1) Obtain multiple types of target SAR images as a training set, and divide the base-class support dataset, base-class query dataset, new-class support dataset, and new-class query dataset; (2) Construct a local contrast aggregation network including a candidate region feature extraction unit, a local feature aggregation unit, a meta-contrast loss unit, and a detection head unit: The local feature aggregation unit is used to effectively aggregate the local features of the support image and the query image, and includes a global cross-attention module and a reweighted local cross-attention module; The meta-contrast loss unit is used to calculate the contrast loss between the local features of the support image and the local features of the query image; (3) Set the loss function of the local contrast aggregation network and perform iterative training on it using the training dataset; (4) Input the test data into the trained local contrast aggregation network and output the detection and recognition results of the SAR target.
2. The method according to claim 1, characterized in that, In the above (1), obtaining multiple types of target SAR images and dividing the base-class support dataset, base-class query dataset, new-class support dataset, and new-class query dataset are implemented as follows: (1a) Obtain multiple SAR images and mark the positions and categories of the targets in each SAR image; (1b) Divide the SAR images and their labels into a base-class dataset and a new-class dataset according to the categories; (1c) Use the images in the base-class dataset as the query set of the base class, and sample the target regions of some images in the base-class dataset to form the support set of the base class; Use all the images in the new-class dataset as the query set and support set of the new class respectively.
3. The method according to claim 1, wherein In the above (2), constructing a local contrast aggregation network including a candidate region feature extraction unit, a local feature aggregation unit, a meta-contrast loss unit, and a detection head unit is implemented as follows: (2a) Select an existing candidate region feature extraction module for extracting the candidate region feature F of the query image q and the candidate region feature F of the support image s ; (2b) Establish a local feature aggregation unit composed of a global cross-attention module and a reweighted local cross-attention module connected; The global cross-attention module takes as input the candidate region features F of the support sample q and the candidate region features F of the query sample s to calculate the global affinity matrix The reweighted local cross-attention module takes as input the candidate region features F of the support sample q and the candidate region features F of the query sample s , and weights them using the global affinity matrix to obtain the aggregated local features F; (2c) Establish a meta-contrast loss unit for calculating the meta-contrast loss between the candidate region features of the support samples and the candidate region features of the query samples (2d) Select an existing detection head unit to obtain the category and detection box of the target; (2e) Connect one output end of the candidate region feature extraction unit to the meta-contrast loss unit, and connect the other output end of the candidate region feature extraction unit to the local feature aggregation unit and the detection head unit in cascade to form a local contrast aggregation network.
4. The method according to claim 3, characterized in that, The global cross-attention module described in step (2b) calculates the global affinity matrix, and its implementation includes the following: (2b1) Respectively perform global average pooling on the candidate region features F of the query image q and the candidate region features F of the support image s to obtain the instance-level query semantic vector f q and the support semantic vector f s ; (2b2) Use the value space projection matrix W V Project the query semantic vector f q into the value space to obtain the value space feature W V f q , and use the key space projection matrix W K Project the support semantic vector f s into the key space to obtain the key space feature W K f s ; (2b3) Obtain the global affinity matrix based on the spatial characteristics of value and the spatial characteristics of key where d k is the feature depth, W V and W K are the projection matrices projected onto the value space and the key space respectively, and softmax is the activation function.
5. The method according to claim 3, characterized in that, The reweighted local cross-attention module described in step (2b) uses the global affinity matrix to weight to obtain the aggregated local features, and its implementation includes the following: (2b4) Flatten the candidate region features F of the query image q and the candidate region features F of the support image s to obtain the query semantic vector and the support semantic vector (2b5) Use the value space projection matrix Project the support semantic vector onto the value space to obtain the value space feature Use the key space projection matrix Project the support semantic vector onto the key space to obtain the key space feature (2b6) Multiply the features in the value space with the features in the key space to obtain the local affinity matrix where, T represents matrix transpose; (2b7) Use the global affinity matrix to weight the local affinity matrix to obtain a re-weighted local affinity matrix where, Multiply is element-wise multiplication of matrices, and softmax is an activation function; (2b8) Reweight the local affinity matrix with the value space features perform matrix multiplication, and add it to the candidate region features F of the query image q to obtain the aggregated feature F: Among them, The projection matrix projected onto the value space.
6. The method according to claim 3, wherein The meta - contrast loss unit described in step (2c) calculates the contrast loss between the candidate region features of the support samples and the candidate region features of the query samples The formula is as follows: where i is the candidate region number, is the feature vector of the i-th candidate region in the query image, is its corresponding class y i 's prototype feature, u i is the intersection over union of the candidate region and the ground truth bounding box, N proposal is the number of candidate regions, N ways is the number of classes, is the normalized query feature vector, is the normalized prototype feature of class y i and τ is the temperature coefficient.
7. The method according to claim 1, characterized in that, Set the loss function of the local contrast aggregation network in (3) It is expressed as follows: Among them, is the classification loss of the detection head module, is the localization loss of the detection head module, is the meta-contrast loss, f(λ) represents the weight of the meta-contrast loss, f(λ)=λ when calculating the loss of the candidate region feature extraction module, f(λ)=0 when calculating the loss of the detection head module, and λ = 1 is set.
8. The method according to claim 1, characterized in that In the above (3), training the local contrast aggregation network is implemented as follows: (3a) Image preprocessing: Perform channel normalization, multi-scale transformation, random rotation, and random flipping on the query image in sequence, Perform random sampling, channel normalization, random rotation, and random flipping on the support image in sequence; (3b) Input the preprocessed image into the local contrast aggregation network, and use the loss function to calculate the gradient of the network, perform backpropagation using the gradient descent method, and update the network parameters; (3c) Repeat the process described in (3b) until the loss function converges to obtain the trained local contrast aggregation network.
9. A few-shot SAR target detection and recognition system based on meta-learning, characterized in that, Including: Dataset generation module: used to obtain multi-class target SAR images and divide them into a base class dataset and a new class dataset; Image preprocessing module: used to perform different image preprocessing operations on query images and support images to enhance the generalization ability of the model; Network construction module: used to construct a local contrast aggregation network to extract and aggregate features; Network training module: used to train the local contrast aggregation network; Network inference module: used to extract features and match the input image through the trained local contrast aggregation network to achieve target detection and recognition.
10. The system according to claim 9, characterized in that, The described network construction module includes: Candidate region feature extraction sub-module: used to extract the global features of the image using a deep network and obtain candidate region features using a candidate region proposal network; Local feature aggregation sub-module: used to effectively aggregate the candidate region features of the query sample and the candidate region features of the support sample based on the candidate region, ensuring that the query sample can focus on the correct support local features and effectively avoiding interference from incorrect categories; Meta-contrast loss sub-module: used to further optimize the feature representation, constrain the projection of the target in the deep feature space, enhance the separability between different category targets, and effectively improve the accuracy of few-shot detection and recognition; Detection head sub-module: used to predict the category and location of the target using the optimized candidate region features to achieve accurate target detection and recognition.
Citation Information
Patent Citations
Small sample SAR target detection and identification method based on meta learning and metric learning
CN116664823A