Meta-DETR small sample rotating target detection method based on directed candidate box generation mechanism

By introducing a directed candidate box generation mechanism in the Meta-DETR model, the inaccurate positioning, classification errors and missed detection problems of the model when detecting the rotation target with an angle are solved, and higher detection accuracy and wider application scenarios are achieved.

CN120070860AActive Publication Date: 2025-05-30BEIJING UNIV OF TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510143842.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-05-30
Estimated Expiration
2045-02-10

AI Technical Summary

Technical Problem

The existing Meta-DETR model performs poorly when detecting rotation targets with angles, inaccurate positioning, classification errors and missed detection problems are prominent, especially in the fields of remote sensing, industrial detection and text detection.

Method used

The Meta-DETR small sample rotation object detection method based on the directed candidate box generation mechanism is adopted. Through steps such as feature extraction, position coding, correlation aggregation module, Encoder encoding, directed candidate box generation, Decoder layer and prediction head, more accurate directed candidate boxes are generated to improve the model's performance in rotation object detection.

Benefits of technology

Through the directed candidate box generation mechanism, the model can position the rotating target more accurately, reduce false detection and missed detection, thereby improving the overall detection accuracy and expanding the application scenarios of the model in the angular object detection task.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070860A_ABST
    Figure CN120070860A_ABST
Patent Text Reader

Abstract

The invention discloses a Meta-DETR small sample rotating target detection method based on a directed candidate box generation mechanism. The method comprises the following seven steps: step 1, feature extraction; step 2, position coding; step 3, a correlation aggregation module; step 4, encoding the features by the Encoder; step 5, generating a directed candidate box; step 6, carrying out a Decoder layer; step 7, predicting a head; according to the method, after the directed candidate frame generation mechanism is introduced into the Meta-DETR, the model can more accurately position the target, the conditions of false detection and missing detection are reduced, the overall detection precision is improved, and the directed candidate frame generation mechanism can utilize the direction information of the target, so that feature alignment is more effective. This is particularly important for targets with specific directionality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of object detection, and more specifically, to a small-sample rotated object detection method for Meta-DETR based on a directed candidate box generation mechanism. Background Art

[0002] Object Detection is one of the core tasks in the field of computer vision, aiming to locate and identify target objects in images. In recent years, object detection methods based on deep learning (such as Faster R-CNN, YOLO, DETR, etc.) have made remarkable progress with the support of large-scale labeled datasets (such as COCO, PASCAL VOC). However, these methods usually rely on large-scale labeled data: a large number of labeled samples are required to train the model, and high-performance hardware is needed to support large-scale training.

[0003] Although deep learning models perform excellently in object detection tasks, their reliance on large-scale labeled data brings the following challenges: Labeling object detection data requires precise bounding box annotation for each target object, which is time-consuming and expensive. In some fields (such as medical images and remote sensing images), labeling requires professional knowledge, further increasing the cost. In real-world scenarios, the target categories often exhibit a long-tail distribution, and the samples of minority classes are scarce, making it difficult for the model to learn the features of these classes. In dynamic scenarios, it may be necessary to detect targets of new categories, but the cost of relabeling and retraining the model is too high. Few-shot Object Detection (FSOD) aims to quickly learn the features of new categories through a small number of labeled samples (usually 1 to 10). Few-shot object detection introduces the idea of few-shot learning into the object detection task, with the goals of: quickly learning and detecting targets of new categories with only a small number of labeled samples. Maintaining the detection performance for known categories while having the generalization ability for new categories. The emergence of the Few-shot Object Detection (FSOD) technology provides a new solution to the current challenges. Few-shot object detection methods can achieve precise recognition and accurate localization of targets under limited labeled data. This is mainly due to the innovations of few-shot object detection methods in model design, optimization, and transfer learning, etc. In terms of model design, few-shot object detection methods focus on the lightweight and efficiency of the model. By adopting a concise model structure and effective feature extraction methods, these algorithms can learn the essential features of the target from limited data. At the same time, by introducing techniques such as the attention mechanism, the model can pay more attention to key information. The Transformer model has made a major breakthrough in the field of natural language processing with its unique self-attention mechanism and has gradually shown broad application prospects in the field of computer vision. Especially in the field of few-shot object detection, the Transformer-based methods are gradually becoming the focus of researchers. Different from traditional convolutional neural networks, the Transformer can dynamically focus on different parts of the image and capture the correlations between them, thus performing well in dealing with targets in occluded, overlapping, or complex scenarios. However, despite these advantages of the Transformer model, in the few-shot object detection task, the existing methods still face many challenges. To solve problems such as high model complexity and insufficient detection accuracy, researchers have begun to explore combining the Transformer with advanced technologies such as meta-learning to improve the performance of few-shot object detection.As an important representative in this field, the Meta-DETR model combines meta-learning and the DETR framework, and initially realizes effective object detection under limited data. Through a carefully designed model structure and training strategy, the Meta-DETR model can effectively learn the feature representation of objects under limited data conditions, and then achieve accurate detection of new-class objects. This characteristic makes the Meta-DETR model have broad application prospects in the field of small-sample object detection, thus providing practical methods and means for solving practical problems in related fields. However, the Meta-DETR model usually performs poorly when detecting objects with angles (such as rotated objects, tilted objects). These objects may appear in scenarios such as remote sensing images, industrial inspections, and text detection (such as tilted text). The specific manifestations are as follows: inaccurate positioning: the detection box cannot accurately cover the object with an angle, and usually represents the object with a horizontal bounding box (HBB) instead of a rotated bounding box (RBB). misclassification: due to the feature differences caused by the angle change, the model may have difficulty correctly classifying the object. missed detection problem: objects with angles may be ignored, especially when the object angle is large or the shape is complex. Therefore, the performance of Meta-DETR in the task of detecting objects with angles can be improved to make it more widely applicable in fields such as remote sensing, industrial inspection, and text detection. Summary of the Invention

[0004] In view of this, the embodiments of the present invention provide a Meta-DETR small-sample rotated object detection method based on a directed candidate box generation mechanism to achieve the detection of rotated objects under a small number of samples.

[0005] To achieve the above purpose, the technical solutions and experimental steps adopted by the present invention are as follows:

[0006] The Meta-DETR small-sample rotated object detection method based on a directed candidate box generation mechanism includes the following 7 steps:

[0007] Step 1. Feature extraction;

[0008] The feature extraction stage is divided into two parts: support set feature extraction and query set feature extraction. The task of support set feature extraction is to extract the features of the target region from the support set images, and the task of query set feature extraction is to extract the full-image features from the query set images to provide input for the object detection task. The feature extractors of the support set and the query set are shared to ensure that the feature representations of the two are in the same semantic space.

[0009] Step 2. Position encoding;

[0010] Position encoding enables the model to distinguish features at different positions in the feature map by providing spatial position information for the features.

[0011] Step 3. Correlation Aggregation Module;

[0012] The Correlation Aggregation Module effectively integrates the category information provided by the support set into the feature representation of the query set, enabling the model to accurately locate and identify the targets corresponding to the support set categories in the query image. Specifically, CAM achieves this goal in the following ways:

[0013] Feature Alignment: Ensure that the category features of the support set are aligned with the image features of the query set in the same semantic space.

[0014] Feature Enhancement: Utilize the category information of the support set to enhance the feature expression of specific categories in the query set.

[0015] Context Association: Capture the association relationship between specific regions in the query set and the support set categories through the attention mechanism.

[0016] Step 4. Encoder encodes the features;

[0017] The enhanced query feature map has combined the category information of the support set, but still needs to be further modeled by the Encoder to capture the global relationship between different positions in the query feature map.

[0018] Step 5. Directed Candidate Box Generation;

[0019] When generating the directed candidate box, first initialize a rotated box represented as (p ix , p iy , p iw , p ih , p iθ ), where i represents one of the pixels in the feature map, and (p ix , p iy ) ∈ [0, 1]. We use the prediction result Δr obtained from the directed candidate box generation mechanism and the initial rotated box p i to obtain the final coordinates of each pixel i.

[0020] Step 6. Decoder Layer;

[0021] Take the obtained directed candidate box as Object Queries and input them into the Decoder layer. In the Decoder, interact with the encoded features output by the Encoder to obtain the output embedding containing the target feature information, and then input the output embedding into the prediction head branch.

[0022] Step 7. Prediction Head

[0023] The prediction head branch includes a class prediction head branch and a bounding box prediction. On the basis of these two branches, a rotation angle prediction branch is added in parallel. Using the feature information contained in the embedding vector, a three-layer fully connected network is adopted to implement the prediction of the target rotation angle, and the ReLU activation function is used for non-linear mapping of features before adjacent two layers of fully connected layers.

[0024] Advantages of the present invention:

[0025] After introducing the directed candidate box generation mechanism into Meta-DETR, the model can more accurately locate the target, reduce the situations of false detection and missed detection, thereby improving the overall detection accuracy. And the directed candidate box generation mechanism can utilize the direction information of the target to make the feature alignment more effective. This is particularly important for targets with specific directions. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.

[0027] Figure 1 It is a schematic diagram of the Meta-DETR few-shot rotated object detection method based on the directed candidate box generation mechanism provided by the embodiment of the present invention;

[0028] Figure 2 It is a schematic diagram of the alignment convolution provided by the embodiment of the present invention;

[0029] Figure 3 It is a flowchart of the Meta-DETR few-shot rotated object detection method based on the directed candidate box generation mechanism provided by the embodiment of the present invention;

[0030] Figure 4 It is a schematic diagram of the directed candidate box generation mechanism provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0031] The following further describes the present invention with reference to the drawings.

[0032] Refer to the attached Figure 1 , the implementation steps of the present invention are as follows

[0033] Step 1. Feature extraction

[0034] The feature extraction stage is divided into two parts: support set feature extraction and query set feature extraction. The task of support set feature extraction is to extract the features of the target region from the support set images, and the task of query set feature extraction is to extract the full-image features from the query set images, providing inputs for the object detection task. The weight-sharing feature extractor first encodes them into the same feature space. Subsequently, the Correlation Aggregation Module (CAM), which will be introduced in detail later, performs the matching between the query features and the support class set.

[0035] Step 2. Position Encoding

[0036] The position encoding provides spatial position information for the query set feature F query so that the model can distinguish the features at different positions in the feature map. The formula is as follows.

[0037]

[0038] In the formula, pos represents the position of the embedding vector in the image feature vector, i represents the index in the embedding vector, and d model represents the output dimension of the model. If the dimension index of the embedding vector is even, the result of the sine formula is used; if the index is odd, the result of the cosine formula is used.

[0039] Step 3. Correlation Aggregation Module

[0040] The Correlation Aggregation Module is different from existing aggregation methods because it can aggregate multiple support classes simultaneously, enabling it to capture the inter-class correlations between them to reduce misclassification and enhance model generalization. Given the query and support features, the weight-sharing multi-head attention module first encodes them into the same feature space, and then obtains the prototype of each support class by applying Region of Interest Align (RoIAlign) and then average pooling on the support features. The Correlation Aggregation Module then performs feature matching and encoding matching. Feature matching is completed through a single-head attention mechanism. Given the support set class feature F class ∈R C×d and the query set F query ∈R H×W×d , the matching coefficient is obtained through the following formula:

[0041]

[0042] where W is the linear projection shared by Q and S, which ensures that they are embedded into the same feature space. Subsequently, the output of feature matching is obtained through the following formula:

[0043] Q F =Aσ(F class )⊙F query

[0044] where σ() represents the sigmoid function and ⊙ represents the Hadamard product, and σ(F class ) serves as the feature filter for each support class, aiming to extract class-related features only from the query features. By applying the matching coefficient A to σ(S), we filter out any features that do not match the support class, generating the feature map Q that highlights the objects belonging to the given support class F .

[0045] To achieve meta-learning that requires class-agnostic prediction, a set of predefined task encodings are introduced, and the given support class is mapped to these task encodings so that the final prediction can be made on the task encodings rather than specific classes. The task encoding T ∈ R is implemented using the sine function C×d , which is the same as the positional encoding. The encoding matching uses the same matching coefficient as the feature matching, and the matched encoding Q E is obtained as follows:

[0046] Q E = AT

[0047] Step 4. The Encoder encodes the features;

[0048] The feature map enhanced by CAM has combined the class information of the support set, but still needs to be further modeled by the Encoder to capture the global relationship between different positions in the query feature map. The Encoder globally models the input enhanced query feature map and integrates the context information. At the same time, considering the requirements of few-shot learning, it enhances the modeling ability of the relationship between the support set and the query set. Through the self-attention mechanism, the Encoder can capture the similarities and differences between the support set and the query set, thus helping the model better identify the targets in the query set that belong to the support set class

[0049] Step 5. Generate directed candidate bounding boxes;

[0050] Meta-DETR is a few-shot object detection model based on Transformer, which combines meta-learning and the DETR framework and can complete object detection tasks with a small number of labeled samples. However, the ObjectQueries of the original Meta-DETR are randomly initialized, independent of image features, and mainly target horizontal bounding box (HBB) object detection. It cannot directly handle rotated objects (OBB), and horizontal bounding boxes contain a large amount of background, sometimes even multiple objects because they tend to be densely arranged. This leads to the pooling operation introducing irrelevant features into the object's features, reducing its matching degree with the ground truth box. Moreover, the intersection over union between the horizontal bounding box and the ground truth box is very small, unable to clearly reflect the matching degree between the two. Therefore, it is necessary to generate oriented candidate boxes. We introduce an oriented candidate box generation mechanism that generates more accurate oriented candidate boxes by learning the angle, center, and size of each oriented candidate box and provides them as object queries to the decoder, which can provide a better position prior for the decoder. When generating oriented candidate boxes, first initialize a rotated box represented as (p ix , p iy , p iw , p ih , p iθ ), where i represents one pixel in the feature map, and (p ix , p iy ) ∈ [0, 1]. In the oriented candidate box generation mechanism, there are two branches: the distance branch and the angle branch. The distance branch predicts the distances d = (l, t, r, b) from the initial rotated box to the four boundaries of the target box, where l: the distance to the left boundary, t: the distance to the upper boundary, r: the distance to the right boundary, b: the distance to the lower boundary, and the loss function uses the regression loss on a logarithmic scale. The angle branch predicts the rotation angle θ of the candidate box and is trained using the mean squared error loss. Combining the predicted values of the distance branch and the angle branch, a rough oriented candidate box B = (x, y, w, h, θ) is generated. After the rough candidate box is generated, there is an alignment problem between the features and the candidate box. For this reason, AlignConv (refer to Figure 3 ) is introduced for feature alignment. In AlignConv, the calculation of the offset field is the core step to achieve feature alignment. The goal of AlignConv is to adjust the sampling points of the feature map through the offset field to align them with the geometry of the candidate box (including the center point, width, height, and rotation angle), thereby solving the geometric mismatch problem between the features and the candidate box. The offset field calculation of AlignConv is based on the geometric parameters B = (x, y, w, h, θ) of the candidate box. For a standard 3×3 convolutional kernel, the sampling point set R is regularly distributed. AlignConv constrains the sampling points within the candidate box, and the position r box, determined by the geometric parameter B of the candidate box. The position r of the sampling point within the candidate box box The calculation formula is:

[0051]

[0052] where s i is the stride of the feature map, p is the position index on the feature map, R is the sampling point set offset field of the standard convolution kernel, and O represents the offset of the sampling point relative to the feature map grid. The calculation formula is:

[0053]

[0054] where r box is the position of the sampling point within the candidate box, p is the position index on the feature map, and r is the sampling point of the standard convolution kernel. Through aligned convolution, the sampling points of the feature map are aligned to the boundaries of the candidate box, thus eliminating the geometric mismatch between the feature and the candidate box. After the feature map is aligned by aligned convolution, the coarse candidate box is further optimized by a lightweight fully convolutional network to generate high-quality oriented candidate boxes.

[0055] Step 6. Decoder layer

[0056] The obtained oriented candidate boxes are used as Object Queries and fed into the Decoder layer. In the Decoder, information interaction is carried out with the encoded features output by the Encoder to obtain the output embedding containing the target feature information, and then the output embedding is input into the prediction head branch.

[0057] Step 7. Prediction head

[0058] The prediction head branch includes a class prediction head branch and a bounding box prediction. On the basis of these two branches, a rotation angle prediction branch is added in parallel. Using the feature information contained in the embedding vector, a three-layer fully connected network is used to implement the prediction of the target rotation angle, and the ReLU activation function is used for non-linear mapping of the features before adjacent two layers of fully connected.

[0059] Experiments and analysis

[0060] 1. Experimental conditions

[0061] The hardware environment for the experiments in this chapter is a CPU of Intel(R) Xeon(R) CPU E5-2678 and a GPU of NVIDIA RTX 3090, with 32G of memory. The models in this chapter are developed using the Python language and implemented using the PyTorch deep learning framework, and the system version is Ubuntu18.04.

[0062] 2. Experimental data

[0063] The present invention uses the MAR20 dataset to train and test the model. The MAR20 dataset has the following three characteristics: (1) Large scale: MAR20 is the largest publicly available remote sensing image military aircraft target recognition dataset, containing 3,842 images, 20 types of military aircraft models, and 22,341 instances. In addition, the dataset provides two annotation methods, horizontal bounding boxes and oriented bounding boxes, corresponding to the horizontal box target recognition task and the oriented box target recognition task respectively, which facilitates researchers to evaluate the performance of different algorithms. (2) Small inter-class differences between different model targets: Since these 20 fine-grained categories all belong to the same large aircraft category, different models of aircraft often have similar characteristics, resulting in a high similarity between fine-grained categories. The differences between different aircraft models are mainly reflected in the differences in local components. For example, the difference between the SU-35 and SU-34 aircraft models is the canard in the front of the head. (3) Large intra-class differences for the same model targets: Due to the influence of factors such as climate, season, light, occlusion, and even atmospheric scattering during the acquisition of remote sensing images, the same model targets often have large visual differences and high intra-class variability. In this experiment, the categories A1, A2, A3, A4, A5, A6, A8, A9, A12, A14, A15, A16, A17, A18, A19 in the MAR20 dataset are used as the base classes, a total of 15 categories. The categories A7, A10, A11, A13, A20 are used as new classes, a total of 5 categories. During the few-shot training process, the number of Object Queries is 300, and the number of layers of both the Transformer encoder and decoder is 6. In the base class training stage, the batch size is 4, the initial learning rate of the model is 0.0002, and the training is iterated 50 times, and the learning rate is decayed by 1 / 10 at the 45th iteration. In the few-shot fine-tuning training stage, the initial learning rate for 1-shot is 0.00015, and the training is iterated 700 times, and the learning rate is decayed to 1 / 10 at the 350th and 600th iterations respectively; the initial learning rate for 3-shot is 0.0001, and the training is iterated 600 times, and the learning rate is decayed to 1 / 10 at the 300th and 550th iterations respectively; the initial learning rate for 5-shot and 10-shot is 0.00005, and the training is iterated 500 times, and the learning rate is decayed to 1 / 10 at the 250th and 450th iterations respectively. Under the condition that the IoU threshold is 0.5, the detection performance of each model for rotated targets on the base classes is shown in Table 1. Under the condition that the IoU threshold is 0.5, the detection performance of each model for rotated targets on the new classes is shown in Table 2. Since there is no existing few-shot rotated target detection method, the detection of the original Meta-DETR model on horizontal annotation data is compared with the detection of this model on rotated annotation data.

[0064] 3. Experimental Analysis

[0065] We use the same evaluation metrics as PASCAL VOC 2007 (Everingham et al., 2010), namely the mean Average Precision (mAP), which is the mean of the Average Precision (AP) of different types of targets, and the Average Precision is defined as the area under the precision-recall curve. The calculations of precision and recall are shown as follows:

[0066]

[0067] In the formula, TP represents true positive, FN represents false negative, and FP represents false positive.

[0068] Under the condition that the IoU threshold is 0.5, the detection performance of each model for rotated targets on the base classes is shown in Table 1. The experimental results show that compared with the original Meta-DETR model, the model proposed in this chapter has a 4.6% improvement in the object detection performance on the base classes. To compare with other rotated object detection methods, it is compared with R-CenterNet, and the detection performance of this model is 16.6% higher. The few-shot detection ability of the model is verified on the new classes of the MAR20 dataset. Since the R-CenterNet model has poor applicability in the few-shot object detection task and there are few rotated object detection models based on few-shot, this experiment compares the detection of the original model Meta-DETR on horizontally labeled data with the detection of this model on rotated labeled data. The object detection performance of each model under the condition that the IoU threshold is 0.5 is shown in Table 4-2. The experimental results show that under 1-shot, 3-shot, 5-shot, and 10-shot, the model proposed in this chapter has improvements of 0.1%, 2.3%, 1.0%, and 9.9% respectively compared with the original model. It can be concluded that the model proposed in this chapter can better achieve the detection of few-shot rotated targets and has good few-shot detection performance.

[0069] In summary, the present invention proposes a Meta-DETR few-shot rotated object detection method based on a directed candidate box generation mechanism. Based on Meta-DETR, a new directed candidate box generation mechanism is proposed. By using the generated directed candidate boxes as object queries in the decoder, it provides a position prior for the subsequent features, and an angle branch is designed to achieve the detection of the rotation angle of few-shot targets, improving the detection accuracy of the original model and expanding the application scenario of the original model.

[0070] Table 1 Experimental results of each model on the base classes of the MAR20 dataset

[0071]

[0072]

[0073] Table 2 Experimental results of each model on the new classes of the MAR20 dataset

[0074]

Claims

1. The Meta-DETR small sample rotation target detection method based on the directed candidate box generation mechanism is characterized by: It includes the following 7 steps: Step 1. Feature extraction; The feature extraction stage is divided into two parts: support set feature extraction and query set feature extraction. The task of support set feature extraction is to extract the features of the target area from the support set image, and the task of query set feature extraction is to extract the full image features from the query set image to provide input for the target detection task. The feature extractors of the support set and query set are shared to ensure that the features of the two are represented in the same semantic space. Step 2. Position encoding; Position encoding enables the model to distinguish features at different positions in the feature map by providing spatial position information for the features; Step 3. Related aggregation module; The relevant aggregation module effectively incorporates the category information provided by the support set into the feature representation of the query set, so that the model can accurately locate and identify objects corresponding to the support set categories in the query image; specifically, CAM achieves this goal in the following ways: Feature alignment: ensure that the category features of the support set are aligned with the image features of the query set in the same semantic space; Feature enhancement: Use the category information of the support set to enhance the feature expression of specific categories in the query set; Contextual association: The association between specific regions in the query set and categories in the support set is captured through the attention mechanism; Step 4. Encoder encodes the features; The enhanced query feature graph has been combined with the category information of the support set, and the global relationship between different positions in the query feature graph still needs to be captured through Encoder modeling; Step 5. Generate directed candidate boxes; When generating a directional candidate box, first initialize a rotation box represented as (p ix , p iy , p iw , p ih , p iθ ), i represents a pixel in the feature map, and (p ix , p iy )∈[0,1], using the prediction result Δr and the initial rotation box p obtained from the directed candidate box generation mechanism i To obtain the final coordinates of each pixel i; Step 6.Decoder layer; The obtained directional candidate boxes are passed as object queries to the Decoder layer, where they interact with the encoded features output by the Encoder to obtain an output embedding containing target feature information, and then the output embedding is input to the prediction head branch; Step 7. Prediction head; The prediction head branch includes the category prediction head branch and the bounding box prediction branch. On the basis of these two branches, the rotation angle prediction branch is added in parallel. The feature information contained in the embedded vector is used to realize the prediction of the target rotation angle using a three-layer fully connected network and the ReLU activation function is used before the adjacent two layers of full connection for nonlinear mapping of features.

2. According to claim 1, the Meta-DETR small sample rotation target detection method based on the directed candidate box generation mechanism is characterized in that: In step 2, the position encoding is performed by the query set feature F query Providing spatial location information enables the model to distinguish the features at different locations in the feature map. The formula is as follows; In the formula, pos represents the position of the embedding vector in the image feature vector, i represents the index in the embedding vector, and d model Represents the output dimension of the model; if the dimension index of the embedding vector is even, the result of the sine formula is used, and if the index is odd, the result of the cosine formula is used.

3. The Meta-DETR small sample rotation target detection method based on directed candidate box generation mechanism according to claim 1 is characterized in that: In step 3, feature matching is done through a single-head attention mechanism. Given the support set category feature F class ∈R C×d and query set F query ∈R H×W×d , the matching coefficient is obtained by the following formula: Where W is a linear projection shared by Q and S, which ensures that they are embedded in the same feature space; Then, the output of feature matching is obtained as follows: Q F =Aσ(F class )⊙F query where σ() represents the sigmoid function and ⊙ represents the Hadamard product, σ(F class ) as a feature filter for each support class, aiming to extract only class-relevant features from the query features; by applying the matching coefficient A to σ(S), we filter out any features that do not match the support class, producing a feature map Q that highlights objects belonging to a given support class F ; To achieve meta-learning that requires class-agnostic predictions, a set of predefined task encodings are introduced, and the given support classes are mapped to these task encodings to make final predictions on the task encodings rather than specific classes. The task encoding T∈R is implemented using a sinusoidal function C×d , the same as position encoding, encoding matching uses the same matching coefficient as feature matching, and the matching encoding Q E Obtained through: Q E =AT Step 4. Encoder encodes the features; The feature graph enhanced by CAM has been combined with the category information of the support set, and still needs to be further modeled by the Encoder to capture the global relationship between different positions in the query feature graph; the Encoder performs global modeling and integrates context information on the input enhanced query feature graph, and at the same time, combined with the needs of small sample learning, enhances the modeling ability of the relationship between the support set and the query set. Through the self-attention mechanism, the Encoder can capture the similarities and differences between the support set and the query set, thereby identifying the targets in the query set that belong to the support set category.

4. The Meta-DETR small sample rotation target detection method based on directed candidate box generation mechanism according to claim 1, characterized in that: In step 5, a directed candidate box generation mechanism is introduced to generate more accurate directed candidate boxes by learning the angle, center, and size of each directed candidate box. These are provided to the decoder as object queries to provide a better position prior for the decoder. When generating directed candidate boxes, a rotation box is initialized as (p ix , p iy , p iw , p ih , p iθ ), i represents a pixel in the feature map, and (p ix , p iy )∈[0,1], there are two branches in the directed candidate box generation mechanism: the distance branch and the angle branch. The distance branch predicts the distance d=(l, t, r, b) from the initial rotation box to the four boundaries of the target box, where 1: the distance to the left boundary, t: the distance to the upper boundary, r: the distance to the right boundary, b: the distance to the lower boundary, and the loss function uses a logarithmic scale regression loss; the angle branch predicts the rotation angle θ of the candidate box and uses the mean square error loss for training. Combining the predicted values ​​of the distance branch and the angle branch, a rough directed candidate box B=(x, yw, h, θ); after the rough candidate box is generated, the aligned convolution is introduced for feature alignment; in the aligned convolution, the calculation of the offset field is the core step to achieve feature alignment; the goal of the aligned convolution is to adjust the sampling points of the feature map through the offset field so that it is aligned with the geometric shape of the candidate box including the center point, width, height and rotation angle; the offset field calculation of the aligned convolution is based on the geometric parameters B = (x, yw, h, θ) of the candidate box; for a standard 3×3 convolution kernel, its sampling point set R is regularly distributed, AlignConv constrains the sampling points within the candidate box, and the position of the sampling point r box , determined by the geometric parameters B of the candidate box; the sampling point position r in the candidate box box The calculation formula is: where s i is the step size of the feature map, p is the position index on the feature map, R is the sampling point set of the standard convolution kernel, and the offset field O represents the offset of the sampling point relative to the feature map grid. The calculation formula is: where r box is the sampling point position in the candidate box, p is the position index on the feature map, and r is the sampling point of the standard convolution kernel; through aligned convolution, the sampling points of the feature map are aligned to the boundary of the candidate box to eliminate the geometric mismatch between the feature and the candidate box; after the feature map is aligned by aligned convolution, the coarse candidate box is optimized through a lightweight full convolutional network to generate a high-quality directional candidate box.

Citation Information

Patent Citations

  • Transform-based traffic scene small sample target detection method and device

    CN116052108A

  • DDETR small sample target detection method based on transfer learning fine tuning

    CN118135353A

  • Small sample target detection method based on correlation region candidate network and converter coding and decoding structure

    CN118781417A

  • Small sample target detection method based on variational coding

    CN119027647A

  • Object detection method and apparatus, device, and storage medium

    WO2024183181A1