DETR model-based real-time small target detection method, system and device, and medium

By introducing feature pyramid module, feature selector and loss function for difficult samples in DETR model, the problem of feature redundancy and IoU sensitivity in small object detection is solved, and the object detection effect with high precision and fast training is achieved.

CN120088575APending Publication Date: 2025-06-03XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510262140.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The prior art has problems of feature redundancy and IoU sensitivity in small object detection, resulting in low detection accuracy and long model training time.

Method used

The real-time small object detection method based on the DETR model is adopted to realize the adaptive fusion of low-level and high-level features through the feature pyramid module, and the feature selector module is used to select features conducive to classification and positioning as candidate vectors, and a specific loss function for difficult samples is designed to improve the model learning efficiency.

Benefits of technology

It significantly improves the accuracy of small object detection and effectively shortens the model training time. Compared with traditional object detection algorithms, it shows significant performance improvements in small object detection scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088575A_ABST
    Figure CN120088575A_ABST
Patent Text Reader

Abstract

The invention discloses a DETR model-based real-time small target detection method, system and device and a medium, and the method comprises the steps: achieving the self-adaptive fusion of low-level and high-level features through a feature pyramid module, and remarkably improving the capability of an Encoder encoder in the feature extraction of a small target object; meanwhile, a feature selector module is used for selecting features beneficial to classification and positioning as candidate vectors, and then query vectors of a Decoder decoder are generated, so that the quality of the query vectors is effectively improved, and the network training process is accelerated; a specific loss function is designed for difficult samples which are difficult to classify or position and are located at an IoU boundary, so that the learning efficiency and the detection accuracy of the model for the difficult samples are improved; the system, the device and the medium are based on the method, in a small target detection scene, the detection precision is improved, and the model training time is effectively shortened; compared with a traditional target detection algorithm, the method shows remarkable performance improvement in a small target detection scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision, and particularly relates to a real-time small target detection method, system, device and medium based on the DETR model. Background Art

[0002] In recent years, the rapid progress of deep learning technology has significantly promoted the leapfrog development of computer technology in many fields. As a core branch in the field of computer vision, object detection mainly aims to classify and identify various objects in an image and accurately provide the position information of these objects. With the continuous improvement of computer hardware performance and the rapid expansion of the dataset scale, object detection technology plays a crucial role in application fields such as medical image diagnosis, autonomous driving technology, and intelligent monitoring systems.

[0003] In actual object detection scenarios, small targets are often more difficult to identify due to reasons such as scarce feature information and background interference. For the patent application with the application number CN202410179944.3 and the title of an object detection method and system for accurately detecting small objects based on deep learning, although multi-scale feature information is fused through a feature pyramid module, the problem of feature redundancy between different scale features in the multi-scale feature fusion process has not been solved; for the patent application of State Grid Chongqing Electric Power Company Electric Power Research Institute and Zhongke Fangcun Zhiwei (Nanjing) Technology Co., Ltd. (application number: CN202410640235.0, title: An unmanned aerial vehicle inspection small target detection method and system based on an improved YOLOv7 model), although the Focal-WIoU loss function is used to enhance the convergence speed of the regression gradient and the weight ratio of high- and low-quality prediction boxes, the problem that the detection accuracy of IoU-sensitive sample targets is greatly affected by IoU fluctuations has not been solved.

[0004] Therefore, conducting research on the small target objects to be detected and improving the algorithm to enhance the small target detection accuracy is a very important and practically valuable research direction in the current object detection field. Summary of the Invention

[0005] To overcome the defects existing in the above-mentioned prior art, the purpose of the present invention is to provide a real-time small target detection method, system, device and medium based on the DETR model. Through the feature pyramid module, the adaptive fusion of low-level and high-level features is realized, significantly enhancing the ability of the Encoder to extract features of small target objects. At the same time, the feature selector module is used to select features conducive to classification and positioning as candidate vectors, and then generate query vectors for the Decoder, effectively improving the quality of the query vectors and accelerating the network training process. For difficult samples that are difficult to classify or locate and are at the IoU boundary, a specific loss function is designed to improve the learning efficiency and detection accuracy of the model for difficult samples. The real-time small target detection method based on the DETR model of the present invention not only improves the detection accuracy but also effectively shortens the model training time in the small target detection scenario. Compared with the traditional target detection algorithm, the present invention shows a significant performance improvement in the small target detection scenario.

[0006] To achieve the above purpose, the technical solution adopted by the present invention is as follows:

[0007] A real-time small target detection method based on the DETR model, comprising the following steps:

[0008] Step 1: Use the backbone network Backbone to extract the features of the image;

[0009] Scale the image to a fixed size, and then input it into the backbone network Backbone for data feature extraction, which is used to provide feature representations for downstream tasks;

[0010] Step 2: Use the Encoder to capture the dependency relationships and context information of the image features extracted by the backbone network Backbone;

[0011] Input the feature information extracted by the backbone network Backbone in Step 1 into the Encoder. Through the self-attention mechanism and feature pyramid module in the Encoder, capture the dependency relationships between features, and adaptively fuse low-level and high-level features as the output features of the Encoder;

[0012] Step 3: Select features conducive to classification and positioning from the output features of the Encoder in Step 2 as candidate vectors, and use the candidate vectors to generate query vectors for decoding;

[0013] Specifically, the prediction-based feature selection module scores the contribution of all the output features of the Encoder to classification and localization; through scoring, the prediction-based feature selection module filters out the features with high contribution to classification and localization, and uses them as candidate vectors; and by adding the loss term of the feature selector to the loss function, the reliability of the vectors filtered by the feature selection module is ensured.

[0014] Step 4: Use the Decoder to convert the query vector into a prediction vector.

[0015] The Decoder matches the query vector with the features output by the Encoder through the cross-attention mechanism; specifically, the cross-attention mechanism calculates the similarity scores between the query vector and the features output by the Encoder to determine which features are most relevant to the query vector; then, the Encoder weights and sums the features output by the Encoder according to these similarity scores to obtain a comprehensive feature representation; by continuously repeating this process, the Decoder gradually learns the position and category information of the target to be predicted; each layer of the attention mechanism receives the features and query vector from the previous layer and continuously updates the query vector through the cross-attention mechanism; the query vector updated by the Decoder is the final prediction vector used for target prediction.

[0016] Step 5: Classify and localize the target according to the prediction vector in Step 4.

[0017] The prediction vector is mapped into the category and position information of the prediction target through the classification neural network and the localization neural network respectively; specifically, the classification neural network determines the category to which the prediction target belongs according to the features in the prediction vector; while the localization neural network obtains the accurate position information of the prediction target according to the position-related features in the prediction vector.

[0018] Step 6: Compare the category and position information of the prediction target in Step 5 with the actual value of the target, and focus on learning the difficult samples with obvious prediction errors.

[0019] By modifying the loss function, the DETR model focuses on learning difficult samples, including samples that are difficult to classify or localize and boundary samples, so as to make the DETR model make more accurate predictions.

[0020] The Encoder described in Step 2 consists of a self-attention layer and a feature pyramid; the specific steps are as follows:

[0021] Represent the last three layers of the feature map output by the backbone network Backbone as {S 3 ,S 4,S 5}, where the size of each feature layer is half of the size of its previous feature layer; then the last layer feature S of the Backbone model is used 5 After serializing in the spatial dimension, it is used as the input of the self-attention layer to capture the dependencies and context information between the input features; subsequently, the output of the self-attention layer is deserialized into the same size as the last layer feature S 5 and is recorded as S 6 ; finally, {S 3 ,S 4 ,S 5 ,S 6} is used as the input of the feature pyramid module to adaptively fuse low-level boundary features and high-level semantic features.

[0022] The structure of the feature pyramid module is as follows:

[0023] The feature pyramid module consists of multiple Adaptive Feature Aggregation Modules (AFAMs), downsampling modules, and upsampling modules; in the feature pyramid, the upsampling module and the downsampling module respectively increase and decrease the feature spatial resolution through the method of bilinear interpolation to ensure that when different features are input into the Adaptive Feature Aggregation Module (AFAM) and the Adaptive Feature Selection Unit (AFSU), the feature size and dimension are consistent; the Adaptive Feature Aggregation Module (AFAM) fuses low-level boundary features and high-level semantic features to achieve adaptive feature integration; an Adaptive Feature Selection Unit (AFSU) is established to enable the Adaptive Feature Aggregation Module (AFAM) to adaptively fuse features of various scales, and the formula is as shown in (2):

[0024] F' a = W a (F a ), F' b = W b (F b ) (1)

[0025] AFSU(F a , F b ) =

[0026] F' a ⊙ F a + F b ⊙ F' b ⊙ (1 - F' a ) + F a (2)

[0027] where F a and F b represent the input features of AFSU, that is, the features before fusion, Fa and F b are input into two linear mappings and the sigmoid function W a and W b , obtaining the feature coefficients F' a and F' b . Subsequently, through the action of these feature coefficients, the Adaptive Feature Selection Unit (AFSU) can utilize F' a and F' b to eliminate the information repeated with F b from the feature F a , retain the unique content of F b , and adaptively fuse the feature F b into the feature F a ;

[0028] After introducing the Adaptive Feature Selection Unit (AFSU), the formula of the Adaptive Feature Fusion Module (AFAM) is shown in (4):

[0029] I a' = AFSU(I a , I b ), I b' = AFSU(I b , I a ) (3)

[0030] AFAM(I a , I b ) = Conv(Concat(I a' , I b' )) (4)

[0031] Among them, I a' represents the new feature after I a fuses the feature I b , I b' represents the new feature after I b fuses the feature I a ; Conv(·) refers to the 3×3 convolutional layer, which integrates batch normalization and the ReLU activation function; Concat(·) is to concatenate two features along the channel dimension to form a larger feature.

[0032] The specific method of screening out the features conducive to classification and localization as candidate vectors and generating query vectors for decoding described in Step 3 is:

[0033] Input the features output by the Encoder into the prediction-based feature selector; after receiving the input, the feature selector first converts the input features into a probability value. Subsequently, based on these calculated probability values, the feature selector performs sorting and screening. The higher the probability value, the greater the contribution of the features output by the Encoder to the target prediction during the decoding process by the Decoder; finally, select the top K features with the highest probability values; these selected features are regarded as features beneficial to classification and positioning, which can significantly improve the performance and accuracy of the model; the specific formula is shown in (5):

[0034] O 1,2...k =TopK(W φ (I 1,2...n )) (5)

[0035] Q 1,2...k =Embedding(O 1,2...k ) (6)

[0036] Among them, I 1,2...n is the feature output by the Encoder, and n represents the number of output features. W φ (·) is the combination of a linear mapping and the sigmoid function, which can map the features to the range from 0 to 1; O 1,2...k are the k features beneficial to classification and positioning selected by the prediction-based feature selector, and then use the embedding technique to construct a query vector, that is, Q 1,2...k ;

[0037] By adding the loss term of the feature selector to the loss function, to ensure that the features selected by the prediction-based feature selector are more beneficial to classifying and positioning the target; the loss term of the specific feature selector is shown in formula (10):

[0038]

[0039] O 1,2...k are the k features beneficial to classification and positioning selected by the prediction-based feature selector, W β (·) is the combination of a linear mapping and the sigmoid function, and O 1,2...k can obtain a probability value cls β in the range from 0 to 1 through W 1,2...k (·). The higher the probability value, the higher the probability that the object is recognized after O 1,2...k is decoded by the Decoder; and further estimate the offset anchor_offset 1,2...k of the target object relative to the anchor point norm_anchor; among them, cls 1,2...k ranges from 0 to 1, anchor1,2...k Indicates the position of the target object, which is the sum of the anchor point and the offset;

[0040] pred 1,2...k Is the predicted value after decoding the features that are beneficial for classification and localization. Based on the predicted value pred 1,2...k And the ground truth gt 1,2...m , calculate the matching cost C i Between the predicted value pred j And the ground truth gt ij , thereby constructing a cost matrix; adjust the rows and columns of the matrix to determine the best matching scheme of the predicted value and the ground truth in the cost matrix, so that the total matching cost is minimized; thereby determining the matching relationship between the predicted value and the actual value. If the predicted value can find the corresponding actual value match, it means the prediction is successful; otherwise, the prediction fails; finally, the loss of the feature selector consists of the localization loss of anchor 1,2...k , box_gt 1,2...k And the classification loss of cls 1,2...k , cls_gt 1,2...k .

[0041] Focus on learning the difficult samples described in step six; specifically, the difficult samples can be defined into the following two categories:

[0042] The first category is samples that are difficult to classify or localize, referring to those that are difficult to classify or localize the target object during the classification or localization process, including those with accurate localization but incorrect classification, and those with accurate classification but obvious deviation in localization; the second category is boundary samples, referring to those where the localization of the target object is exactly at the threshold edge; whether such objects can be successfully detected depends on the Intersection over Union (IoU) threshold of the final prediction;

[0043] For samples that are difficult to classify or localize, when calculating the classification loss, introduce the localization loss between the predicted value and the ground truth as a weight factor and perform a weighted operation with the classification loss; the specific classification loss is shown in formula (11):

[0044]

[0045] The predicted target position set and the ground truth target position set are respectively represented as box_pred 1,2...k , box_gt 1,2...k ; while the predicted classification set and the ground truth classification set are respectively represented as cls 1,2...k , cls_gt 1,2...k ; iou 1,2...k Is the localization loss between the predicted value and the ground truth target.

[0046] Focus on learning the difficult samples described in step six, specifically:

[0047] For samples with the localization score at the threshold edge, whether the detection is successful depends on the IoU threshold set during the prediction process; when the IoU threshold is increased, samples that could originally be detected may also be missed;

[0048] By adjusting the IoU threshold between the predicted target and the ground truth target and optimizing the weight of the loss function of this sample, the model can focus more on samples that are difficult to detect, thereby improving the accuracy of overall object detection; the specific formula is shown in (14):

[0049] η i = η i-1 ×(1 - e -i / epochs ) (12)

[0050]

[0051] Among them, u represents the average value of the intersection over union of all bounding boxes during a single iteration of training; and u mean is the mean IoU obtained by weighted average calculation of time series data, and this mean value can be dynamically adjusted according to the progress of training iterations; η is the decay coefficient, and i represents the number of iterations;

[0052] During object detection, samples with IoU greater than u mean and less than u mean - 0.1 are regarded as easily recognizable samples, and in other cases, they are defined as boundary samples. The loss function of formula (14) amplifies the loss value of boundary samples so that the model can more effectively focus on the learning of boundary samples during object recognition.

[0053] A real-time small object detection system based on the DETR model, including:

[0054] A feature extraction module, used in step one, to extract data features from the input image through the backbone network Backbone, for providing feature representations for downstream tasks;

[0055] A feature pyramid module, used in step two, first calculates the feature coefficients between different features input to the feature pyramid module through an adaptive feature selection unit, and fuses the features according to the feature coefficients; then splices the fused new features along the channel dimension through an adaptive feature fusion module, so that the spliced features contain semantic information of different levels, realizing the adaptive fusion of low-level boundary features and high-level semantic features;

[0056] A prediction-based feature selection module, which is used in Step 3. Through the feature selection module, the features output by the Encoder are evaluated, and several features beneficial to classification and localization are selected to initialize the query vector. This can make the query vector have practical physical meaning and higher quality, avoid the interference of randomly initialized noise information during the decoding process of the Decoder, and improve the accuracy of object detection while achieving the rapid convergence of the model.

[0057] The decoder module, which is used in Step 4. Through the cross-attention mechanism, the query vector is matched with the features output by the Encoder, and gradually the query vector learns the position and category information of the target to be predicted; it provides a vector for prediction for downstream tasks.

[0058] The target prediction module, which is used in Step 5. Through the classification neural network and the localization neural network, the prediction vector is respectively mapped into the category and position information of the predicted target.

[0059] The hard sample learning module, which is used in Step 6. By introducing the localization score into the classification loss function and dynamically adjusting the loss weight of the localization boundary samples, the DETR model can focus on learning hard samples that are difficult to classify or localize and are near the IoU boundary. This realizes the consistency constraint of the positive sample classification accuracy and the localization accuracy in the object detection process and reduces the sensitivity of the IoU boundary class samples to the localization threshold.

[0060] A real-time small object detection device based on the DETR model, comprising:

[0061] A memory, which is used to store computer programs;

[0062] A processor, which is used to implement the real-time small object detection method based on the DETR model described in Steps 1 to 6 when executing the computer program.

[0063] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, it can implement real-time small object detection based on the DETR model by the method described in Steps 1 to 6. Compared with the prior art, the present invention has the following advantages:

[0064] 1. In the feature pyramid module of the present invention in Step 2, by introducing the Adaptive Feature Aggregation Module (AFAM), the feature coefficients between two different features are successfully learned. By element-wise multiplying the coefficients with the features, this module realizes the function of adaptively fusing shallow and deep features. Compared with the traditional feature pyramid module, this innovation not only effectively supplements the lack of high-level semantic features in spatial boundary information but also makes up for the deficiency of low-level boundary features in semantic information.

[0065] 2. In the third step of the present invention, the feature selector module evaluates the output features of the Encoder encoder and selects a number of high-quality output features of the Encoder encoder for initializing the query vector. To ensure that the selected features have higher quality and stronger representativeness, a loss function for the feature selector module is specifically designed. This method successfully solves the problem of random initialization of query vectors in traditional DETR models.

[0066] 3. In the sixth step of the present invention, for difficult samples that are difficult to classify or locate and are near the IoU boundary, by introducing the localization score into the classification loss function, the consistency constraint between the classification accuracy of positive samples and the localization accuracy in the object detection process is ensured. In addition, a strategy for dynamically adjusting the loss weight of localization boundary samples is proposed to improve the learning ability of the object detection system for the IoU boundary, thereby reducing the sensitivity of IoU boundary class samples to the localization threshold.

[0067] In summary, the real-time small object detection method based on the DETR model of the present invention not only improves the detection accuracy but also effectively shortens the model training time in the small object detection scenario. Compared with traditional object detection algorithms, this method shows significant performance improvement in the small object detection scenario. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] Figure 1 is a schematic diagram of the model of the present invention.

[0069] Figure 2 is a schematic diagram of the adaptive feature fusion module of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0070] The present invention will be further described in detail below with reference to the accompanying drawings.

[0071] See Figure 1 , a small object real-time detection method based on the DETR model, the specific steps are as follows:

[0072] Step 1: Use the Backbone to extract the features of the image, scale the image to a fixed size, and then input it into the backbone network Backbone for data feature extraction to provide rich feature representations for downstream tasks.

[0073] Step 2: Use the Encoder encoder to capture the dependency relationship and context information of the image features extracted by the backbone network Backbone, and input the feature information extracted by the backbone network Backbone in Step 1 into the Encoder encoder. The Encoder encoder consists of a self-attention layer and a feature pyramid; the specific steps are as follows:

[0074] Represent the last three layers of the feature maps output by the backbone network Backbone as {S 3 , S 4 , S 5}, where the size of each feature layer is half of the size of its previous feature layer; then serialize the last layer feature S 5 of the backbone network Backbone model in the spatial dimension as the input of the self-attention layer to capture the dependencies and context information between the input features; subsequently, deserialize the output of the self-attention layer into the same size as the last layer feature S 5 and record it as S 6 ; finally, use {S 3 , S 4 , S 5 , S 6} as the input of the feature pyramid module to adaptively fuse low-level boundary features and high-level semantic features.

[0075] Among them, the feature pyramid module consists of multiple adaptive feature fusion modules (AFAM), downsampling modules, and upsampling modules; in the feature pyramid, the upsampling module and the downsampling module respectively increase and decrease the feature spatial resolution through the method of bilinear interpolation to ensure that when different features are input into the adaptive feature fusion module (AFAM) and the adaptive feature selection unit (AFSU), the feature size and dimension are consistent; the adaptive feature fusion module (AFAM) fuses low-level boundary features and high-level semantic features to achieve adaptive feature integration; an adaptive feature selection unit (AFSU) is established to enable the adaptive feature fusion module (AFAM) to adaptively fuse features of various scales, and the formula is as shown in (2):

[0076] F' a = W a (F a ), F' b = W b (F b ) (1)

[0077] AFSU(F a , F b ) =

[0078] F' a ⊙ F a + F b ⊙ F' b ⊙ (1 - F' a ) + F a (2)

[0079] Among them, F a and F bRepresent the input features of AFSU, i.e., the features before fusion, F a and F b are input into two linear mappings and the sigmoid function W a and W b , obtaining the feature coefficients F' a and F' b . Subsequently, through the action of these feature coefficients, the Adaptive Feature Selection Unit (AFSU) can utilize F' a and F' b to eliminate the information repeated with F b from the feature F a , retain the unique content of F b , and adaptively fuse the feature F b into the feature F a . After introducing the Adaptive Feature Selection Unit (AFSU), the formula of the Adaptive Feature Fusion Module (AFAM) is shown in (4):

[0080] I a' = AFSU(I a , I b ), I b' = AFSU(I b , I a ) (3)

[0081] AFAM(I a , I b ) = Conv(Concat(I a' , I b' )) (4)

[0082] where, I a ' represents the new feature after I a fuses the feature I b , I b ' represents the new feature after I b fuses the feature I a ; Conv(·) refers to the 3×3 convolutional layer, which integrates batch normalization and the ReLU activation function; Concat(·) is to concatenate two features along the channel dimension to form a larger feature;

[0083] Step 3: Select features conducive to classification and localization and generate query vectors. Use the prediction-based feature selection module to select the features conducive to classification and localization from the features output by the Encoder as candidate vectors, and use them to generate query vectors for decoding.

[0084] Specifically, the features output by the Encoder are input into a prediction-based feature selector; after receiving the input, the feature selector first converts the input features into a probability value. Subsequently, based on these calculated probability values, the feature selector performs sorting and screening. The higher the probability value, the greater the contribution of the features output by the Encoder to the target prediction during the decoding process by the Decoder. Finally, the top K features with the highest probability values are selected; these selected features are regarded as features beneficial to classification and localization, which can significantly improve the performance and accuracy of the model. The specific formula is shown in (5):

[0085] O 1,2...k =TopK(W φ (I 1,2...n )) (5)

[0086] Q 1,2...k =Embedding(O 1,2...k ) (6)

[0087] Among them, I 1,2...n is the feature output by the Encoder, and n represents the number of output features. W φ (·) is the combination of a linear mapping and a sigmoid function, which can map the features to the range from 0 to 1. O 1,2...k are the k features beneficial to classification and localization selected by the prediction-based feature selector. Subsequently, a query vector is constructed using the embedding technique, that is, Q 1,2...k .

[0088] To ensure that the selected features have better quality, the loss term of the feature selector is included in the loss function during the training phase. First, use the features I 1,2...n generated by the Encoder to predict whether these features can accurately predict the object and its position after being processed by the Decoder. Then, the loss function of the feature selector is constructed based on the difference between the predicted value and the actual value.

[0089] The loss term of the specific feature selector is shown in formula (10).

[0090]

[0091] O 1,2...k are the k features beneficial to classification and localization selected by the prediction-based feature selector. Based on these features, predict whether the object cls 1,2...k can be recognized after being processed by the Decoder, and further estimate the offset anchor_offset 1,2...k of the target object relative to the anchor point norm_anchor, where cls 1,2...kThe value range is from 0 to 1, anchor 1,2...k represents the position of the target object, which is the sum of the anchor point and the offset.

[0092] pred 1,2...k is the predicted value after decoding the features beneficial to classification and localization. Based on the predicted value pred 1,2...k and the ground truth gt 1,2...m , through the Hungarian matching algorithm, calculate the predicted value pred i and the ground truth gt j the matching cost C between them ij , thus constructing a cost matrix. By adjusting the rows and columns of the matrix, determine the best matching scheme of the predicted value and the ground truth in the cost matrix, so that the total matching cost is minimized. In this way, the matching relationship between the predicted value and the actual value can be determined. If the predicted value can find the corresponding actual value match, it means the prediction is successful; otherwise, the prediction fails. Finally, the loss of the feature selector consists of the localization loss of anchor 1,2...k , box_gt 1,2...k and the classification loss of cls 1,2...k , cls_gt 1,2...k .

[0093] Step 4: Use the Decoder decoder to convert the query vector into a predicted vector;

[0094] The Decoder decoder matches the query vector with the features output by the Encoder through the cross-attention mechanism; specifically, the cross-attention mechanism determines which features are most relevant to the query vector by calculating the similarity scores between the query vector and the features output by the Encoder; then, the Decoder decoder weights and sums the features output by the Encoder according to these similarity scores to obtain a comprehensive feature representation; by continuously repeating this process, the Decoder decoder gradually learns the position and category information of the target to be predicted; each layer of the attention mechanism receives the features and the query vector from the previous layer and continuously updates the query vector through the cross-attention mechanism; the query vector updated by the Decoder decoder is the final predicted vector used for target prediction;

[0095] Step 5: Classify and localize the target according to the predicted vector in Step 4

[0096] The predicted vector is mapped into the category and position information of the predicted target through the classification neural network and the localization neural network respectively; specifically, the classification neural network determines the category to which the predicted target belongs according to the features in the predicted vector; while the localization neural network obtains the accurate position information of the predicted target according to the position-related features in the predicted vector;

[0097] Step 6: Focus on learning difficult samples. The first type is samples that are difficult to classify or locate, referring to those for which it is difficult to classify or locate the target object during the classification or location process, including those with accurate location but incorrect classification, and those with accurate classification but significant deviation in location. The second type is boundary samples, referring to those where the location of the target object is exactly at the threshold edge; whether such objects can be successfully detected depends on the Intersection over Union (IoU) threshold of the final prediction.

[0098] For samples that are difficult to classify or locate, when calculating the classification loss, the location loss between the predicted value and the true value is introduced as a weight factor and weighted with the classification loss; this method aims to ensure the consistency of prediction information when predicting the classification and location of the target, thereby improving the accuracy of target object location and classification; the specific classification loss is shown in formula (11).

[0099]

[0100] The predicted target location set and the true target location set are respectively denoted as box_pred 1,2...k , box_gt 1,2...k . While the predicted classification set and the true classification set are respectively denoted as cls 1,2...k , cls_gt 1,2...k . iou 1,2...k is the location loss between the predicted value and the true target.

[0101] For samples whose location score is exactly at the threshold edge, whether such samples can be successfully detected highly depends on the IoU threshold set during the prediction process. When the IoU threshold is increased, samples that could originally be detected may also encounter missed detection. To address this issue, the weight of the loss function for this sample can be optimized according to the IoU between the predicted target and the true target. This strategy aims to make the model more focused on difficult-to-detect samples, thereby improving the overall accuracy of target detection. The specific formula is shown in (14).

[0102] η i = η i-1 ×(1 - e -i / epochs ) (12)

[0103]

[0104] Among them, u represents the average value of the IoU of all bounding boxes during a single iteration of training. And u meanThe mean IoU is obtained by calculating the weighted average of time series data, and this mean can be dynamically adjusted according to the progress of training iterations. η is the decay coefficient and i represents the iteration number, which indicates that in the early stage of training, the model mainly focuses on current observations; while in the later stage of training, the model is more inclined to rely on historical data.

[0105] During the object detection process, samples with IoU greater than u mean are regarded as easily recognizable samples, while samples with IoU values close to u mean are defined as boundary samples. Given the deficiency of the model in recognizing boundary samples, this loss function specifically amplifies the loss values of boundary samples so that the model can more effectively focus on learning boundary samples.

[0106] The present invention further verifies its effect through the following experiments.

[0107] A real-time small object detection method based on the DETR model, VisDrone2019 is a drone aerial photography dataset, and the resolution of some images is as high as 2000×1500. The object density in the VisDrone2019 dataset is very high, and most objects are very small, with an average pixel less than 32×32 pixels. It is a commonly used public dataset for small object detection algorithms. This dataset includes 10 categories: pedestrians, crowds, bicycles, cars, vans, trucks, tricycles, sunshade tricycles, buses, and motorcycles.

[0108] The operating system used in the experiment is Ubuntu18.04, and the deep learning framework is PyTorch. The specific configurations involved in the experiment are shown in Table 1.

[0109] Table 1 Experimental configuration table

[0110]

[0111]

[0112] In order to verify the effectiveness of the adaptive feature fusion module (AFAM), feature selector and difficult sample learning strategy proposed in this invention, six groups of ablation experiments are designed. All experiments use ResNet-18 as the backbone network (Bonebone). The results of the ablation experiments conducted on the VisDrone2019 dataset are detailed in Table 2. Experimental group 1 represents the baseline model; experimental groups 2 to 4 introduce AFAM, feature selector and difficult sample learning strategy based on the baseline model; experimental group 5 combines AFAM and feature selector; experimental group 6 is the final model using the above three modules. The evaluation index adopts a percentage system, where mAP0.5:0.95 means that the mean average precision (mAP) is calculated at intervals of 0.05 in the range of intersection over union (IoU) from 0.5 to 0.95, and the average value is taken as the final evaluation result.

[0113] Table 2 Ablation experiment results of VisDrone2019 dataset

[0114] Experimental group AFAM Feature selector Hard sample learning mAP0.5 (%) mAP0.5:0.95 (%) 1 36.8 21.0 2 √ 39.0 22.8 3 √ 37.7 21.9 4 √ 38.2 22.2 5 √ √ 39.6 23.3 6 √ √ √ 40.5 24.0

[0115] The experimental results show that the baseline model performs the worst in various accuracy indicators. In experimental groups 2-4, the application of any module alone can significantly improve the accuracy of the model, which verifies the effectiveness of the adaptive feature fusion module, feature selector and difficult sample learning strategy in improving the accuracy of the model. In experimental group 6, the final model that comprehensively applies all modules achieved 40.5% in the mAP0.5 indicator, an increase of 3.7% compared to the baseline model, and performed best among all experimental groups. This shows that after the comprehensive application of each module, the detection performance of the model reaches its peak. The present invention verifies the effectiveness of the adaptive feature fusion module, feature selector and difficult sample learning strategy through experiments, and shows that the improved model has higher accuracy in small target object detection.

[0116] In order to verify the accuracy difference between the method of the present invention and other public target detection algorithms, a public algorithm comparison experiment was set up. Table 3 lists the results of the comparison experiment on the VisDrone2019 dataset. As can be seen from the table, the algorithms involved in the comparison are YOLOv8-m, YOLOv9-m, YOLOv10-m, YOLOv11-m, DETR, RT-DETR and the model proposed in the present invention. The parameter amount (params) and the number of floating-point operations (GFLOPs) are used to measure the parameter scale and computational complexity of each model respectively.

[0117] Table 3. Comparison results of model performance on the VisDrone2019 dataset

[0118] Model mAP0.5 (%) params (M) GFLOPs YOLOv8-m 38.2 25.8 78.7 YOLOv9-m 38.9 20.0 76.5 YOLOv10-m 37.8 16.5 63.5 YOLOv11-m 38.8 20.0 67.5 DETR 34.3 26.9 88.7 RT-DETR 36.8 20.2 58.6 The present invention 40.5 19.2 81.3

[0119] The experimental results show that: under the condition that the number of parameters and the amount of computation are comparable to those of the DETR series models, the YOLO series models show higher accuracy in the small object detection task. However, after applying the module proposed in the present invention to the DETR model, the mAP0.5 index reaches 40.5%, which is 1.6% higher than the highest accuracy model in the YOLO series models. This result confirms that the method proposed in the present invention has higher detection accuracy in the small object detection task compared with other publicly disclosed object detection algorithms, showing its remarkable advancement.

Claims

1. A real-time small target detection method based on DETR model, characterized in that: The following steps are involved: Step 1: Use the backbone network to extract image features; The image is scaled to a fixed size and then input into the backbone network for data feature extraction to provide feature representation for downstream tasks; Step 2: Use the Encoder encoder to capture the dependency and context information of the image features extracted by the backbone network; The feature information extracted by the backbone network in step 1 is input into the encoder. The self-attention mechanism and feature pyramid module in the encoder capture the dependency between features, and adaptively fuse low-level and high-level features as the output features of the encoder. Step 3: Select features that are conducive to classification and positioning from the output features of the encoder in step 2 as candidate vectors, and use the candidate vectors to generate query vectors for decoding; Specifically, the prediction-based feature selection module scores the contribution of all encoder output features to classification and positioning; through scoring, the prediction-based feature selection module selects features with high contribution to classification and positioning and uses them as candidate vectors; and by adding the loss term of the feature selector to the loss function, the reliability of the vectors selected by the feature selection module is ensured; Step 4: Use the Decoder to convert the query vector into a prediction vector; The Decoder matches the query vector with the features output by the Encoder through the cross-attention mechanism. Specifically, the cross-attention mechanism determines which features are most relevant to the query vector by calculating the similarity score between the query vector and the features output by the Encoder. Then, the Encoder performs a weighted summation of the features output by the Encoder based on these similarity scores to obtain a comprehensive feature representation. By repeating this process, the Decoder gradually learns the location and category information of the target to be predicted. Each layer of the attention mechanism receives the features and query vector from the previous layer, and continuously updates the query vector through the cross-attention mechanism. The query vector updated by the Decoder is the final prediction vector used for target prediction. Step 5: Classify and locate the target according to the prediction vector of step 4 The prediction vector is mapped into the category and location information of the prediction target through the classification neural network and the positioning neural network respectively; specifically, the classification neural network determines the category of the prediction target according to the features in the prediction vector; while the positioning neural network obtains the accurate location information of the prediction target according to the location-related features in the prediction vector; Step 6: Compare the category and location information of the predicted target in step 5 with the actual value of the target, and focus on learning the difficult samples with obvious prediction errors; By modifying the loss function, the DETR model is used to focus on learning difficult samples, including samples that are difficult to classify or locate and boundary samples, so that the DETR model can make more accurate predictions.

2. A real-time small target detection method based on DETR model according to claim 1, characterized in that: The Encoder described in step 2 consists of a self-attention layer and a feature pyramid; the specific steps are as follows: The last three layers of the feature map output by the backbone network Backbone are represented as {S3, S4, S5}, where the size of each feature layer is half the size of its previous feature layer; then the last layer feature S5 of the backbone network Backbone model is serialized according to the spatial dimension as the input of the self-attention layer to capture the dependencies and contextual information between the input features; then, the output of the self-attention layer is deserialized into the same size as the last layer feature S5 and recorded as S6; finally, {S3, S4, S5, S6} is used as the input of the feature pyramid module, and the low-level boundary features and high-level semantic features are adaptively fused through the feature pyramid module.

3. The Encoder according to claim 2 is composed of a self-attention layer and a feature pyramid, characterized in that: The feature pyramid module structure is as follows: The feature pyramid module is composed of multiple adaptive feature fusion modules (AFAM), downsampling modules and upsampling modules. In the feature pyramid, the upsampling module and downsampling module respectively increase and decrease the feature space resolution through the method of bidirectional linear interpolation, ensuring that the feature size and dimension remain consistent when different features are input into the adaptive feature fusion module (AFAM) and the adaptive feature selection unit (AFSU). Adaptive Feature Fusion Module (AFAM) fuses low-level boundary features and high-level semantic features to achieve adaptive feature integration; establish The adaptive feature selection unit (AFSU) enables the adaptive feature fusion module (AFAM) to adaptively fuse features of various scales. The formula is shown in (2): F' a =W a (F a ),F' b =W b (F b ) (1) Among them, F a and F b represents the input features of AFSU, i.e. the features before fusion, F a and F b It is input to two linear maps and the sigmoid function W a and W b , and obtain the characteristic coefficient F' a and F' b Then, through the action of these feature coefficients, the adaptive feature selection unit (AFSU) can use F' a and F' b , from the feature F b Elimination and F a Repeated information, keep F b unique content, and adaptively transform the feature F b Fusion to feature F a middle; After introducing the adaptive feature selection unit (AFSU), the adaptive feature fusion module (AFAM) formula is shown in (4): I a' =AFSU(I a ,I b ),I b' =AFSU(I b ,I a ) (3) AFAM(I a ,I b )=Conv(Concat(I a' ,I b' )) (4) Among them, I a' Representative I a Incorporates Feature I b The new features after I b' Representative I b Incorporates Feature I a The new features after that; Conv(·) refers to the 3×3 convolution layer, which integrates batch normalization and ReLU activation function; Concat(·) connects two features along the channel dimension to form a larger feature.

4. A real-time small target detection method based on DETR model according to claim 1, characterized in that: The specific method of selecting the features that are conducive to classification and positioning as candidate vectors and using the candidate vectors to generate query vectors for decoding as described in step 3 is: Input the features output by the Encoder encoder into the prediction-based feature selector; After receiving the input, the feature selector first converts the input feature into a probability value. Then, the feature selector sorts and filters according to these calculated probability values. The higher the probability value, the greater the contribution of the encoder output feature to the target prediction during the decoding process of the decoder. Finally, the top K features with the highest probability value are selected. These selected features are regarded as features that are conducive to classification and positioning, which can significantly improve the performance and accuracy of the model. The specific formula is shown in (5): ABOUT 1,2...k =TopK(W φ (AND 1,2...n )) (5) Q 1,2...k =Embedding(O 1,2...k ) (6) Among them, I 1,2...n W is the feature output by the Encoder, and n represents the number of output features. φ (·) is a combination of linear mapping and sigmoid function, which can map features to the range of 0 to 1; 1,2...k The k features that are useful for classification and positioning are selected by the prediction-based feature selector, and then the query vector is constructed using embedding technology, i.e., O 1,2...k ; By adding the loss term of the feature selector to the loss function, it is ensured that the features selected by the feature selector based on the prediction are more conducive to the classification and positioning of the target; the specific loss term of the feature selector is shown in formula (10): O 1,2...k are k features selected by the prediction-based feature selector that are beneficial for classification and positioning, W β (·) Combination of linear mapping and sigmoid function, O 1,2...k By W β (·) can get a probability value cls in the range of 0 to 1 1,2...k The higher the probability value, the higher the 1,2...k The higher the probability of identifying the object after decoding by the Decoder, the higher the probability of identifying the object; and further estimate the offset anchor_offset of the target object relative to the anchor point norm_anchor 1,2...k ; Among them, cls 1,2...k The value range is 0 to 1, anchor 1,2...k Indicates the position of the target object, which is the sum of the anchor point and the offset; pred 1,2...k It is the predicted value after decoding of the features that are conducive to classification and positioning. 1,2...k and the true value gt 1,2...m , calculate the predicted value pred i and the true value gt j Matching cost C ij , thereby constructing a cost matrix; adjusting the rows and columns of the matrix to determine the best matching scheme between the predicted value and the true value in the cost matrix, so that the total matching cost is minimized; thereby determining the matching relationship between the predicted value and the actual value. If the predicted value can find the corresponding actual value match, it means that the prediction is successful; otherwise, the prediction fails; finally, the loss of the feature selector is determined by the anchor 1,2...k 、box_gt 1,2...k The positioning loss and cls 1,2...k ,cls_gt 1,2...k The classification loss is composed of .

5. A real-time small target detection method based on DETR model according to claim 1, characterized in that: Step 6 focuses on learning the difficult samples; specific difficult samples can be defined into the following two categories: The first category is samples that are difficult to classify or locate, which refers to samples that are difficult to classify or locate during the classification or positioning process, including those with accurate positioning but wrong classification, and samples with accurate classification but obvious positioning deviation; the second category is boundary samples, which refers to the positioning of the target object just at the edge of the threshold; whether such objects can be successfully detected depends on the final predicted Intersection over Union (IoU) threshold; For samples that are difficult to classify or locate, when calculating the classification loss, the positioning loss of the predicted value and the true value is introduced as a weight factor and weighted with the classification loss. The specific classification loss is shown in formula (11): The target position prediction set and the true target position set are represented as box_pred 1,2...k 、box_gt 1,2...k ; The predicted classification set and the true classification set are represented as cls 1,2...k ,cls_gt 1,2...k ;iou 1,2...k is the localization loss between the predicted value and the true target.

6. A real-time small target detection method based on DETR model according to claim 1 or 5, characterized in that: The focus of learning difficult samples described in step 6 is as follows: For samples whose positioning scores are at the edge of the threshold, the success of their detection depends on the IoU threshold set during the prediction process; when the IoU threshold is increased, samples that could have been detected may also be missed; By adjusting the IoU threshold between the predicted target and the true target, the weight of the sample loss function is optimized, so that the model can focus more on samples that are difficult to detect, thereby improving the overall target detection accuracy; the specific formula is shown in (14): or i =the i-1 ×(1-e -i / epochs ) (12) Among them, u represents the average value of the intersection-over-union ratio of all bounding boxes during a single iteration of training; and u mean It is the mean IoU calculated by weighted averaging of time series data, which can be dynamically adjusted according to the progress of training iterations; η is the decay coefficient, and i represents the number of iterations; In the object detection process, IoU is greater than u mean and less than u mean Samples with a value of -0.1 are considered easy-to-recognize samples, and samples in other cases are defined as boundary samples. The loss function in formula (14) amplifies the loss value of boundary samples so that the model can focus on learning boundary samples more effectively during target recognition.

7. A real-time small target detection system based on DETR model, characterized in that: include: The feature extraction module is used in step 1 to extract data features from the input image through the backbone network to provide feature representation for downstream tasks. The feature pyramid module is used in step 2. First, the feature coefficients between different features input to the feature pyramid module are calculated through the adaptive feature selection unit, and the features are fused according to the feature coefficients; then, the fused new features are spliced ​​along the channel dimension through the adaptive feature fusion module, so that the spliced ​​features contain semantic information at different levels, realizing adaptive fusion of low-level boundary features and high-level semantic features; The prediction-based feature selection module is used in step 3. The feature selection module evaluates the output features of the encoder and selects several features that are conducive to classification and positioning to initialize the query vector. This can make the query vector have actual physical meaning and higher quality, avoid the decoder being disturbed by randomly initialized noise information during the decoding process, achieve rapid model convergence, and improve the accuracy of target detection; The decoder module is used in step 4 to match the query vector with the features output by the encoder through the cross-attention mechanism, gradually allowing the query vector to learn the location and category information of the target to be predicted; and provide a vector for prediction for downstream tasks; The target prediction module is used in step 5 to map the prediction vector into the category and location information of the predicted target respectively through the classification neural network and the positioning neural network; The difficult sample learning module is used in step six. By introducing the positioning score into the classification loss function and dynamically adjusting the positioning boundary sample loss weight, the DETR model can focus on learning difficult samples that are difficult to classify or locate and are located near the IoU boundary, thereby achieving the consistency constraint between the positive sample classification accuracy and the positioning accuracy in the target detection process and reducing the sensitivity of the IoU boundary class samples to the positioning threshold.

8. A real-time small target detection device based on the DETR model, characterized in that: include: Memory for storing computer programs; A processor is used to implement the real-time small target detection method based on the DETR model described in steps 1 to 6 when executing the computer program.

9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method described in steps 1 to 6 can realize real-time small target detection based on the DETR model.

Citation Information

Patent Citations

  • Target detection method and system for accurately detecting small object based on deep learning

    CN118053117A

  • Unmanned aerial vehicle inspection small target detection method and system based on improved YOLOv7 model

    CN118397488A