A target detection method based on OTS-DETR lightweight model
By introducing the foreground selection network and selection loss function in the OTS-DETR lightweight model, the problem of high computational complexity of the object detection method and inconsistent detection results is solved, and more efficient and accurate object detection is achieved.
Patent Information
- Application Number
- CN202410157799.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-04
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-02-04
AI Technical Summary
The existing object detection methods have high computational complexity and are difficult to obtain the best object detection results. This is mainly because the encoder fails to effectively distinguish the foreground tokens from background tokens, resulting in the inconsistency of redundant calculations and classifications and regression.
A target detection method based on OTS-DETR lightweight model is proposed. The background tokens and foreground tokens are divided by the foreground selection network, and the selection loss function and the IoU-aware selection loss function are introduced to optimize the model to reduce the computational complexity and improve the detection accuracy.
It significantly reduces the computational complexity of target detection, improves prospect quality and target prediction accuracy, achieves higher classification and interchange scores, and obtains the best target detection results.
Smart Images

Figure CN118262087B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target detection based on lightweight models, and in particular to a target detection method based on an OTS-DETR lightweight model. Background Art
[0002] Object detection is one of the important research directions in the field of computer vision. Its main goal is to predict the bounding box and category of an object. In the past decade, the classic convolution-based object detection method has made significant progress. However, the accuracy of the traditional convolution-based object detection method is not high, so how to improve the accuracy of object detection has become a research focus in this field.
[0003] Carion et al. first proposed an end-to-end object detector based on Transformer, named DETR (DEtection TRansformer). The introduction of DETR has successfully applied the idea of transformer to the field of target detection. Due to its simple structure, DETR and similar models have become one of the most popular models in the field of target detection. The uniqueness of DETR is that it eliminates the manually designed anchor points and NMS components in classic target detection, and directly finds the one-to-one optimal result through a bipartite graph matching strategy, thereby solving the bottleneck problem that NMS may bring.
[0004] The research focus of the existing technology is to design object detection methods that can accelerate the convergence of DETR models and improve the accuracy. For example, Deformable-DETR introduces multi-scale information and proposes a deformable attention mechanism, which significantly improves the efficiency of the attention mechanism. Conditional DETR and Anchor DETR optimize the query and enhance the accuracy of the model. DAB-DETR introduces anchor boxes and optimizes the prediction boxes layer by layer. DN-DETR introduces certain positive and negative samples to skip the uncertainty in bipartite graph matching, thereby improving the accuracy of object recognition. DINO integrates a variety of common effective methods and successfully surpasses the accuracy of most classic models. Co-DETR introduces a hybrid matching strategy and achieves state-of-the-art results. PnP-DETR abstracts the entire feature into a fine foreground object feature vector and a small number of coarse background context feature vectors. IMFA searches for key points based on decoder layer predictions to sample multi-scale features and aggregates the sampled features with single-scale features. Sparse DETR preserves the 2D spatial structure of the label by querying sparsity, making it suitable for Deformable DETR to utilize multi-scale features. Lite DETR proposes an efficient encoder that updates high-level and low-level features in an interleaved manner. This method reduces feature labeling and enables efficient detection. Although the target detection method based on the DETR series of models has achieved remarkable results, such excellent performance is usually accompanied by huge computational costs. The main reason is that the encoder in the current target detection method fails to effectively distinguish between foreground tokens and background tokens during the calculation process, resulting in unnecessary redundant calculations in the target detection process, resulting in high computational complexity. At the same time, there is also the problem of inconsistency between classification and regression in the target detection process, which makes it difficult to obtain the best target detection results. Summary of the invention
[0005] The purpose of the present invention is to solve the problems of high computational complexity and difficulty in obtaining the best target detection results in existing target detection methods, and propose a target detection method based on the OTS-DETR lightweight model.
[0006] The specific process of a target detection method based on the OTS-DETR lightweight model is as follows:
[0007] Obtain an image of the target to be detected, input the image of the target to be detected into the trained OTS-DETR lightweight model, and obtain the target detection result;
[0008] The trained OTS-DETR lightweight model is obtained in the following way:
[0009] Step 1: Use the original image and the annotated image corresponding to the original image to form a training set;
[0010] Step 2: Use the training set to train the OTS-DETR lightweight model to obtain a trained OTS-DETR lightweight model;
[0011] The OTS-DETR lightweight model includes: a backbone network, a foreground selection network, an encoder, a decoder, and a prediction head;
[0012] The backbone network uses a pre-trained backbone network to obtain an image feature set S and convert the image feature S in the image feature set S into l Divided into high-level features and low-level features, all image features are input into the foreground selection network;
[0013] The backbone network uses a pre-trained backbone network to obtain an image feature set S and convert the image feature S in the image feature set S into l It is divided into high-level features and low-level features, specifically:
[0014] First, the original image is input into the pre-trained backbone network to obtain the features S4, S3, and S2 output by the C2, C3, and C4 layers;
[0015] Then, S2 is downsampled to obtain feature S1;
[0016] Finally, S1, S2, and S3 are used as high-level features, and S4 is used as a low-level feature;
[0017] in, l∈[1,L], l is the multi-scale feature map layer label, L is the total number of multi-scale feature map layers, C is S l The number of channels, H l YesS l Height, W l YesS l Width;
[0018] The foreground selection network divides background tokens and foreground tokens using image features, and inputs the divided tokens into an encoder;
[0019] The encoder includes: a high-level feature encoder and a low-level feature encoder; the encoder includes: a high-level feature encoder and a low-level feature encoder; the high-level feature encoder and the low-level feature encoder are both composed of a cross attention module Cross Attention; the encoder is used to encode the high-level features and low-level features for dividing background tokens and foreground tokens, and input the encoded high-level features and low-level features into the decoder;
[0020] The decoder is used to decode the encoded high-level features and low-level features, and input the decoded features into the prediction head;
[0021] The prediction head predicts the detection target using the decoded features.
[0022] Furthermore, the encoder includes: a cross attention module Cross Attention.
[0023] Furthermore, the foreground selection network uses image features to divide background tokens and foreground tokens, and inputs the divided tokens into the encoder, specifically:
[0024] Step 1: Get the set S of foreground scores of each layer of image features F :
[0025] First, S1 is input into the multi-layer perceptron MLP F Get the prospect score of S1 in
[0026] Then, using the prospect score of S1 Get Prospect score
[0027]
[0028] Among them, α l’ is a learnable modulation coefficient, UP(.) indicates upsampling by bilinear interpolation, l′∈[1,L-1], l′ is the label of other image features except the first image feature;
[0029] Finally, according to and Get S l The set S of foreground scores F ;
[0030] Step 2: Use S F Get the selection score set S S ;
[0031] Step 3. Sort the features in S from large to small according to the selection scores obtained in Step 2. According to the preset percentage, take the first K features as foreground tokens, and the other tokens as background tokens. Take the high-level features in the saved foreground tokens as the query Q, and input all tokens as the key K and value V into the high-level feature encoder; take the low-level features in the saved foreground tokens as the query Q, and input all tokens as the key K and value V into the low-level feature encoder.
[0032] Furthermore, the use of S in Step 2 F Get the selection score set S S , specifically:
[0033] S s =S F ×MLP C (S)
[0034] Among them, S F yes The set of S is S l The collection of S s is the selection score set.
[0035] Furthermore, the decoder includes: a feedforward neural network module FFN, a cross attention module CrossAttention, and a self-attention module Self attention.
[0036] Furthermore, the foreground selection network is supervised and trained by selecting a loss function, and the selection loss function is obtained by:
[0037] A11. Use the annotated image corresponding to the original image to assign labels to tokens. Foreground tokens are marked as 1 and background tokens are marked as 0. Specifically:
[0038]
[0039] in, is the label of the token at position (i, j), is the maximum chessboard distance between a point (x, y) in the original image and the center of the bounding box, (x, y) is Corresponding to the position of the original image, YesS l The token with coordinates (i, j) in D Bbox (x, y, w, h) is the actual target box, w is the width of the actual target box, h is the height of the actual target box, is the starting feature of the object predicted by the l-th layer feature, is the terminal feature of the object predicted by the l-th layer feature, and
[0040] A12. Use the label of the token assigned by A11 to obtain the selection loss L s .
[0041] Furthermore, the maximum chessboard distance between a point (x, y) in the original image and the center of the bounding box is:
[0042]
[0043] Furthermore, the label of the token assigned by A11 is used in A12 to obtain the selection loss L s , specifically:
[0044] L s =-(1-P t )γ·log(P t )
[0045]
[0046] Among them, P t is the probability value of the predicted result being the true category, and γ is an adjustable parameter.
[0047] Furthermore, the optimization objective of the OTS-DETR lightweight model is obtained as follows:
[0048] B1. Get the binary cross loss weight w′:
[0049] w′=α*(S C ) 2 *(1-t′)+IoU
[0050] Where t′ is the target in the original image, α is the weight adjustment parameter, α∈(0,1), IoU is the intersection over union score between the predicted box and the target box, S C is the classification score;
[0051] B2. Use the binary cross loss weight w′ obtained in B1 to obtain the optimization target of the OTS-DETR lightweight model.
[0052] Furthermore, the binary cross loss weight w′ obtained by B1 in B2 obtains the optimization target of the OTS-DETR lightweight model as follows:
[0053]
[0054]
[0055] in, is the final prediction result, is the predicted category, is the predicted bounding box, y is the actual result, c represents the actual category, b is the actual bounding box, and Binary CrossEntropy() is the binary cross entropy.
[0056] The beneficial effects of the present invention are:
[0057] The object detection method based on the OTS-DETR lightweight model (Optimal Tokens Selection for Efficient DETR) proposed in the present invention obtains a selection score by using a foreground selection network based on sharing and fusion, thereby dividing tokens into foreground tokens and background tokens. By retaining the foreground and discarding the background, the amount of calculation in the encoding process can be significantly reduced, thereby reducing the computational complexity of object detection. The present invention also proposes a selection loss function to supervise the foreground selection, so that the encoder can select better foreground tokens, thereby improving the foreground quality and making the target prediction accuracy higher. The present invention also introduces an IoU-aware selection loss function in the detection head, and uses the IoU coefficient to achieve classification and regression alignment, so that the OTS-DETR lightweight model has a higher classification and intersection-over-union score, thereby improving the accuracy of complex scene object detection, thereby obtaining the best object detection result. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 The relationship between the average accuracy and GFLOPs of different detection models on the COCO dataset;
[0059] Figure 2 GFLOPs occupied by encoder and decoder in different DETR models;
[0060] Figure 3 The framework diagram of OTS-DETR proposed in the present invention;
[0061] Figure 4 Select a frame image for the foreground;
[0062] Figure 5 Select the structural diagram for the top-down perspective. DETAILED DESCRIPTION
[0063] The DETR model mainly consists of three parts: backbone network, encoder and decoder. In DETR, the backbone network inherits the pre-trained network. This design not only simplifies the encoding process, but also allows the use of the backbone network of the pre-trained model, thereby shortening the time required to train the model and improving the training efficiency. Benefiting from the researchers' various optimizations of the matching method and the introduction of anchor points and denoising strategies in decoding, the DETR-based detector has achieved high accuracy. However, the DETR model still faces the problem of high computational cost, which limits its wide application in practical applications and leads to the inability to give full play to its high accuracy advantage. The present invention finds that there is a large amount of background information in the encoder, which far exceeds the foreground information, and most of the calculations of this background information are useless. Therefore, reducing background information can be regarded as an effective strategy, which is expected to reduce the computational cost of GFLOPs. The present invention proposes an OTS-DETR lightweight model to achieve target detection. Figure 2 The GFLOPs occupied by encoding and decoding of different models are shown. Figure 2 It can be seen that the computational cost of the encoder in Deformable DETR is 8.8 times that of the decoder, and the computational cost of the encoder in DINO is 7.0 times that of the decoder. Lightweight in the present invention means that the target detection model proposed in the present invention has a small amount of computation and achieves high accuracy. Next, the present invention is explained in conjunction with specific embodiments.
[0064] Specific implementation method 1: This implementation method is a target detection method based on the OTS-DETR lightweight model. The specific process is as follows:
[0065] Obtain an image of the target to be detected, input the image of the target to be detected into the trained OTS-DETR lightweight model, and obtain the target detection result;
[0066] The trained OTS-DETR lightweight model is obtained in the following way:
[0067] Step 1: Use the original image and the annotated image corresponding to the original image to form a training set;
[0068] The training set is the COCO2017 dataset;
[0069] Step 2: Use the training set to train the OTS-DETR lightweight model to obtain a trained OTS-DETR lightweight model;
[0070] like Figure 3 As shown, the OTS-DETR lightweight model includes: a backbone network, a foreground selection network, an encoder, a decoder, and a prediction head;
[0071] The backbone network uses a pre-trained backbone network to obtain an image feature set S and convert the image feature S in the image feature set S into l Divided into high-level features and low-level features, all image features are input into the foreground selection network;
[0072] Among them, l∈[1,L], l is the multi-scale feature map layer label, L=4 is the total number of multi-scale feature map layers, C is for S l The number of channels, H l YesS l Height, W l YesS l The width of the network is ResNet-50.
[0073] The foreground selection network uses image features to divide background tokens (features) and foreground tokens, and inputs the divided tokens into the encoder;
[0074] The encoder includes a high-level feature encoder (composed of a cross attention module Cross Attention) and a low-level feature encoder (composed of a cross attention module Cross Attention);
[0075] The encoder updates high-level features (corresponding to small-resolution feature maps) and low-level features (corresponding to large-resolution feature maps) in an interlaced manner, and inputs the updated high-level features and low-level features into the decoder;
[0076] The decoder is used to decode the updated high-level features and low-level features, and input the decoded features into the prediction head;
[0077] The decoder includes: a feedforward neural network module FFN, a cross attention module Cross Attention, and a self attention module Self attention;
[0078] The prediction head predicts the detection target using the decoded features.
[0079] In this embodiment, the high-level encoder is executed five times and the low-level encoder is executed once.
[0080] Specific implementation method 2: The backbone network uses a pre-trained backbone network to obtain an image feature set S, and the image feature S in the image feature set S l are high-level features and low-level features, specifically:
[0081] First, the original image is input into the pre-trained backbone network to obtain the features S4, S3, and S2 output by the C2, C3, and C4 layers;
[0082] Then, S2 is downsampled to obtain feature S1;
[0083] Finally, S1, S2, and S3 are used as high-level features, and S4 is used as a low-level feature.
[0084] Specific implementation method three: Figure 4 As shown, the foreground selection network uses image features to divide background tokens and foreground tokens, and inputs the divided tokens into the encoder, specifically:
[0085] Step 1. Get the foreground score of each layer of image features:
[0086] First, S1 is input into the multi-layer perceptron MLP F Get the prospect score of S1 in
[0087] Then, using the prospect score of S1 Get Prospect score
[0088]
[0089] Among them, α l’ is a learnable modulation coefficient, UP(.) indicates upsampling by bilinear interpolation, l′∈[1,L-1], l′ is the label of other image features except the first image feature;
[0090] Finally, according to and Get S l Prospect score The collection S F ;
[0091] Step 2: The accuracy can be significantly improved by simply multiplying the foreground score and the category score. The collection S F Get the selection score set S S :
[0092] S s =S F ×MLP C (S)S s =S F ×MLP C (S)
[0093] Among them, S F yes The set of S is S l The collection of Ss is the selection score set;
[0094] Step 3. Sort the features in S from large to small according to the selection scores obtained in Step 2. According to the preset percentage, take the first K features as foreground tokens, take the other tokens as background tokens, take the high-level features in the saved foreground tokens as query Q, and input all tokens as key K and value V into the high-level feature encoder; take the low-level features in the saved foreground tokens as query Q, and input all tokens as key K and value V into the low-level feature encoder.
[0095] In this embodiment, in the multi-scale feature map, high-level features have richer semantic information, while low-level features have higher resolution. There are associations and differences between different scales, which complement each other and have different functions. Although the feature map of the existing Sparse DETR is multi-scale, when performing foreground selection, Sparse DETR does not fully utilize the multi-scale information and the associations and differences between different scales. The present invention introduces a top-down modulation mechanism based on the shared and fused foreground selection network, which modulates low-level features by high-level features. Specifically, the foreground score multilayer perceptron (MLP) is first used. F ) layer obtains the foreground score of S1 Since the feature S of the l+1 layer l′+1 With higher resolution, the present invention has a l’ Prospect score Upsampling, we get to a higher resolution. Then, use For S l’+1 Modulate and get the modulated S l′+1 Finally, using the modulated S l′+1 Through the foreground scoring multilayer perceptron (MLP F ) layer, obtain S l’+1 Prospect score This mechanism helps to more fully integrate multi-scale information and improve the accuracy of foreground selection. The structure of the top-down foreground selector is as follows: Figure 5 shown.
[0096] In order to select foreground tokens with richer information, the present invention proposes a ranking method of selection scores to determine whether they are foreground tokens. In order to make the selection score carry richer information, the new method combines the classification score, so that the classification score and foreground score of the detector can be fused into the selection score. The classification scorer shares parameters with the class_embed in the prediction head in DETR, which not only reduces the number of parameters of the model, but also achieves a strong association between the encoder and the prediction head, thereby improving the accuracy of target detection. The category scorer shares parameters with the class_embed (embedded class) in the prediction head in DETR, which can prevent the class_embed in the foreground selection network based on sharing and fusion and the class_embed in the detection head from training different models, resulting in reduced accuracy of the trained model.
[0097] Specific implementation method 4: The foreground selection network is supervised and trained by selecting a loss function, and the selection loss function is obtained by the following method:
[0098] A11. Use the annotated image corresponding to the original image to assign labels to tokens. Foreground tokens are marked as 1 and background tokens are marked as 0. Specifically:
[0099] The annotated image corresponding to the original image includes: ground truth and label;
[0100]
[0101]
[0102] in, is the label of the token at position (i, j), is the maximum chessboard distance between a point (x, y) in the original image and the center of the bounding box, (x, y) is Corresponding to the position of the original image, YesS l The token with coordinates (i, j) in D Bbox (x, y, w, h) is the actual target box, w is the width of the actual target box, h is the height of the actual target box, is the starting feature of the object predicted by the l-th layer feature, is the terminal feature of the object predicted by the l-th layer feature, Head
[0103] Each The position (x, y) corresponding to the original image can be expressed as:
[0104] A12. Use the label of the token assigned by A11 to obtain the selection loss L s :
[0105] L s =-(1-P t ) γ ·log(P t )
[0106]
[0107] Among them, γ is an adjustable parameter used to adjust the focus loss, P t is the probability value of the predicted result being the true category.
[0108] The selection loss function proposed in the present invention can supervise the information-rich foreground tokens selected by the foreground selector on the one hand, and can achieve more comprehensive supervision of the detection head on the other hand. The present invention uses groundtruth (real information) and label supervision to select scores. For the selection score of each token, the present invention provides a binary label, where the foreground is marked as 1 and the background is marked as 0. This method helps to guide the foreground selection process more accurately and improve the performance and stability of the model.
[0109] Specific implementation method 5: The optimization target of the OTS-DETR lightweight model is obtained by the following method:
[0110] B1. Get the binary cross loss weight w′:
[0111] w′=α*(S C ) 2 *(1-t′)+IoU
[0112] Where t′ is the target in the original image, α is the weight adjustment parameter, α∈(0,1), IoU is the intersection over union score between the predicted box and the target box, S C is the classification score;
[0113] B2. Use the binary cross loss weight w′ obtained in B1 to obtain the optimization target of the OTS-DETR lightweight model:
[0114]
[0115]
[0116] in, is the final prediction result, is the predicted category, is the predicted bounding box, y is the actual result, c represents the actual category, b is the actual bounding box, and Binary CrossEntropy() is the binary cross entropy; the detection target output by the prediction head is processed using the bipartite graph matching method to obtain
[0117] The present invention truly achieves the consistency between the category score and the IoU score, so that the IoU score and the category score have high scores or low scores at the same time. In solving the problem of aligning high confidence scores with high IoU scores, CNN-based object detectors have proposed a variety of solutions. For example, the confidence score is adjusted by introducing a weight coefficient about IoU, or the IoU prediction branch is directly integrated into the classification loss. These methods solve the alignment problem to a certain extent, but they are all NMS-based methods and cannot be directly applied to the DETR model. Align DETR establishes a strong correlation between the classification score and the positioning accuracy, and solves the alignment problem by introducing the relevant IoU score weight. The present invention introduces an alignment strategy for IoU scores and classification scores, thereby performing IoU-aware selection. IoU-aware selection cleverly integrates labels, IoU scores and classification scores. This paper directly uses IoU scores to replace the labels in the traditional ground truth, that is, the labels are no longer 0 and 1, but the corresponding IoU scores are used according to the correct category, and zero is used for incorrect categories. The present invention calculates the binary cross entropy loss between the new label and the category score as the classification loss. The present invention introduces the classification score as the weight modulation loss to make full use of the classification score.
[0118] Embodiment: In order to verify the beneficial effects of the present invention, the following experiments were carried out:
[0119] All experiments in this example were completed on eight NVIDIA RTX 3090 GPUs. The MS-COCO 2017 detection dataset was used to perform all experiments, and the main indicators of the model, including mean average precision (mAP) and billion floating-point operations (GFLOPs), were verified.
[0120] This embodiment uses Lite DETR as the baseline method, and Lite DETR updates high-level and low-level features in an interlaced manner. This embodiment chooses to update the high-level features five times in a row, and then updates the parameters of the low-level features. In the six feature updates, the foreground proportions are: 1.0, 0.8, 0.6, 0.4, 0.3, and 0.4, respectively. The model was trained for 36 epochs, and the learning rate was reduced by 0.1 times at the 30th epoch. This embodiment uses two backbone networks, ResNet-50 and Swin-T, which are pre-trained on the ImageNet-1K dataset. In addition, this embodiment also performs image enhancement, adjusts the short side of the picture to between 380 and 800, while ensuring that the longest side is less than 1333 pixels, and performs random cropping with a probability of 0.5.
[0121] The OTS-DETR lightweight model uses lite-DETR as a baseline and conducts a series of experiments on different CNN backbone networks, including resnet50 and swin-tiny, to verify the effectiveness of the OTS-DETR lightweight model. The OTS-DETR lightweight model is mainly compared with the real-time detection yolo model and the lightweight DETR model, as shown in Table 1.
[0122] Table 1 Comparison with existing object detectors on COCO val 2017
[0123]
[0124]
[0125] According to Table 1, OTS-DETR not only has high accuracy, but also can significantly reduce the computational cost. When using Resnet-50 as the backbone network, the OTS-DETR lightweight model can achieve 50.6AP at 130GFLOPs; when using swin-tiny as the backbone network, the OTS-DETR lightweight model can achieve 53.7AP at 138GFLOPs. This shows that the OTS-DETR lightweight model significantly surpasses the accuracy of other models under basically the same GFLOPs.
[0126] Comparison with the real-time detector YOLO series models. For fairness, the OTS-DETR lightweight model is only compared with YOLO series models with comparable GFLOPs and AP. Including YOLOv5-L, YOLOv5-X, PPYOLOE-X, YOLOv6-L, YOLOv7-X, YOLOv8-L and YOLOv8-X. Compared with YOLOv5-X, PPYOLOE-X, YOLOv6-L and YOLOv7-X, OTS-DETR achieved an improvement of 3 / 1.4 / 0.9 / 0.8AP respectively, and the computational cost was reduced by 33% / 33% / 8% / 27%, while the number of parameters was reduced by 46% / 52% / 17% / 34%. Compared with YOLOv8-L, OTS-DETR improved by 0.8AP, reduced the computational cost by 16%, and the number of parameters was basically the same. Although it is slightly lower than YOLOv8-X by 0.2AP, the computational cost is significantly reduced by 47%, and the number of parameters is reduced by 31%. Figure 1 As shown in the figure, OTS-DETR achieved the highest accuracy under the same GFLOPs, achieving a balance between computation and accuracy. In the absence of additional training data, the relationship between the average accuracy (Y-axis) and GFLOPs (X-axis) of different detection models on the COCO dataset. All models except the yolo series use ResNet-50 or Swin-Tiny as the backbone network.
[0127] Comparison with end-to-end detectors,Table 1 shows that the OTS-DETR lightweight model achieves state-of-the-art performance among all end-to-end detectors with the same backbone network. Compared with the baseline lite-DETR, OTS-DETR improves 0.4 / 0.4AP and reduces the computational cost by 7% / 8% under different backbone networks of R50 and swin-T. Compared with the DINO model, although the OTS-DETR lightweight model slightly reduces the AP, the computational cost is reduced by 45% / 43% under different backbone networks of R50 and swin-T. Compared with H-DETR, OTS-DETR improves 0.6 / 0.4AP and reduces the computational cost by 42% / 41% under different backbone networks of R50 and swin-T. Compared with Focus, it improves 0.2AP and reduces the computational cost by 14% under the R50 backbone network. Compared with IMFA-DETR, the R50 backbone network significantly improves 5.1AP, the computational cost increases slightly, but the number of parameters decreases by 11%. Compared with Efficient-DETR, the R50 backbone network improves 5.5AP and reduces the computational cost by 38%. With similar GFLOPs, the OTS-DETR lightweight model using the swin-Tiny backbone network significantly improves 0.6AP over the RT-DETR-R50 model, and OTS-DETR has a faster training convergence speed.
[0128] In order to verify the effectiveness of each part of OTS-DETR, this paper conducts ablation experiments on the baseline lite-DETR, selects resnet R50 as the backbone network, and introduces the foreground selection network, selection loss function and IOU-aware selection function. The experimental results are based on the COCO val 2017 dataset. As shown in Table 2, all models are built based on lite-DETR and trained for 36 epochs.
[0129] Table 2 Analysis results of superimposing different numbers of each of the OTS-DETR lightweight model
[0130]
[0131] As can be seen from Table 2, when only the foreground selection network is introduced, the average accuracy only reaches 49.6AP. This is because OTS-DETR only uses foreground tokens, resulting in less information obtained in the encoding. At this time, GFLOPs is reduced by 11GFLOPs, proving the effectiveness of the foreground selection module. After adding the selection loss function, the selection loss can make the selected foreground part more consistent with the prediction head. Although some calculations are increased, higher calculation accuracy is achieved, which proves the effectiveness of the selection loss function. When the OTS-DETR lightweight model adds the IoU-aware selection function, the accuracy increases by 0.5AP, and the calculation amount is not increased, achieving the best effect, proving the effectiveness of the IoU-aware selection function.
[0132] Therefore, it can be seen that the OTS-DETR lightweight model proposed in the present invention achieves the purpose of reducing the amount of calculation by retaining foreground tokens and discarding background tokens. Subsequently, the selection loss proposed in the present invention can select better foreground tokens. The OU-aware selection function introduced in the present invention solves the problem of regression and classification alignment.
Claims
1. A target detection method based on the OTS-DETR lightweight model, characterized in that The specific process of the method is: Obtain an image of the target to be detected, input the image of the target to be detected into the trained OTS-DETR lightweight model, and obtain the target detection result; The trained OTS-DETR lightweight model is obtained in the following way: Step 1: Use the original image and the annotated image corresponding to the original image to form a training set; Step 2: Use the training set to train the OTS-DETR lightweight model to obtain a trained OTS-DETR lightweight model; The OTS-DETR lightweight model includes: a backbone network, a foreground selection network, an encoder, a decoder, and a prediction head; The backbone network uses a pre-trained backbone network to obtain an image feature set S and convert the image feature S in the image feature set S into l Divided into high-level features and low-level features, all image features are input into the foreground selection network; The backbone network uses a pre-trained backbone network to obtain an image feature set S and convert the image feature S in the image feature set S into l It is divided into high-level features and low-level features, specifically: First, the original image is input into the pre-trained backbone network to obtain the features S4, S3, and S2 output by the C2, C3, and C4 layers; Then, S2 is downsampled to obtain feature S1; Finally, S1, S2, and S3 are used as high-level features, and S4 is used as a low-level feature; in, l∈[1,L], l is the multi-scale feature map layer label, L is the total number of multi-scale feature map layers, C is S l The number of channels, H l YesS l Height, W l YesS l Width; The foreground selection network uses image features to divide background tokens and foreground tokens, and inputs the divided tokens into the encoder, specifically: Step 1: Get the set S of foreground scores of each layer of image features F : First, S1 is input into the multi-layer perceptron MLP F Get the prospect score of S1 in Then, using the prospect score of S1 Get Prospect score Among them, α l’ is a learnable modulation coefficient, UP(.) indicates upsampling by bilinear interpolation, l′∈[1,L-1], l′ is the label of other image features except the first image feature; Finally, according to and Get S l The set S of foreground scores F ; Step 2: Use S F Get the selection score set S S : S s =S F ×MLP C (S) Among them, S F yes The set of S is S l The collection of S s is the selection score set; Step 3. Sort the features in S from large to small according to the selection scores obtained in Step 2. According to the preset percentage, take the first K features as foreground tokens, and take the other tokens as background tokens. Take the high-level features in the saved foreground tokens as the query Q, and input all tokens as the key K and value V into the high-level feature encoder; take the low-level features in the saved foreground tokens as the query Q, and input all tokens as the key K and value V into the low-level feature encoder; The encoder includes: a high-level feature encoder and a low-level feature encoder; the high-level feature encoder and the low-level feature encoder are both composed of a cross attention module Cross Attention; the encoder is used to encode the high-level features and low-level features for dividing background tokens and foreground tokens, and input the encoded high-level features and low-level features into a decoder; The decoder is used to decode the encoded high-level features and low-level features, and input the decoded features into the prediction head; The prediction head predicts the detection target using the decoded features.
2. The target detection method based on the OTS-DETR lightweight model according to claim 1, characterized in that: The decoder includes: a feedforward neural network module FFN, a cross attention module Cross Attention, and a self attention module Self attention.
3. A target detection method based on the OTS-DETR lightweight model according to claim 2, characterized in that: The foreground selection network is supervised by the selection loss function, and the selection loss function is obtained by the following method: A11. Use the annotated image corresponding to the original image to assign labels to tokens. Foreground tokens are marked as 1 and background tokens are marked as 0. Specifically: in, is the label of the token at position (i,j), is the maximum chessboard distance between a point (x,y) in the original image and the center of the bounding box, (x,y) is Corresponding to the position of the original image, YesS l The token with coordinates (i, j) in D Bbox (x, y, w, h) is the actual target box, w is the width of the actual target box, h is the height of the actual target box, is the starting feature of the object predicted by the l-th layer feature, is the terminal feature of the object predicted by the l-th layer feature, and A12. Use the label of the token assigned by A11 to obtain the selection loss L s .
4. The target detection method based on the OTS-DETR lightweight model according to claim 3, characterized in that: The maximum chessboard distance between a point (x,y) in the original image and the center of the bounding box is:
5. The target detection method based on the OTS-DETR lightweight model according to claim 4, characterized in that: The label of the token assigned by A11 in A12 is used to obtain the selection loss L s , specifically: L s =-(1-P t ) γ ·log(P t ) Among them, P t is the probability value of the predicted result being the true category, and γ is an adjustable parameter.
6. The target detection method based on the OTS-DETR lightweight model according to claim 5, characterized in that: The optimization objective of the OTS-DETR lightweight model is obtained as follows: B1. Get the binary cross loss weight w′: w′=α*(S C ) 2 *(1-t′)+IoU Where t' is the target in the original image, α is the weight adjustment parameter, α∈(0,1), IoU is the intersection over union score between the predicted box and the target box, S C is the classification score; B2. Use the binary cross loss weight w′ obtained in B1 to obtain the optimization target of the OTS-DETR lightweight model.
7. The target detection method based on the OTS-DETR lightweight model according to claim 6, characterized in that: The binary cross loss weight w′ obtained by B1 in B2 is used to obtain the optimization target of the OTS-DETR lightweight model, as follows: in, is the final prediction result, is the predicted category, is the predicted bounding box, y is the actual result, c represents the actual category, b is the actual bounding box, and Binary CrossEntropy() is the binary cross entropy.
Citation Information
Patent Citations
Improved Faster R-CNN network floating object detection method based on video features
CN110472628A
End-to-end panoramic image segmentation method based on query vector
CN113706572A