Target detection method based on cascade query optimization
By adopting cascade query optimization and joint loss function methods in the object detection algorithm, the problems of insufficient accuracy and cascade errors in complex scenarios and small object detection in the prior art are solved, and higher detection accuracy and performance are achieved.
Patent Information
- Application Number
- CN202510190349.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-02-20
AI Technical Summary
Existing object detection algorithms have problems of insufficient accuracy and cascade errors when dealing with complex scenarios and small targets, making it difficult to take into account both accuracy and speed.
A target detection method based on cascade query optimization is proposed. By selective aggregation of intermediate queries, the negative impact of cascade errors is reduced, and a joint loss function is designed to combine classification loss and regression loss to enhance the correlation between category scores and positioning accuracy.
It significantly improves the overall detection accuracy of the model, reduces the impact of cascade errors, and improves the performance of object detection, especially in complex scenarios.
Smart Images

Figure CN120164015A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision object detection algorithms, and particularly relates to an object detection method based on cascade query optimization. Background Art
[0002] With the rapid development of deep learning, object detection, as a core task in computer vision, has been widely applied in fields such as autonomous driving, security monitoring, and medical imaging, and has achieved remarkable breakthroughs. Object detection algorithms are mainly divided into two categories: two-stage methods and one-stage methods. Two-stage methods (such as Faster R-CNN) generate candidate regions for fine classification and regression. Although they have high accuracy, they have high computational overhead and limited efficiency, especially in complex backgrounds and real-time scenarios where it is difficult to balance accuracy and speed.
[0003] YOLO (You Only Look Once), as a one-stage method, improves the detection speed by directly predicting the object location and category, but it is still limited by the anchor design and its accuracy is inferior in complex scenarios. FCOS proposes an anchor-free architecture, simplifies the detection process, and improves accuracy and reduces computational complexity.
[0004] Although one-stage and two-stage methods have improved in terms of accuracy and efficiency, convolutional neural network CNN still has limitations in extracting global context information, especially when dealing with long-distance context and small objects. To this end, Feature Pyramid Network (FPN) and Deformable Convolution improve the global perception ability through multi-scale information fusion and dynamic convolution kernel adjustment, but the problem has not been completely solved.
[0005] To further improve the global perception ability, the Facebook AI team proposed DETR (DEtectionTRansformer). DETR models the global context through the multi-head self-attention mechanism and can capture the global information in the image, promoting the research progress in the field of object detection. However, DETR still faces the following challenges in practical applications: First, its multi-stage decoding process is prone to cascading errors (as shown in Figure 1 ), and at the same time, the contradiction between class scores and localization accuracy weakens the stability of the detection results; second, DETR has insufficient detection accuracy in some complex scenarios, further limiting its performance in these scenarios. Summary of the Invention
[0006] The object of the present invention is to overcome the defects existing in the above-mentioned prior art, and a target detection method based on cascaded query optimization is proposed. By selecting and aggregating intermediate queries, this method effectively reduces the negative impact of cascading errors, thereby significantly improving the overall detection accuracy of the model. In addition, the present invention also proposes an innovative joint loss function, which alleviates the contradiction between class scores and localization accuracy by combining classification loss and regression loss, further enhancing the performance of target detection.
[0007] To achieve the above object of the invention, the following technical solutions are adopted in the present invention:
[0008] A target detection method based on cascaded query optimization, comprising the following steps:
[0009] S1. Obtain a target detection data set, extract a training set and a test set, and preprocess the images of the training set and the test set to obtain their respective corresponding normalized images;
[0010] S2. Input the normalized image corresponding to the training set into the target detection model optimized by cascaded query for training to obtain a trained target detection model;
[0011] The target detection model optimized by cascaded query includes a spatial feature extraction module, a global feature extraction module, and a cascaded query optimization module;
[0012] The training process of step S2 includes the following steps:
[0013] S21. Input the normalized image into the spatial feature extraction module for deep feature extraction to generate a spatial feature with dimensions (B, C, H, W), and embed position encoding into the spatial feature; where B, C, H, and W are the number, height, width, and number of channels of the spatial feature respectively;
[0014] S22. Input the spatial feature with position encoding into the global feature extraction module, flatten it and then input it into the Transformer encoder to generate a global feature with context awareness;
[0015] The process of inputting the global feature into the cascaded query optimization module includes: First, construct a query set, which is dynamically updated and generated by the query vector q in the multi-head self-attention mechanism during the decoding process; Subsequently, calculate the scores of each query vector in the query set through the query scoring module to form a corresponding query score set S; Then, screen the top K indices with the highest scores from the query score set S, filter according to the indices in the query set, splice the selected queries into a tensor, and input it together with the global feature into the decoder layer to generate a query result; Then, decouple the query result generated by the decoder layer and feedback it back to the query set; After the end of the last decoding stage, process all the queries in the query set through the classification head and the regression head to generate a prediction result;
[0016] S24. Design the joint loss function of the object detection model and optimize it by the gradient descent method; During the training process, continuously adjust the model parameters to minimize the loss function to obtain a trained object detection model;
[0017] S3. Input the standardized image corresponding to the test set into the trained object detection model to obtain the object classification and localization coordinates in the image to be predicted, draw the object box in the image according to the localization coordinates, and label the object category.
[0018] Compared with the prior art, the present invention has the following beneficial effects:
[0019] (1) The object detection method based on cascaded query optimization of the present invention effectively reduces the negative impact brought by cascaded errors by selecting and aggregating intermediate queries, thereby significantly improving the accuracy of model prediction;
[0020] (2) The joint loss function adopted by the object detection model of the present invention enhances the internal correlation between the class score and the localization accuracy by combining the class loss and the regression loss, and further improves the accuracy of object detection. Description of the Drawings
[0021] Figure 1 is a schematic diagram of cascaded errors in the prior art;
[0022] Figure 2 is a flowchart of the object detection method according to the embodiment of the present invention;
[0023] Figure 3 is a diagram of the global feature extraction module according to the embodiment of the present invention;
[0024] Figure 4 is a diagram of the cascaded query optimization module according to the embodiment of the present invention;
[0025] Figure 5 is a diagram of the query scoring module according to the embodiment of the present invention;
[0026] Figure 6 This is a comparison chart of the average precision changes during the training process of the object detection model of the embodiment of the present invention and the DAB-DETR model;
[0027] Figure 7 This is a comparison chart of the detection effects of the object detection model (right side) of the embodiment of the present invention and the DAB-DETR model (left side) on the test set images. Detailed implementation manners
[0028] In order to more clearly illustrate the embodiments of the present invention, the specific implementation manners of the present invention will be described below with reference to the accompanying drawings. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings, and other implementation manners can also be obtained.
[0029] As Figure 2 shown, the object detection method based on cascade query optimization of the embodiment of the present invention includes implementation steps T1 to T6:
[0030] Step T1, obtaining an experimental training set and a test set, and the specific steps are as follows:
[0031] T1.1, downloading a publicly available object detection dataset from the Internet, and extracting training set and test set images and corresponding annotation files (including image basic attributes, object categories, real box information, etc.);
[0032] T1.2, randomly horizontally flipping the images of the training set and the test set, then randomly cropping them into sub-images of different sizes and aspect ratios, and then scaling the cropped images to one of a set of multiple sizes;
[0033] T1.3, normalizing the scaled images according to given means and standard deviations to obtain standardized images.
[0034] The object detection model with cascade query optimization of the embodiment of the present invention includes a spatial feature extraction module, a global feature extraction module, and a cascade query optimization module.
[0035] Step T2, using the spatial feature extraction module to extract the spatial features of the image and embedding the position encoding into the spatial features to provide spatial position information. The specific steps are as follows:
[0036] T2.1, selecting ResNet50 as the spatial feature extraction module, and generating high-dimensional and multi-level spatial features by performing convolution calculations, pooling operations, and activation functions layer by layer to capture details in the image;
[0037] T2.2. Embed the spatial features through position encoding; specifically, add the position encoding to the spatial features pixel by pixel to enhance the model's perception ability of the target spatial position information;
[0038] Step T3. Construct a global feature extraction module. As shown in Figure 3 , input the flattened spatial features into the Transformer encoder to generate global features. The specific steps are as follows:
[0039] T3.1. First, flatten the spatial features into a sequence representation, and then input this sequence into the Transformer encoder. Capture the correlations between sequence elements through the multi-head self-attention mechanism, and combine the position encoding to enhance the global context perception ability of the features. Then, through the processing of the feed-forward neural network, the feature points are mapped to a higher-level representation, further improving the expression ability of the features, thus helping the model to more effectively learn complex feature patterns;
[0040] T3.2. By stacking multiple Transformer encoder layers, gradually strengthen the global context information, and finally generate global features with global context perception, providing richer representations for subsequent tasks.
[0041] In step T4, construct a cascaded query optimization module. As shown in Figure 4 , the specific steps are as follows:
[0042] T4.1. First, construct a query set, and calculate the scores of each query vector in the set through a query scoring module, thereby obtaining a query score set. The structure of the query scoring module is as shown in Figure 5 .
[0043] T4.1.1. Calculate all query scores S in the query set. The formula is as follows:
[0044] S j =Score(q j )
[0045] S={S 0 ,S 1 ,S 2 ,…}
[0046] where q j represents the j-th query vector, S j represents the score of the j-th query, and Score(.) is the scoring head module. Specifically, the core of this module is a linear layer used to perform weighted mapping on the feature representations of all queries. For each query vector q j , process it sequentially through the linear layer to generate a set of mapped feature vectors y j, the formula is as follows:
[0047] y j = W j q j + b j
[0048] where W j is the weight matrix and b j is the bias matrix; then, the mapped feature vector y j is processed through the Sigmoid function to limit all values in the feature vector between 0 and 1, and finally the processed feature vector z j is obtained, and the formula is as follows:
[0049]
[0050] Subsequently, a summation operation is performed on these eigenvalues one by one to obtain the comprehensive score S j of this query, and the formula is as follows:
[0051]
[0052] where M is the number of eigenvalues, is the i-th eigenvalue in the feature vector z j . The comprehensive score effectively reflects the overall quality of the query and its correlation with potential targets. Through the query scoring module, unified modeling and scoring of the query are achieved, providing a clear quantitative indicator for subsequent screening tasks.
[0053] T4.2. First, select the K indices with the highest scores from the query score set, then filter the queries from the query set according to the indices and concatenate them into a tensor, and then input this tensor together with the global features into the Transformer decoder to obtain the query result; finally, decouple the query result and feedback it to the query set for use in the next stage. The specific steps are as follows:
[0054] T4.2.1. First, select the K indices with the highest scores from the score set S, and the formula is as follows:
[0055] topk = Topk(S, K)
[0056] where Topk refers to selecting the K indices with the highest scores from the score set S, and topk are the K selected indices with the highest scores;
[0057] T4.2.2. Filter K queries from the query set according to topk, and then concatenate the K filtered queries to generate a tensor, and the formula is as follows:
[0058] Qt = gather(q t , topk)
[0059] query t = concat(Q t )
[0060] Among them, is the query set in the decoding stage of the t-th layer. The gather operation refers to retrieving the corresponding queries from the query set q t according to the topk indices. Concat is the concatenation operation. Then, both the key and value are repeated K times. Among them, the key is the global feature after embedding position encoding, and the value is the global feature. After the three are processed, they are passed into the multi-head attention mechanism for calculation. The formula is as follows:
[0061] key t = repeat(key, K)
[0062] value t = repeat(value, K)
[0063] Next, the processed query t , key t and value t need to be passed into the decoder for attention calculation. The formula is as follows:
[0064] output = multihead_attention(query t , key t , value t )
[0065] Among them, multihead_attention() is the multi-head attention calculation, and output is the output of the attention mechanism.
[0066] T4.2.3. The output results need to be separated and added to the query set for use in the next stage; after the last decoding stage ends, all the queries in the query set are processed through the classification head and regression head to generate the prediction results.
[0067] In step T5, train the object detection model optimized based on cascaded queries. The specific steps are as follows:
[0068] T5.1. First, according to the prediction results of the model and the real targets, construct a cost matrix and calculate the matching cost (such as IoU) between the predicted bounding boxes and the real bounding boxes; then, apply the Hungarian algorithm to optimize the cost matrix to find the optimal one-to-one matching between each predicted bounding box and the real bounding box
[0069]
[0070]
[0071] Among them, σ(i) is the index of the predicted bounding box that matches the i-th ground truth bounding box, is the pairwise matching loss that combines the classification loss and the regression loss ;
[0072] T5.2. Set the combined loss as the total loss of the model. The calculation process of the combined loss is as follows: First, define t as the weighted arithmetic mean of the confidence score s and the IoU score u, and the formula is as follows:
[0073] t = γ·s + (1 - γ)·u
[0074] where γ is a hyperparameter used to control the ratio of the confidence score and the IoU score; if γ = 0, then t = u, and the loss target will depend entirely on the IoU score; if γ = 1, then t = s, and the loss target will depend entirely on the confidence score; then, sort a set of predicted t i values from largest to smallest to obtain the sorted rank r i , and then calculate the weight w i of the positive samples according to r i , and the formula is as follows:
[0075] w i = exp(-r i / τ)
[0076] where τ is a hyperparameter, and the role of w i is to assign a weight to each positive sample to reflect its importance in learning; subsequently, use t i and w i to improve the classification loss, and assign different weights to positive and negative samples, and the formula is as follows:
[0077]
[0078] where N pos is the number of positive samples, N neg is the number of negative samples, BCE() is the binary cross-entropy loss; t i is weighted down by the w i factor, and it produces a weaker target for negative samples. For negative samples, use the loss weight to focus on those background samples that the model misclassifies as positive samples. For consistency, the regression loss is also weighted by w iThe weight is reduced, and the formula is as follows:
[0079]
[0080]
[0081] Among them, is the IoU loss function applied to the predicted bounding box and the ground truth bounding box b i . The L1 loss is , and λ iou , λ L1 are hyperparameters used to balance the loss. The total loss of the model is:
[0082]
[0083] Among them, N gt is the number of ground truth objects, and λ cls is a hyperparameter.
[0084] T5.3. Use the training set to train the object detection model optimized based on cascaded queries; configure the training parameters with the total batch size set to N;
[0085] T5.4. First, input the images in the training set into the spatial feature extraction module to obtain the spatial features required for the model prediction results and their corresponding ground truth object sets. Next, input the extracted spatial features into the global feature extraction module to obtain global features. Then, these global features are fed into the cascaded query optimization module to generate prediction results. Finally, calculate the loss value between the prediction results and the corresponding ground truth labels through the joint loss function;
[0086] T5.5. Use the Adam optimizer to calculate the gradients of the loss function and update the model parameters using these gradients to minimize the loss value and accelerate the convergence of the model. By setting an appropriate learning rate lr, the optimizer can control the step size of each parameter update, thereby accelerating convergence in the early stage of training and preventing unstable training caused by overly large update steps in the later stage of training. Finally, the optimizer improves the prediction accuracy of the object detection model by continuously updating the model parameters;
[0087] T5.6. Determine whether the current model has converged. If so, obtain the trained object detection model and execute T6; otherwise, continue training and return to T5.4.
[0088] Step T6. Use the trained object detection model to perform object detection on the input pictures.
[0089] First, it is input into the spatial feature extraction module to obtain the spatial features required for the model prediction results and their corresponding true target sets. Next, the extracted spatial features are input into the global feature extraction module to obtain global features. Then, these global features are fed into the cascaded query optimization module to generate prediction results, and the target boxes are drawn in the image and the target categories are labeled through the prediction results.
[0090] The detection effect of the above object detection method according to the embodiments of the present invention can be further verified by the following experiments:
[0091] I. Experimental conditions
[0092] The computing hardware uses an Intel Xeon Silver 4210R processor, and the GPU configuration is 4 NVIDIA GeForce RTX 3090 graphics cards; the operating system is Ubuntu 18.04. The deep learning framework selects PyTorch 2.0.1, and the computing acceleration framework uses CUDA 11.8 and cuDNN 8.6.0. In addition, the present invention takes DAB-DETR as the benchmark model for improvement and optimization.
[0093] II. Experimental content
[0094] Experiment 1: The object detection model in the present invention and the DAB-DETR model are respectively trained using the COCO 2017 training dataset, and their average precision during the training process is respectively recorded. The average precision during the training process is as Figure 6 shown. As the number of iterations increases, the AP of the object detection model (red line) in the present invention gradually increases and stabilizes at 44.6%, while DAB-DETR (green line) can only reach 42.2%, which is lower than the accuracy rate of the present invention.
[0095] From Figure 6 the curve comparison in, as the number of training iterations increases, the object detection model in the present invention effectively reduces the negative impact of cascaded errors through cascaded query optimization, and alleviates the contradiction between the class score and the localization accuracy through the joint loss function. Therefore, the AP gradually improves.
[0096] Table 1 shows the comparison results of various performance indicators of the two models. AP 50 represents the average precision calculated when IoU is 0.50. AP 75 represents the average precision calculated when IoU is 0.75. AP S , AP M and AP L respectively evaluate the detection performance for small, medium, and large objects.
[0097] Table 1 Comparison results of various performance indicators of the two models
[0098] model number of training rounds AP <![CDATA[AP 50 > <![CDATA[AP 75 > <![CDATA[AP S > <![CDATA[AP M > <![CDATA[AP L > DAB-DETR 50 42.2 63.1 44.7 21.5 45.7 60.3 the model of the present invention 50 44.6(+2.4) 64.5 47.8 25.1 48.7 62.1
[0099] As can be seen from Table 1, the object detection model of the present invention is superior to DAB-DETR in different metrics, indicating that its performance in object detection tasks has been comprehensively improved.
[0100] Experiment 2: Randomly select 3 images from the above test dataset and input them into the object detection model and DAB-DETR model of the present invention for object detection. The detection results are as Figure 7 shown.
[0101] From Figure 7 it can be seen that the object detection model of the present invention can identify more objects and has a higher confidence level compared with the DAB-DETR model.
[0102] The object detection method based on cascade query optimization of the present invention reduces the negative impact of cascade errors and improves the accuracy of model prediction by selecting and aggregating intermediate queries. In addition, the present invention proposes a joint loss function, which enhances the correlation between class scores and localization accuracy by combining class loss and regression loss, further improving the detection accuracy.
[0103] The above is only a detailed description of the preferred embodiments and principles of the present invention. For those of ordinary skill in the art, according to the idea provided by the present invention, there will be changes in the specific implementation manners, and these changes should also be regarded as the protection scope of the present invention.
Claims
1. A target detection method based on cascade query optimization, characterized in that: The following steps are involved: S1. Obtain the target detection data set and extract the training set and test set, preprocess the images of the training set and test set to obtain their corresponding standardized images; S2, inputting the standardized image corresponding to the training set into the cascade query optimized target detection model for training, to obtain a trained target detection model; The target detection model of the cascade query optimization includes a spatial feature extraction module, a global feature extraction module and a cascade query optimization module; The training process of step S2 comprises the following steps: S21, the standardized image is input into the spatial feature extraction module to perform deep feature extraction to generate a spatial feature with dimensions (B, C, H, W) and embed the position code into the spatial feature; wherein B, C, H, W are the number, height, width, and number of channels of the spatial feature, respectively; S22, inputting the spatial features with position encoding into the global feature extraction module, flattening them and inputting them into the Transformer encoder to generate global features with context-awareness; S23, the process of inputting the global features into the cascade query optimization module for processing includes: first, constructing a query set, which is generated by dynamically updating the query vector q in the multi-head self-attention mechanism during the decoding process; then, calculating the score of each query vector in the query set through the query scoring module to form a corresponding query score set S; then, filtering the K indexes with the highest scores from the query score set S, filtering the query set according to the index, splicing the filtered queries into a tensor, and inputting the query into the decoder layer together with the global features to generate query results; then, decoupling the query results generated by the decoder layer and feeding them back to the query set; after the last decoding stage, processing all queries in the query set through the classification head and the regression head to generate prediction results; S24. Design a joint loss function for the target detection model and optimize it by gradient descent method; during the training process, continuously adjust the model parameters to minimize the loss function to obtain a trained target detection model; S3. Input the standardized image corresponding to the test set into the trained target detection model to obtain the target classification and positioning coordinates in the image to be predicted. Draw the target box in the image through the positioning coordinates and mark the target category.
2. The target detection method according to claim 1, characterized in that: The step S1 specifically includes the following steps: S11, downloading a public target detection dataset from the Internet, and extracting training set and test set images and corresponding annotation files, where the annotation files include basic image attributes, target categories, and true frame information; S12, randomly flipping the images of the training set and the test set horizontally, and then randomly cropping them into sub-images of different sizes and aspect ratios, and then scaling the cropped images to one of the set multiple sizes; S13, normalizing the scaled image according to a given mean and standard deviation to obtain a standardized image.
3. The target detection method according to claim 1, characterized in that: The spatial feature extraction module is ResNet50.
4. The target detection method according to claim 1, characterized in that: The position code is embedded in the spatial feature by adding the position code and the spatial feature pixel by pixel.
5. The target detection method according to claim 1, characterized in that: The step S22 specifically includes: The spatial features are flattened into a sequence representation, and then the sequence is input into the Transformer encoder. The correlation between sequence elements is captured through a multi-head self-attention mechanism, and the global context perception ability of the features is enhanced by combining position encoding. Then, after being processed by a feedforward neural network, the feature points are mapped to a higher-level representation. By stacking multiple Transformer encoder layers, the global context information is gradually strengthened, and finally a global feature with global context awareness is generated.
6. The target detection method according to claim 1, characterized in that: In step S23, the score set S is queried, and the formula is as follows: S j =Score(q j ) S={S 0 ,S 1 ,S 2 ,…} Among them, q j represents the jth query vector, S j represents the score of the j-th query vector. Score(.) is the scoring head module, which is a linear layer used to perform weighted mapping on the feature representations of all query vectors. For each query vector q j , processed by linear layers in turn, generating a set of mapped feature vectors y j , the formula is as follows: y j =W j q j +b j Among them, W j is the weight matrix, b j is the bias matrix; then, the mapped eigenvector y j The Sigmoid function is used to process all the values in the feature vector to be between 0 and 1, and finally the processed feature vector z is obtained. j , the formula is as follows: Then, the eigenvalues are summed one by one to obtain the comprehensive score S of the query. j , the formula is as follows: Where M is the number of eigenvalues, is the eigenvector z j The i-th eigenvalue in .
7. The target detection method according to claim 6, characterized in that: In step S23, the K indexes with the highest scores are selected from the query score set S, and the formula is as follows: topk=Topk(S,K) Among them, Topk refers to selecting the K indexes with the highest scores from the score set S, and topk is the K indexes with the highest scores selected; According to topk, K queries are filtered out from the query set, and then the filtered K queries are concatenated to generate a tensor. The formula is as follows: Q=gather(q,topk) query = concat(Q) Where q = {q 1 ,q 2 ,...} is the query set, the gather operation refers to taking out the corresponding query from the query set q according to the topk index, and concat is the concatenation operation; then both key and value are repeated K times, where key is the global feature after embedding position encoding, and value is the global feature; after the three are processed, they are passed to the multi-head attention mechanism for calculation, and the formula is as follows: key=repeat(key,K) value=repeat(value,K) Next, the processed query, key, and value are passed to the decoder for attention calculation. The formula is as follows: output=multihead_attention(query,key,value) Among them, multihead_attention() is the multi-head attention calculation, and output is the output of the attention mechanism.
8. The target detection method according to claim 7, characterized in that: In step S24, a cost matrix is constructed based on the prediction results of the target detection model and the real target, and the matching cost IoU between the predicted box and the real box is calculated; then, the Hungarian algorithm is applied to optimize the cost matrix to find the optimal one-to-one match between each predicted box and the real box. Among them, σ(i) is the predicted box index that matches the i-th ground-truth box, To combine classification loss and regression loss Pairwise matching loss of ; The joint loss is set as the total loss of the model. The joint loss calculation process is as follows: First, define t as the weighted arithmetic mean of the confidence score s and the IoU score u, as follows: t=γ·s+(1-γ)·u Among them, γ is a hyperparameter used to control the ratio of confidence score and IoU. If γ = 0, then t = u, and the loss target will depend entirely on IoU. If γ = 1, then t = s, and the loss target will depend entirely on the confidence score. Then, according to a set of predicted t i Sort the values from large to small to get the ranking r i , then according to r i Calculate the weight w of the positive sample i , the formula is as follows: w i =exp(-r i / τ) Among them, τ is a hyperparameter, w i The role of is to assign a weight to each positive sample to reflect its importance in learning; then, use t i and w i The classification loss is improved by assigning different weights to positive samples and negative samples. The formula is as follows: Among them, N pos is the number of positive samples, N neg is the number of negative samples, BCE() is the binary cross entropy loss, is the loss weight; in, For predicting bounding boxes and the ground-truth bounding box b i The IoU loss function is: is the L1 loss, λ iou , L1 is a hyperparameter used to balance the loss; the total loss of the model is: Among them, N gt is the number of real objects, λ cls is a hyperparameter.
9. The target detection method according to claim 8, characterized in that: In step S24, during the training process, the target detection model based on cascade query optimization is trained using the training set, and the training parameters are configured such that the total batch size is set to N; First, the images in the training set are input into the spatial feature extraction module to obtain the spatial features required for the model prediction results and their corresponding real target sets; next, the extracted spatial features are input into the global feature extraction module to obtain global features; Then, the global features are input into the cascade query optimization module to generate prediction results; finally, the loss value between the prediction results and the corresponding true labels is calculated through the joint loss function; The Adam optimizer is used to calculate the gradient of the loss function and use the gradient to update the model parameters to minimize the loss value and accelerate the convergence of the model; Determine whether the current model has converged; if so, obtain the trained target detection model; if not, continue training.
Citation Information
Patent Citations
Small target detection method and system based on improved Deformable-DAB-DETR
CN118097335A
Three-dimensional target detection method and device, electronic equipment and storage medium
CN118298418A
Remote sensing target detection method and system based on dynamic adaptive query
CN119152191A
System and method of open-world semi-supervised satellite object detection
US20240395015A1