DETR-based double-domain pseudo-label generation cross-domain target detection method

By introducing a DETR-based dual-domain pseudo-label generation method in unsupervised field adaptive object detection, the problem of difficulty in feature alignment between the source domain and the target domain and instability in pseudo-label quality is solved, and the generalization ability and pseudo-label quality of the model on the target domain are improved.

CN120147676AActive Publication Date: 2025-06-13CHINA UNIV OF GEOSCIENCES (WUHAN)

Patent Information

Application Number
CN202510214760.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-13
Estimated Expiration
2045-02-26

AI Technical Summary

Technical Problem

In the unsupervised field adaptive target detection, the difficulty of feature alignment between the source domain and the target domain and the unstable pseudo-label quality leads to insufficient generalization ability of the model on the target domain.

Method used

A cross-domain object detection method based on DETR is proposed. By constructing a teacher-student model framework, multi-scale decoding query clustering module, dual-domain collaboration module and dual-domain verification module, high-quality pseudo-labels are generated, false positives and false negatives are reduced, and cross-domain detection performance is enhanced.

Benefits of technology

The generalization ability of the object detection model on the target domain is improved, the pseudo-label quality is significantly improved, the dependence on precise feature alignment is reduced, and the performance of the model in complex scenarios is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147676A_ABST
    Figure CN120147676A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer vision, and provides a DETR-based double-domain pseudo-tag generation cross-domain target detection method, a multi-scale decoding query clustering module performs multi-scale clustering on queries output by a decoder, and can accurately capture feature information of targets of different scales, so that the confidence coefficient of pseudo tags is effectively evaluated. The method has a strong cross-domain migration capability, can generate a pseudo label with higher quality between the style data of the target domain and the style data of the source domain, and optimizes the confidence of the pseudo label, thereby improving the accuracy and robustness of target detection, and remarkably reducing the problems of missing detection and false detection of pseudo labels with different scales. The double-domain cooperation and double-domain verification module enhances the reliability of pseudo label generation by combining pseudo labels under different confidence levels of style images of a source domain and a target domain, and effectively filters noise. The adaptability of cross-domain data is improved, it is ensured that the model can smoothly migrate among different fields, and the detection precision and generalization ability are further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and particularly relates to a cross-domain object detection method based on DETR for generating dual-domain pseudo-labels. Background Art

[0002] Object detection technology has made remarkable progress with the development of deep learning. However, most object detection models are difficult to effectively generalize to the target domain with different visual features and data distributions after being trained on specific datasets (such as COCO, Pascal VOC, etc.). The visual differences between the source domain and the target domain (such as object appearance, background environment, lighting conditions, etc.) lead to the "inter-domain difference" or "domain shift" problem, thus affecting the performance of the object detection model in the target domain. To solve this problem, domain-adaptive object detection has emerged, and its goal is to transfer the knowledge of the source domain to the target domain through unsupervised learning or weakly supervised learning, and reduce the negative impact of inter-domain differences on the object detection model.

[0003] Domain-adaptive object detection methods are mainly divided into two categories: feature alignment methods and pseudo-label generation methods. Feature alignment methods reduce the feature distribution differences between the source domain and the target domain through means such as adversarial training, so as to extract domain-invariant features and achieve cross-domain detection. However, it is difficult to perfectly achieve feature alignment when there are significant visual differences, and the supervision signal of the source domain overly affects the learning of the target domain, resulting in insufficient generalization ability of the model. The pseudo-label generation method is to first train a detector on the source domain, then generate pseudo-labels for the target domain, and use the pseudo-labels to self-train the model. The teacher-student framework is a common strategy in this method. The teacher model is responsible for generating pseudo-labels, the student model is trained through the pseudo-labels, and the parameters of the teacher model are updated by exponential moving average to maintain the stability of the pseudo-labels. Current pseudo-label generation methods have achieved good results in some scenarios, but due to inter-domain differences, the quality of pseudo-labels still has a large uncertainty, and the existence of false positive and false negative labels seriously affects the training effect of the model.

[0004] Although progress has been made in domain adaptation object detection methods, several challenges still remain. First, feature alignment methods are difficult to achieve perfect alignment when there are large differences between domains, which results in the model's performance on the target domain not reaching the ideal effect. Second, the quality of pseudo-label generation methods depends on the confidence estimation of the teacher model, and the confidence of the teacher model is easily affected by the differences between the source domain and the target domain, leading to inaccurate pseudo-labels. Many existing teacher-student models transfer knowledge based on convolutional neural networks, but convolutional neural network models lack global context information and relationship modeling between instances, which makes the detection results in the target domain, especially the candidate boxes with low confidence, prone to misjudgment. In addition, the pseudo-labels generated by the teacher model often have large uncertainties when the target domain labels are missing, further affecting the training effect of the student model. Therefore, researching a method that can improve the generalization ability of the object detection model and improve the quality of pseudo-labels is crucial for unsupervised domain adaptation object detection tasks. Summary of the Invention

[0005] Aiming at the problems of difficult feature alignment and unstable pseudo-label quality between the source domain and the target domain in the unsupervised domain adaptation object detection task, the present invention proposes a cross-domain object detection method based on DETR for generating dual-domain pseudo-labels. Its purpose is to solve the problems of difficult feature alignment and unstable pseudo-label quality between the source domain and the target domain in unsupervised domain adaptation object detection. Furthermore, it improves the generalization ability of the object detection model on the target domain, reduces false positives and false negatives by generating high-quality pseudo-labels, and then enhances the cross-domain detection performance and reduces the dependence on precise feature alignment.

[0006] To achieve the above purpose, the present invention provides a cross-domain object detection method based on DETR for generating dual-domain pseudo-labels, which is applied to the unsupervised domain adaptation object detection task. The method includes:

[0007] S1: Construct a teacher-student model framework on a DETR-like detector, which is composed of a trained student model and an inferring teacher model;

[0008] S2: Obtain source domain and target domain images, train the CUT model, and infer the corresponding source domain and target domain style images through this CUT model;

[0009] S3: Use a pre-trained model trained on the COCO dataset, and further train the student model and the teacher model with the labeled source domain images and target domain style images to obtain an initialized student model and teacher model;

[0010] S4: Construct a multi-scale decoding query clustering module. The teacher model infers the target-domain and source-domain style images to obtain decoding queries, and calculates the query similarity through this multi-scale decoding query clustering module as the confidence. Each query represents a potential object, containing category information and bounding box prediction information.

[0011] S5: Construct a dual-domain collaboration module to fuse the high-confidence results inferred from the target-domain and source-domain style images as reliable pseudo-labels.

[0012] S6: Construct a dual-domain verification module. From the difficult samples inferred from the target-domain and source-domain style images, extract the highly overlapping samples with the IOU value of the detection boxes of the same category exceeding the threshold as reliable pseudo-labels, and fuse them with the reliable pseudo-labels generated by the dual-domain collaboration module to obtain the final pseudo-labels.

[0013] S7: Use the final pseudo-labels as the labels of the target-domain images for the training of the student model. At the same time, the teacher model is continuously updated, and finally a trained teacher-student model is obtained for cross-domain object detection.

[0014] A storage device that stores instructions and data for implementing the above-mentioned cross-domain object detection method based on DETR with dual-domain pseudo-label generation.

[0015] A cross-domain object detection device based on DETR with dual-domain pseudo-label generation, including: a processor and a storage device; the processor loads and executes the instructions and data in the storage device for implementing the above-mentioned cross-domain object detection method based on DETR with dual-domain pseudo-label generation.

[0016] The beneficial effects of the present invention are as follows: The cross-domain object detection method based on DETR with dual-domain pseudo-label generation proposed by the present invention is used for unsupervised domain adaptation object detection. The introduced multi-scale decoding query clustering module can accurately capture the feature information of targets at different scales by performing multi-scale clustering on the queries output by the decoder, effectively improving the robustness of the confidence evaluation of pseudo-labels in multi-scale object detection tasks, and thus effectively evaluating the confidence of pseudo-labels. This method has strong cross-domain transfer ability, can generate higher-quality pseudo-labels between the target domain and the source-domain style data, optimize the confidence of pseudo-labels, and further improve the accuracy and robustness of object detection, significantly reducing the missed detection and false detection problems of pseudo-labels at different scales, and enhancing the performance of the model in complex scenarios.

[0017] The introduced dual-domain collaboration and dual-domain verification modules enhance the reliability of pseudo-label generation and the robustness of cross-domain learning by combining pseudo-labels under different confidence levels of source-domain and target-domain style images, improve the adaptability of cross-domain data, ensure that the model can be smoothly transferred between different domains, and further improve the detection accuracy and generalization ability. It effectively improves the quality of pseudo-labels, reduces noise and redundancy, and enhances the adaptability and detection performance of the model in the target domain.

[0018] By introducing a multi-scale decoding query clustering module, dual-domain collaboration, and dual-domain verification modules, the process of pseudo-label generation and filtering is optimized, false positives and false negatives are reduced, and the generalization ability of the cross-domain detection model is enhanced. Through this method, the performance of the object detection model in the target domain is improved, the quality of pseudo-labels is significantly enhanced, the dependence on precise feature alignment is reduced, and thus the effect of unsupervised domain adaptation object detection tasks is improved. Brief Description of the Drawings

[0019] Figure 1 It is a framework diagram of a cross-domain object detection method based on DETR with dual-domain pseudo-label generation proposed by the present invention;

[0020] Figure 2 It is a flowchart of the multi-scale decoding query clustering module in an embodiment of the present invention;

[0021] Figure 3 It is a flowchart of the dual-domain collaboration calculation method in an embodiment of the present invention;

[0022] Figure 4 It is a flowchart of the dual-domain verification calculation method in an embodiment of the present invention;

[0023] Figure 5 It is a schematic diagram of the process of category-based non-maximum suppression in an embodiment of the present invention;

[0024] Figure 6 It is a schematic diagram of the operation of the hardware device in an embodiment of the present invention. Detailed Embodiments

[0025] The principles and features of the present invention are described below with reference to the accompanying drawings. The examples given are only for explaining the present invention and are not intended to limit the scope of the present invention.

[0026] Embodiment 1

[0027] This embodiment discloses a cross-domain object detection method based on DETR with dual-domain pseudo-label generation, which is applied to unsupervised domain adaptation object detection tasks. The method includes:

[0028] S1: Build a teacher-student model framework on a DETR-like detector, which consists of a trained student model and an inferring teacher model;

[0029] S2: Obtain source domain and target domain images, train the CUT model, and infer the corresponding source domain and target domain style images through this model;

[0030] S3: Use a pre-trained model trained on the COCO dataset, and further train the student model and the teacher model with the annotated source domain images and target domain style images to obtain the initialized student model and teacher model.

[0031] S4: Build a multi-scale decoding query clustering module. The teacher model infers the target domain and source domain style images to obtain decoding queries, and calculates the query similarity through this module as the confidence; each query represents a potential object, containing class information and bounding box prediction information;

[0032] S5: Build a dual-domain collaboration module, and fuse the high-confidence results inferred from the target domain and source domain style images together as reliable pseudo-labels;

[0033] S6: Build a dual-domain verification module. From the difficult samples inferred from the target domain and source domain style images, extract the highly overlapping samples with the IOU value of the same-class detection boxes exceeding the threshold as reliable pseudo-labels, and fuse them with the reliable pseudo-labels generated by the dual-domain collaboration module to obtain the final pseudo-labels;

[0034] S7: Use the pseudo-labels as the labels of the target domain images for the training of the student model, and at the same time the teacher model is continuously updated, and finally the trained model is obtained.

[0035] The following describes each step in detail.

[0036] Further, in step S1, build a teacher-student model framework on a DETR-like detector, as Figure 1 shown. This framework consists of two DAB Deformable DETR models, both using ResNet-50 as the feature extractor. The specific steps are as follows:

[0037] S101: Build a student model, and update the gradient through backpropagation and loss values. The teacher model is updated by performing an exponential moving average (EMA) on the parameters of the student model. The formula is:

[0038]

[0039] where, θ t and θ srespectively represent the parameter weights of the teacher model and the student model, k is the number of training iterations, and α is the smoothing coefficient of EMA, which is used to control the speed of parameter update.

[0040] S102: This framework is mainly used for the teacher model to generate pseudo-labels for subsequent training of the student model in the target domain for object localization and classification tasks.

[0041] Furthermore, in step S2, source domain and target domain images are obtained, the CUT model is trained, and the model is used to generate corresponding source domain and target domain style images. The specific steps are as follows:

[0042] S201: Obtain image data of the source domain and target domain with different visual styles, and train the CUT model based on these data.

[0043] S201: After training is completed, use the model to perform style conversion inference on the source domain and target domain images, and generate corresponding source domain style images and target domain style images respectively.

[0044] Furthermore, in step S3, use the coco pre-trained model to perform supervised training on the source domain and target domain style data. The specific steps are as follows:

[0045] S301: Download the detection model trained on the COCO dataset, and modify the detection head network according to the requirements of the actual cross-domain task to obtain the pre-trained model.

[0046] S302: Further supervise and train the pre-trained model with the annotated source domain images and target domain style images, and select the student model or teacher model with the best performance during the process.

[0047] Furthermore, in step S4, construct a multi-scale decoding query clustering module, as Figure 2 shown, the teacher model infers the target domain and source domain style images, obtains the decoding query, and calculates the query similarity as the confidence through this module. The specific steps are as follows:

[0048] S401: Given the target image, the encoder of the teacher model first extracts the image features and generates the encoded features, as shown in the formula:

[0049]

[0050] where, is the i-th image in the target domain, is the encoded feature, Encoder is the encoder, and Backbone is the image feature extractor.

[0051] Next, the decoder uses the features output by the encoder and the object Generate a set of decoded object queries, as shown in the formula:

[0052]

[0053] Wherein, represents the decoded object query, N is the number of queries, and d is the dimension of each query. Each query represents a potential object, containing class information and bounding box prediction information. After being processed by the detection head, the inferred target information is obtained, as shown in the formula:

[0054]

[0055] Wherein, is the detection confidence, is the predicted class label, is the refined bounding box;

[0056] S402: Select the detection confidence exceeding the threshold τ r of the query as a reliable decoded object Scale information represents the area ratio of the object, as shown in the formula:

[0057]

[0058] Wherein, is the area of the refined bounding box, is the total area of the input image, represents the i-th target domain image;

[0059] S403: Aggregate similar object features in the feature space to form a consistent and compact set of class-scale prototypes. Extract the class information and scale information from the reliable decoded object in formulas (4) and (5) Then aggregate the features with the same class c and the same scale r, and generate the class-scale prototype by feature averaging The formula is:

[0060]

[0061] Wherein, represents the reliable object output by the decoder, and represent the class and scale information of the corresponding object respectively. The symbol is the indicator function, which is when the condition is satisfied, and 0 otherwise. N t is the total number of target domain images, For each target domain image the total number of objects in it. is the class-scale prototype, where N class is the number of classes, N scale is the number of scale levels, and d is the feature dimension. By clustering the features of the same class and scale, the generated class-scale prototype aims to accurately represent the object features under different class and scale conditions. During the teacher-student mutual learning stage, the class-scale prototype will be continuously updated as the training progresses. The update process is the same as described in formula (6), and the updated prototype is generated by aggregating and averaging the features of the same class c and scale level r so as to achieve continuous optimization;

[0062] S404: Design a new confidence evaluation method by calculating the cosine similarity between the decoded query generated by the decoder and the pre-built class-scale prototype As shown in formula (7), select the prototype most similar to the query as the confidence source of the query, and regard the maximum similarity as the clustering confidence of the object. At the same time, assign the class index related to the class-scale prototype with the highest cosine similarity as the class label of the query, so as to re-evaluate the confidence of the query and possibly change its initial class.

[0063]

[0064] where, is the clustering confidence score calculated by the maximum similarity, is the class label of the j-th object assigned by the maximum cosine similarity.

[0065] Furthermore, in step S5, construct a dual-domain collaboration module, as Figure 3 shown, and fuse the high-confidence inference results of the target domain and source domain style images into reliable pseudo-labels. The steps are as follows:

[0066] S501: Adopt a selection strategy based on clustering confidence scores in the target domain and source domain style images to screen high-confidence pseudo-labels. In the target domain, only the objects with clustering confidence scores greater than the predefined reliability threshold are selected as candidate target pseudo-labels The formula is:

[0067]

[0068] where, is the bounding box of the j-th object in the i-th target domain image, is the class label of the j-th object, is the total number of objects in each target domain image;

[0069] S502: Execute the multi-scale decoding query clustering module in the source domain style image to generate the class-scale prototypes of the source domain style Subsequently, according to the clustering confidence scores in the source domain style image greater than the preset threshold to filter the source domain style images and obtain reliable pseudo-labels Its formula is:

[0070]

[0071] where, is the bounding box of the j-th object in the i-th source domain style image, is the class label of the j-th object in the i-th source domain style image, is each source domain style image the total number of objects in.

[0072] S503: As Figure 3 shown, the target domain image and the corresponding source domain style image only differ in image style, and the classes and positions of the foreground objects remain the same. Therefore, the reliable pseudo-labels generated in the target domain and source domain style images can be combined to generate the final high-confidence target domain pseudo-labels Then apply the class-based non-maximum suppression (NMS) method to these two sets of pseudo-labels to eliminate redundant and overly overlapping predictions, thereby obtaining the final reliable target domain pseudo-labels Its formula is:

[0073]

[0074] where, NMS class represents the class-based non-maximum suppression operation, and ∪ represents the union of the target domain pseudo-labels and the source domain style pseudo-labels the union of.

[0075] Furthermore, in step S6, construct a dual-domain verification module, as Figure 4 shown, mine difficult pseudo-labels and fuse them with the reliable pseudo-labels generated by the dual-domain collaboration module to obtain the final pseudo-labels.

[0076] The specific steps are as follows:

[0077] S601: The dual-domain verification module selects reliable pseudo-labels with low clustering confidence scores by leveraging the consistency between the target image and the corresponding source-domain style image. First, for the target-domain data, n candidate bounding boxes with clustering confidence lower than a preset threshold are selected, and for the corresponding source-domain style image, n candidate bounding boxes with clustering confidence lower than the threshold are selected, denoted as and respectively. Although the clustering confidence of these samples is low, they may still retain valuable information;

[0078] S602: The target-domain image and its corresponding source-domain style image differ in visual style, but the categories and positions of the foreground objects are consistent. Then, the intersection over union (IoU) metric is used to quantify the and overlap between samples of the same category in and the pseudo-labels with higher overlap are selected as reliable hard samples

[0079]

[0080] where, represents the bounding box of the k-th object in the i-th source-domain style image, represents the class label of the k-th object in the i-th source-domain style image, represents the bounding box of the k-th object in the i-th target-domain image. When the intersection over union (IoU) exceeds the overlap threshold δ oss , the sample is retained.

[0081] S603: Since a single foreground object may be captured by multiple decoded object queries with different confidence scores, resulting in the same candidate bounding box being selected in both reliable pseudo-labels and hard pseudo-labels. To reduce confusion, and are subjected to class-based non-maximum suppression (NMS), as shown in Figure 5 , and the formula is:

[0082]

[0083] where, represents the final set of pseudo-labels for the i-th target-domain image after merging the target-domain and source-domain style data, and NMS class represents the class-level non-maximum suppression operation applied to remove redundant labels.

[0084] Further, in step S7, the pseudo-label is used as the label of the target domain image to guide the training of the student model. Meanwhile, the teacher model is continuously updated to obtain the finally trained model. The specific steps are as follows:

[0085] S701: The training of the teacher-student model is divided into two stages. In the initial stage, a dual-source domain supervised training strategy is adopted. The source domain data and the target domain style data are used to gradually guide the model to adapt to the target domain, enhance domain-specific learning, and promote effective transfer. For the supervised training loss, the formula is:

[0086]

[0087] where represents the detection loss, including the source domain detection loss and the target stylized detection loss Each detection loss includes the L1 loss and the Generalized Intersection over Union (GIoU) loss.

[0088] S702: The source domain and the target domain style data share the same foreground objects and labels, so the model needs to achieve consistent predictions between the two. To strengthen this goal, a consistency loss is introduced. The consistency loss is calculated by the L2 distance between the source domain detection loss and the target stylized detection loss The formula is: The formula is:

[0089]

[0090] S703: In the teacher-student co-learning stage, a distillation loss is introduced to enhance the generalization ability of the model. The teacher model generates pseudo-labels t from the target domain data I (s←t) and the source domain stylized data I and guides the training of the student model on the target data, and finally obtains the distillation loss The formula is:

[0091]

[0092] S704: The final total loss is the weighted sum of the detection loss, the consistency loss, and the distillation loss. The distillation loss is only introduced in the teacher-student co-learning stage. Appropriate weighting coefficients can balance the influence of each loss term on the model optimization.

[0093] The overall total loss expression is:

[0094]

[0095] where α, β, and γ are the weighting coefficients of each loss term, and Only used in the teacher-student co-learning stage.

[0096] The following specifically describes the relevant details of the method:

[0097] (1) Multi-scale decoding query clustering: This method aims to enhance the detection performance of the source domain training model on multi-scale targets in different target domains. This module first extracts features from the multi-scale object queries output by the decoder, and groups these queries according to the scale of their proportion of the image area. The average value of the features calculated for different categories and different scales is used to generate multi-scale category prototypes, and each prototype represents the feature center of a class of targets at different scales. These prototypes not only retain the key information of the category but also can adapt to the changes of the target at different scales. Then, this module calculates the cosine similarity between each decoded query and various category prototypes, and uses the obtained similarity value as the confidence score of the query. This confidence score is different from the detection score generated by the traditional detection head and can generate a more accurate confidence score for a specific domain, thereby more effectively distinguishing reliable pseudo-labels, difficult pseudo-labels, and unreliable pseudo-labels in different domains. In addition, this similarity calculation based on multi-scale prototypes can effectively capture the features of targets at different scales in different target domain images, thereby improving the performance of the model in cross-domain detection tasks.

[0098] (2) Dual-domain collaboration and dual-domain verification: This module improves the quality of pseudo-labels by combining the complementary information of target domain and source domain style images. In the dual-domain collaboration module, the target domain and source domain style images are inferred through a multi-scale decoder to generate high-confidence pseudo-labels, and these labels are screened based on the clustering confidence. In this process, the style differences between the source domain and the target domain are taken into account to ensure the reliability of the pseudo-labels. In the dual-domain verification module, the consistency between the target domain and source domain style images is used to further screen out difficult samples with high overlap and use them as reliable pseudo-labels. Finally, the pseudo-labels generated by these two modules are fused. By introducing a category-based non-maximum suppression (NMS) method, redundant and overly overlapping pseudo-labels are removed to ensure that the finally generated pseudo-labels are more accurate and representative. The combination of this dual-domain collaboration and verification can effectively improve the detection performance of the target domain, reduce pseudo-label noise, and enhance the generalization ability of the model.

[0099] To verify the effectiveness of the method of the present invention, the present invention conducts experiments on four mainstream datasets for unsupervised domain adaptation object detection, including Cityscapes, Foggy Cityscapes, Bdd100k, and Sim10k:

[0100] Cityscapes is a high-resolution dataset containing urban street scenes, mainly used for scene understanding, with 2,975 training images and 500 validation images, suitable for cross-weather and cross-scene adaptation research.

[0101] Foggy Cityscapes is a target dataset generated based on the Cityscapes dataset through a haze synthesis algorithm, used to evaluate cross-weather generalization under low visibility, with special attention paid to scenes with a haze density of 0.02.

[0102] BDD100k is a large-scale dataset containing 36,278 daytime driving images, used for cross-scene adaptation experiments to evaluate the adaptability to 7 types of targets under different daytime lighting and visual conditions.

[0103] Sim10k is a synthetic dataset containing 10,000 virtual street scene images annotated with vehicle instances, used for adaptation from synthetic to real scenes to evaluate the detection ability of the model from virtual environments to real-world urban environments.

[0104] Evaluation metrics: This method uses AP50 as the evaluation metric, which represents the average precision of the model in the object detection task when the intersection over union (IoU) threshold is 0.5. By taking the weighted average of the AP50 values for all classes, the mAP metric is obtained to comprehensively evaluate the overall performance of the multi-class object detection model.

[0105] Experimental metrics:

[0106] Table 1: Evaluate the experimental performance on Cityscapes to Foggy Cityscapes.

[0107]

[0108]

[0109] Table 2: Evaluate the experimental performance on Cityscapes to Bdd100k.

[0110]

[0111] Table 3: Evaluate the experimental performance on Sim10k to Cityscapes.

[0112]

[0113]

[0114] From the experimental results in Table 1, Table 2, and Table 3, it can be seen that the proposed method is significantly superior to the existing methods, improving the detection performance under different cross-domain tasks and demonstrating the superiority of the proposed method.

[0115] Embodiment 2

[0116] A cross-domain object detection device 401 for generating dual-domain pseudo-labels based on DETR, as Figure 6 shown, includes: a processor 402 and a storage device 403; the processor 402 loads and executes the instructions and data in the storage device 403 to implement the cross-domain object detection method for generating dual-domain pseudo-labels based on DETR.

[0117] Embodiment 3

[0118] A storage device that stores instructions and data for implementing the cross-domain object detection method for generating dual-domain pseudo-labels based on DETR.

[0119] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A cross-domain target detection method based on DETR dual-domain pseudo-label generation, applied to the unsupervised domain adaptive target detection task in computer vision, the method comprising: S1: Build a teacher-student model framework on the DETR-like detector, which consists of a trained student model and an inference teacher model; S2: Obtain source and target domain images, train the CUT model, and infer the corresponding source and target domain style images through the CUT model; S3: Use the pre-trained model trained on the COCO dataset, and further train the student model and teacher model with labeled source domain images and target domain style images to obtain the initialized student model and teacher model; S4: Construct a multi-scale decoding query clustering module. The teacher model infers the target domain and source domain style images to obtain decoding queries, and calculates the query similarity as confidence through the multi-scale decoding query clustering module. Each query represents a potential object, including category information and bounding box prediction information. S5: Build a dual-domain collaboration module to fuse the high-confidence results of style image reasoning in the target domain and the source domain as reliable pseudo labels; S6: Construct a dual-domain verification module to extract high-overlapping samples whose IOU values ​​of the same category detection boxes exceed the threshold from the difficult samples inferred from the target domain and source domain style images as reliable pseudo labels, and merge them with the reliable pseudo labels generated by the dual-domain collaboration module to obtain the final pseudo labels; S7: The final pseudo-label is used as the label of the target domain image for training the student model. At the same time, the teacher model is continuously updated, and finally a trained teacher-student model is obtained for cross-domain target detection.

2. The cross-domain target detection method based on DETR for dual-domain pseudo-label generation as claimed in claim 1, characterized in that: Both the teacher model and the student model are composed of DAB Deformable DETR, where DAB Deformable DETR uses ResNet-50 as the feature extractor. The student model updates the gradient through back propagation and loss value, and the teacher model updates the student model parameters through EMA. The formula is: in, and They represent the parameter weights of the teacher model and the student model at the kth iteration respectively, and α is the smoothing coefficient of EMA, which is used to control the speed of parameter updating.

3. The cross-domain target detection method based on DETR for dual-domain pseudo-label generation as claimed in claim 1, characterized in that: The specific implementation process of step S2 is as follows: S201: Acquire image data of source domain and target domain with different visual styles, and train a CUT model based on these data; S202: After the training is completed, the CUT model is used to perform style conversion reasoning on the images in the source domain and the target domain to generate corresponding source domain style images and target domain style images, respectively.

4. The cross-domain target detection method based on DETR dual-domain pseudo-label generation as claimed in claim 1, characterized in that: The specific implementation process of step S3 is as follows: S301: Download the detection model trained on the COCO dataset, and modify the detection head network according to the requirements of the actual cross-domain task to obtain a pre-trained model; S302: Further supervised training is performed on the pre-trained model through labeled source domain images and target domain style images, and the student model or teacher model with the best performance is selected in the process.

5. The cross-domain target detection method based on DETR for dual-domain pseudo-label generation as claimed in claim 1, characterized in that: The specific implementation process of step S4 is as follows: S401: Given a target image, the encoder of the teacher model first extracts image features and generates encoded features: in, is the i-th image in the target domain, is the encoding feature, Encoder is the encoder, and Backbone is the image feature extractor; The decoder then uses the encoded features output by the encoder and objects Generate a set of decoded query objects: in, Represents the decoded object query, N is the number of queries, d is the dimension of each query, and Encoder is the decoder; Decoded object query After being processed by the detection head, the inferred target information is obtained: in, To test the confidence level, To predict the class label, is the refined bounding box, Detection Head represents the detection head; S402: Select detection confidence Exceeding the threshold τ r The query is used as a reliable decoding object Scale information Indicates the area ratio and scale information of an object The calculation formula is: in, is the refined bounding box area, is the total area of ​​the input image, represents the i-th target domain image; S403: Aggregate similar object features in the feature space to form a consistent and compact set of category-scale prototypes; extract reliable decoding objects from formulas (4) and (5) The predicted class label of and scale information Then, features with the same category c and the same scale r are aggregated, and the category-scale prototype is generated by feature averaging in, represents the reliable object output by the decoder, and Respectively represent the category and scale information of the corresponding object; the symbol X{·} is an indicator function, when the condition is met, χ{·}=1, otherwise it is 0; N t is the total number of target domain images, For each target domain image The total number of objects in is the category scale prototype, where N class is the number of categories, N scale is the number of scale levels, d is the feature dimension; by clustering features of the same category and scale, the generated category scale prototype aims to accurately represent the features of objects under different categories and scales; in the teacher-student mutual learning stage, the category scale prototype It will be continuously updated as the training progresses; the updating process is the same as described in formula (6), and the updated prototype is generated by aggregating and averaging the features of the same category c and scale level r This enables continuous optimization; S404: Decoding query generated by computing decoder and pre-built category scale prototype A new confidence evaluation method is designed based on the cosine similarity between the query and the object. The specific process is as follows: As shown in formula (7), the prototype most similar to the query is selected as the confidence source of the query, and the maximum similarity is regarded as the clustering confidence of the object; at the same time, the category index related to the category scale prototype with the highest cosine similarity is assigned as the category label of the query, so as to re-evaluate the confidence of the query and possibly change its initial category: in, is the clustering confidence score calculated by maximum similarity, is the category label of the jth object assigned by the maximum cosine similarity.

6. The cross-domain target detection method based on DETR for dual-domain pseudo-label generation as claimed in claim 1, characterized in that: The specific implementation process of step S5 is as follows: S501: A selection strategy based on cluster confidence scores is used to filter high-confidence pseudo labels in the target domain and source domain style images. In the target domain, only cluster confidence scores are used. Greater than a predefined reliability threshold The target is selected as the candidate target pseudo label in, is the bounding box of the jth object in the i-th target domain image, is the category label of the jth object, is the total number of objects in each target domain image; S502: Execute the multi-scale decoding query clustering module in the source domain style image to generate the category scale prototype of the source domain style Then, based on the cluster confidence score in the source domain style image Greater than the preset threshold To filter the source domain style image and obtain reliable pseudo labels in, is the bounding box of the jth object in the i-th source domain style image, is the category label of the jth object in the i-th source domain style image, For each source style image The total number of objects in S503: Combine the reliable pseudo labels generated in the target domain and the source domain style images to generate the final high-confidence target domain pseudo labels Then, a category-based non-maximum suppression method is applied to these two sets of pseudo-labels to eliminate redundant and overly overlapping predictions, thereby obtaining the final target domain reliable pseudo-labels. Among them, NMS class represents the category-based non-maximum suppression operation, ∪ represents the target domain pseudo label and source-domain style pseudo-labels The union of .

7. The cross-domain target detection method based on DETR and dual-domain pseudo-label generation as claimed in claim 1, characterized in that: The specific implementation process of step S6 is as follows: S601: The dual-domain verification module selects reliable pseudo labels with lower cluster confidence scores by exploiting the consistency between the target image and the corresponding source domain style image: First, select the cluster confidence scores below the preset threshold for the target domain data. The n candidate bounding boxes and the clustering confidence of the corresponding source domain style image are lower than the threshold The n candidate bounding boxes are denoted as and S602: Quantify using the intersection over union (IoU) metric and The overlap between samples of the same category in , and pseudo labels with higher overlap are selected as reliable difficult samples The formula is: in, represents the bounding box of the kth object in the i-th source style image, represents the category label of the kth object in the i-th source domain style image, Represents the bounding box of the kth object in the i-th target domain image; when the intersection over union (IoU) exceeds δ oss When , the sample is retained, where δ oss is the overlap threshold; S603: Since a single foreground object may be captured by multiple decoded object queries with different confidence scores, resulting in the same candidate bounding box being selected in both reliable pseudo-labels and difficult pseudo-labels, in order to reduce confusion, and Perform category-based non-maximum suppression, the formula is: in, represents the final pseudo-label set of the i-th target domain image after merging the target domain and source domain style data, NMS class represents the category-level non-maximum suppression operation applied to remove redundant labels, represents the reliable pseudo-label of the target domain, represents high-confidence target domain pseudo-label.

8. The cross-domain target detection method based on DETR and dual-domain pseudo-label generation as claimed in claim 1, characterized in that: The specific implementation process of step S7 is as follows: S701: The training of the teacher-student model is divided into two stages. In the initial stage, a dual-source domain supervised training strategy is adopted to gradually guide the model to adapt to the target domain using source domain data and target domain style data, enhance domain-specific learning and promote effective transfer; for supervised training loss, the formula is: in, represents the detection loss, represents the source domain detection loss, represents the target domain stylized detection loss; S702: The source domain and target domain style data share the same foreground objects and labels, so the model needs to achieve consistent predictions between the two. To strengthen this goal, the consistency loss is introduced, and the source domain detection loss is used. and target stylized detection loss The L2 distance between them is used to calculate the consistency loss The formula is: S703: In the teacher-student joint learning stage, distillation loss is introduced to enhance the generalization ability of the model; the teacher model is trained from the target domain data I t and source domain stylized data I (s←t) Generate pseudo labels And guide the student model to train on the target data, and finally get the distillation loss The formula is: S704: The final total loss is the weighted sum of detection loss, consistency loss and distillation loss, where distillation loss is only introduced in the teacher-student joint learning stage; appropriate weighting coefficients are used to balance the impact of each loss term on model optimization; The overall total loss expression is: Among them, α, β and γ are the weighted coefficients of each loss term, and Only used during the teacher-student joint learning phase.

9. A storage device, characterized in that: The storage device stores instructions and data for implementing the cross-domain target detection method based on DETR for dual-domain pseudo-label generation as described in any one of claims 1 to 8.

10. A cross-domain target detection device based on DETR with dual-domain pseudo-label generation, characterized by: include: Processor and storage device; the processor loads and executes instructions and data in the storage device to implement the cross-domain target detection method based on DETR dual-domain pseudo-label generation as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Multi-stage guided small target semi-supervised learning detection method based on uncertainty

    CN116563738A

  • Target detection method and system for open scene

    CN117853792A

  • Semi-supervised medical image segmentation method based on clustering fusion cross learning

    CN118397272A

  • Domain adaptation using pseudo-labelling and model certainty quantification for video data

    US20220301287A1

  • Source-free cross domain detection method with strong data augmentation and self-trained mean teacher modeling

    US20230154167A1

Cited By

  • Continuous test adaptive method for class balance teacher based on SAM guidance

    CN120451200A

  • Entropy-difference-guided source-domain-free transfer learning image target detection method

    CN120783033A

  • Entropy difference guided unsupervised domain adaptation learning image object detection method

    CN120783033B

  • Daytime and nighttime cross-domain target detection method and system based on reliable teacher model

    CN120976526A

  • Cross-domain adaptive target detection method based on data deviation and training deviation

    CN121033371A