Efficient method for unsupervised domain adaptive target detection
By designing the domain query module and category prototype alignment module in the object detector, using adversarial loss and contrast learning, extracting domain invariant features and aligning category features of different fields, the performance degradation of traditional object detectors under domain offset is solved, and excellent detection performance on multiple domain adaptation benchmark data sets is achieved.
Patent Information
- Application Number
- CN202411953308.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-12-27
AI Technical Summary
The performance of traditional object detectors significantly decreases when facing domain offsets, and existing methods are difficult to effectively solve this problem, especially in the context of unsupervised domain adaptation.
The domain query module and category prototype alignment module were designed. Through adversarial loss and contrast learning, domain invariant features are extracted and category features of different fields are aligned, and the UDA-DETR model is constructed to alleviate domain offset.
By extracting more domain-invariant features and achieving cross-domain category feature alignment, UDA-DETR significantly improves detection performance on the target domain and effectively alleviates the performance degradation caused by domain offset.
Smart Images

Figure CN119992044A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to an efficient method for unsupervised domain adaptive target detection. Background Art
[0002] Object detection has always been a hot area in computer vision research, aiming to predict the bounding box coordinates and category labels of objects of interest in images, which is crucial for daily applications such as security monitoring and unmanned driving. Initially, most object detection methods used two-stage detectors, involving the design of anchor boxes and prior boxes, the extraction of regions of interest, and tedious post-processing steps such as non-maximum suppression (NMS). As research progressed, one-stage detectors based on Transformer were proposed and quickly attracted attention. This type of method removes the manually designed prior boxes and tedious post-processing steps, thereby greatly simplifying the detection process.
[0003] Although the latest detection methods have performed well in simplifying the process and improving performance, they usually rely on a large amount of labeled data for training. In addition, the performance of most detectors will drop significantly when there are significant feature differences between the training data (source domain) and the test data (target domain). This phenomenon is called domain shift. In the face of domain shift, the performance of most detectors usually drops significantly. Manually re-labeling data is a direct way to solve domain shift, but this method is often not feasible due to high cost and complexity. Therefore, utilizing unlabeled target domain data becomes a more realistic option, which promotes the development of unsupervised domain adaptation object detection (UDAOD) methods. In unsupervised domain adaptation, labeled source domain data and unlabeled target domain data are usually used to train the model, and the labeled data is only available in the source domain. By minimizing the domain difference through adversarial learning, the supervised detection model trained on the source domain can be better generalized and extended to the target domain for knowledge transfer and representation.
[0004] Some studies have shown that DETECTION Transformer (DETR) exhibits excellent performance in solving the domain shift phenomenon, and is usually better than pure convolution-based detectors. CNN and Transformer are essentially different: CNN captures visual features on feature maps through convolution blocks, while Transformer models features of data images through attention mechanisms. DETR innovatively integrates CNN and Transformer. The system consists of three parts, including a CNN-based backbone network and a Transformer-based encoder and decoder. The combination of the two significantly improves the feature extraction capability and eliminates tedious post-processing steps such as anchor design and non-maximum suppression (NMS), thereby greatly simplifying the training process.
[0005] Currently, many domain adaptation detectors based on DETR focus on the alignment of image-level features. These methods use adversarial feature alignment in the backbone network part to improve the detector's ability to extract domain-invariant features by confusing the source of domain features. In addition, some methods try to align instance-level features in the decoder part, usually using an image-level alignment mechanism. The latest research shows that the improvement effect of domain feature alignment applied only on the CNN backbone of DETR for domain adaptation is very limited.
[0006] Therefore, a new solution to the above problems needs to be proposed. Summary of the invention
[0007] The purpose of the present invention is to provide an efficient method for unsupervised domain adaptive object detection, aiming to solve the problem of performance degradation of traditional object detectors caused by domain shift.
[0008] To achieve the above object, the present invention provides the following technical solution: an efficient method for unsupervised domain adaptive target detection, comprising at least the following steps:
[0009] S1: Design a domain query module. Under the constraint of adversarial loss, the model focuses on the image features of a specific domain, so that the detector can extract more domain-invariant features and achieve domain-adaptive object detection through adversarial training.
[0010] S2: Design a category prototype alignment module, which focuses on aligning instance features of objects of the same category from different fields. In each model iteration, small batches of category prototype features are continuously updated, and these local category prototype features are stored to gradually generate category prototype features representing the entire data set. At the same time, through contrastive learning, the feature representations of the same category are made closer, while the feature representations of different categories are further apart, which effectively improves the model's ability to distinguish between objects in different fields.
[0011] S3: Design a dataset-level category prototype alignment strategy to make full use of the information of the entire dataset to improve the global representation ability of features and the distinction between categories, and enhance the model's ability to obtain contextual information at the global feature level;
[0012] S4: Build the UDA-DATR model based on the domain query module, category prototype alignment module and dataset-level category prototype alignment strategy;
[0013] S5: The total loss function of the UDA-DETR model is optimized by comprehensive supervision detection loss, adversarial loss and contrast loss. The purpose of optimization is to improve the detection performance on the target domain and alleviate the domain shift problem. The adversarial loss includes the backbone network, encoder and decoder.
[0014] S6: Object detection using the optimized UDA-DETR model.
[0015] Furthermore, the designing of the domain query module in S1 at least comprises the following steps:
[0016] Token-wise multi-scale feature alignment is performed inside the encoder of the model;
[0017] This module consists of a domain query and a cross-attention layer. The domain query is similar to the object queries in the decoder part. It is a set of trainable vectors used to identify domain-specific features of objects.
[0018] Inside the encoder, feature tokens provide key-value pairs, and the domain query is input as a query vector into the cross-attention layer to update the domain query.
[0019] The domain query tends to focus on feature tokens that are discriminative for domain classification, that is, those domain-specific features;
[0020] use represents the domain query after being updated by the cross-attention layer in the i-th encoder layer; Indicates the domain query before update; Represents the feature token in the i-th layer encoder;
[0021] The cross-attention layer receives two inputs, query and key-value pairs;
[0022] In the attention mechanism, the query focuses on specific keys in the key-value pair according to the attention weight and maps itself to the linear combination of the corresponding values;
[0023] Linear is a simple linear mapping layer used to perform this mapping;
[0024] The optimization method of domain query is shown in formula (1):
[0025]
[0026] The entire domain query module is optimized through binary cross entropy loss, as shown in formula (2):
[0027]
[0028] represents the total loss of the domain query module, which is used to train the domain classifier; d represents the domain label, where the value 0 represents the source domain image and the value 1 represents the target domain image;
[0029] Represents the domain prediction of the encoder domain classifier for the final domain query. The prediction value is in the range of [0,1], indicating the probability that the image belongs to the target domain;
[0030] By minimizing the loss function The domain query is trained to highlight the features that are most critical for distinguishing the source domain from the target domain; this setting forces the detector to extract more domain-invariant features to minimize the adversarial loss. This enables domain adaptation.
[0031] Furthermore, the designing of the category prototype alignment module in S2 at least comprises the following steps:
[0032] Receive object queries from the decoder, and aggregate the corresponding semantic information of each object query through the decoder's cross-attention mechanism;
[0033] For each object query, first determine its predicted category;
[0034] Then, the object queries of the same category are merged and the center of the merged aggregated features is calculated to obtain the central feature representation of each category. This process is shown in formula (3):
[0035]
[0036] Where Q n represents the decoded object representation of the N initial object queries output by the decoder; C n represents the predicted category of the nth object query; c represents the category index;
[0037] When C n = c, [C n =c] is 1, otherwise it is 0;
[0038] In this way, I get the feature prototype representation F for each category c ;
[0039] Then, the category prototype feature F cInput into the adversarial domain discriminator to distinguish whether the feature center is from the source domain or the target domain;
[0040] The domain discriminator is trained by minimizing the binary cross entropy loss while making the category feature prototype F c Align in two domains to achieve the purpose of deceiving the discriminator. The optimization method is shown in formula (4):
[0041]
[0042] Where D d (F c ) represents the output of the domain discriminator, d represents the domain label, a value of 0 represents the source domain, and a value of 1 represents the target domain.
[0043] Furthermore, the dataset-level category prototype alignment strategy includes at least the following steps:
[0044] The small batch category prototype features extracted in each iterative training are stored, and the average value of the stored local category prototype features is calculated to model them as dataset-level category prototype features, as shown in formula (5):
[0045]
[0046] Represents the dataset-level category prototype feature representation of category c, n c represents the number of object queries of category c in the current training iteration, and Then it represents the cumulative number of object queries of category c;
[0047] During the model training process, the prototype features of each category are continuously updated so that they can reflect the overall features of the category including the source domain and the target domain;
[0048] In this way, the model can learn more comprehensive and representative feature representations, thereby improving the performance of cross-domain object detection;
[0049] In order to further enhance the inter-category distinguishability of the model, contrastive learning is introduced to bring the features of the same category closer and exclude the features of different categories. The calculation method is shown in formula (6):
[0050]
[0051] F Si and F Ti Represent the central features in the source domain and the target domain respectively, represents the global category prototype feature, and C represents the total number of categories;
[0052] For each category, the central feature F Si and F Ti , and calculate them through the dot product with the global center feature F Ti The similarity between
[0053] Then, the softmax function is used to calculate the relative probability of the similarity between the local prototype feature of each category and the global prototype feature. For the local prototype feature of each category, the logarithmic probability of its similarity with the global prototype feature is calculated, and the logarithmic probabilities of all categories are summed up. Finally, the logarithmic probabilities of all categories are averaged to obtain the contrastive learning loss L con .
[0054] Furthermore, the S5 at least includes the following steps:
[0055] Set the source domain dataset with annotation information to and target domain dataset without annotation information
[0056] Among them, x represents the image itself, y = (b, c) represents the annotation information corresponding to the image, t represents the image from the target domain, s represents the image from the source domain, i represents the current image, and N s Represents the number of source domain images, N t represents the number of target domain images, and the annotation information includes the bounding box position b and the category label c;
[0057] The proposed UDA-DATR model is trained on the source domain and its performance is evaluated on the target domain. This setting aims to improve the detection performance of the model in the target domain by effectively utilizing the annotated data in the source domain, thereby alleviating the impact of domain shift.
[0058] Deformable-DETR is selected as the basic detector of the model, and the Deformable-DETR includes at least a backbone network, an encoder and a decoder;
[0059] The source domain image is obtained by calculating the supervised loss L in supervised learning. sup To optimize the model, the loss is calculated on the source domain, as shown in formula (7):
[0060]
[0061] x s represents the source domain image, y s Represents the label value corresponding to the source domain image, represents the image bounding box regression loss, represents the giou loss of the source domain image, represents the category loss of the source domain image;
[0062] We further summarize the overall goal of UDA-DETR, which consists of three adversarial losses: A contrastive loss L con and a supervised detection loss composition;
[0063] Combining the above, we can construct the total loss function L UDA-DETR , as shown in formula (8):
[0064]
[0065] Compared with the prior art, the present invention has the following beneficial effects:
[0066] The present invention designs a domain query module, which enables the detector to extract more domain-invariant features by adversarially aligning the features output by the encoder. In addition, a category prototype alignment module is proposed, which can extract the category prototype features of each category from the decoder to achieve category-aware cross-domain feature alignment. In each training iteration, the local category prototype gradually generates the global category prototype, and the detector is guided to achieve global feature alignment through contrastive learning and adversarial loss. A large number of experimental results show that the proposed model achieves excellent detection performance on multiple domain adaptation benchmark datasets. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for describing the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0068] Figure 1 This is a schematic diagram of the model architecture of the present invention;
[0069] Figure 2 This is a visualization comparison of the detection effect diagram of the present invention and other advanced domain adaptation methods;
[0070] Figure 3 This is a diagram of the detection effect achieved when the present invention uses different modules;
[0071] Figure 4 This is a thermal diagram of the model used by the present invention for detection in severe weather scenarios. DETAILED DESCRIPTION
[0072] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.
[0073] When there is a domain gap between the training data (source domain data) and the data to be detected in actual applications (target domain data), the performance of many object detectors will drop significantly. To address this problem, many unsupervised domain adaptation detectors narrow the gap by minimizing the domain difference in the category space and aligning the global features of the two domains. Although these techniques have achieved some success, they usually use a category-agnostic way to perform global feature alignment, ignoring the differences in the distribution of features of different categories in different domains, resulting in classification and localization errors.
[0074] The method proposed in the present invention essentially proposes an efficient unsupervised domain adaptation detection architecture (UDA-DETR), which pays more attention to global image features and local instance features from different domains to achieve more accurate domain adaptive detection.
[0075] See also Figure 1 , an efficient method for unsupervised domain adaptive object detection, comprising at least the following steps:
[0076] S1: Design a domain query module. Under the constraint of adversarial loss, the model focuses on the image features of a specific domain, so that the detector can extract more domain-invariant features and achieve domain-adaptive object detection through adversarial training.
[0077] S2: Design a category prototype alignment module, which focuses on aligning instance features of objects of the same category from different fields. In each model iteration, small batches of category prototype features are continuously updated, and these local category prototype features are stored to gradually generate category prototype features representing the entire data set. At the same time, through contrastive learning, the feature representations of the same category are made closer, while the feature representations of different categories are further apart, which effectively improves the model's ability to distinguish between objects in different fields.
[0078] S3: Design a dataset-level category prototype alignment strategy to make full use of the information of the entire dataset to improve the global representation ability of features and the distinction between categories, and enhance the model's ability to obtain contextual information at the global feature level;
[0079] S4: Build the UDA-DATR model based on the domain query module, category prototype alignment module and dataset-level category prototype alignment strategy;
[0080] S5: The total loss function of the UDA-DETR model is optimized by comprehensive supervision detection loss, adversarial loss and contrast loss. The purpose of optimization is to improve the detection performance on the target domain and alleviate the domain shift problem. The adversarial loss includes the backbone network, encoder and decoder.
[0081] S6: Object detection using the optimized UDA-DETR model.
[0082] The UDA-DATR model framework is as follows Figure 1 As shown, for the input image data x, it is first processed by the backbone network to obtain the corresponding feature map z = R C×H×W , C represents the number of channels, H represents the height, and W represents the width. Then the feature map is flattened into z P =R C×L , L = H × W represents the number of tokens, and is input into the encoder to obtain higher-dimensional semantic information. Next, the decoder performs target detection based on the set N object queries (N = 300), and finally obtains the detection loss In addition, we perform category-independent image-level feature alignment in the backbone network and obtain the loss In the encoder part, the domain query model is embedded and token-wise multi-scale feature alignment is performed to obtain the loss In the decoder part, the category prototype feature alignment is designed and combined with contrastive learning to obtain and L con .
[0083] Designing a domain query module in S1 includes at least the following steps:
[0084] Token-wise multi-scale feature alignment is performed inside the encoder of the model;
[0085] This module consists of a domain query and a cross-attention layer. The domain query is similar to the object queries in the decoder part. It is a set of trainable vectors used to identify domain-specific features of objects.
[0086] Inside the encoder, feature tokens provide key-value pairs, and the domain query is input as a query vector into the cross-attention layer to update the domain query.
[0087] Domain queries tend to focus on feature tokens that are discriminative for domain classification, that is, those domain-specific features;
[0088] use represents the domain query after being updated by the cross-attention layer in the i-th encoder layer; Indicates the domain query before update; Represents the feature token in the i-th layer encoder;
[0089] The cross-attention layer receives two inputs, query and key-value pairs;
[0090] In the attention mechanism, the query focuses on specific keys in the key-value pair according to the attention weight and maps itself to the linear combination of the corresponding values;
[0091] Linear is a simple linear mapping layer used to perform this mapping;
[0092] The optimization method of domain query is shown in formula (1):
[0093]
[0094] The entire domain query module is optimized through binary cross entropy loss, as shown in formula (2):
[0095]
[0096] represents the total loss of the domain query module, which is used to train the domain classifier; d represents the domain label, where the value 0 represents the source domain image and the value 1 represents the target domain image;
[0097] Represents the domain prediction of the encoder domain classifier for the final domain query. The prediction value is in the range of [0,1], indicating the probability that the image belongs to the target domain;
[0098] By minimizing the loss function The domain query is trained to highlight the features that are most critical for distinguishing the source domain from the target domain; this setting forces the detector to extract more domain-invariant features to minimize the adversarial loss. This enables domain adaptation.
[0099] The design category prototype alignment module in S2 includes at least the following steps:
[0100] Receive object queries from the decoder, and aggregate the corresponding semantic information of each object query through the decoder's cross-attention mechanism;
[0101] For each object query, first determine its predicted category;
[0102] Then, the object queries of the same category are merged and the center of the merged aggregated features is calculated to obtain the central feature representation of each category. This process is shown in formula (3):
[0103]
[0104] Where Q n represents the decoded object representation of the N initial object queries output by the decoder; C n represents the predicted category of the nth object query; c represents the category index;
[0105] When C n = c, [C n =c] is 1, otherwise it is 0;
[0106] In this way, I get the feature prototype representation F for each category c ;
[0107] Then, the category prototype feature F c Input into the adversarial domain discriminator to distinguish whether the feature center is from the source domain or the target domain;
[0108] The domain discriminator is trained by minimizing the binary cross entropy loss while making the category feature prototype F c Align in two domains to achieve the purpose of deceiving the discriminator. The optimization method is shown in formula (4):
[0109]
[0110] Where D d (F c ) represents the output of the domain discriminator, d represents the domain label, a value of 0 represents the source domain, and a value of 1 represents the target domain.
[0111] The dataset-level category prototype alignment strategy includes at least the following steps:
[0112] The small batch category prototype features extracted in each iterative training are stored, and the average value of the stored local category prototype features is calculated to model them as dataset-level category prototype features, as shown in formula (5):
[0113]
[0114] Represents the dataset-level category prototype feature representation of category c, n crepresents the number of object queries of category c in the current training iteration, and Then it represents the cumulative number of object queries of category c;
[0115] During the model training process, the prototype features of each category are continuously updated so that they can reflect the overall features of the category including the source domain and the target domain;
[0116] In this way, the model can learn more comprehensive and representative feature representations, thereby improving the performance of cross-domain object detection;
[0117] In order to further enhance the inter-category distinguishability of the model, contrastive learning is introduced to bring the features of the same category closer and exclude the features of different categories. The calculation method is shown in formula (6):
[0118]
[0119] F Si and F Ti Represent the central features in the source domain and the target domain respectively, represents the global category prototype feature, and C represents the total number of categories;
[0120] For each category, the central feature F Si and F Ti , and calculate them through the dot product with the global center feature F Ti The similarity between
[0121] Then, the softmax function is used to calculate the relative probability of the similarity between the local prototype feature of each category and the global prototype feature. For the local prototype feature of each category, the logarithmic probability of its similarity with the global prototype feature is calculated, and the logarithmic probabilities of all categories are summed up. Finally, the logarithmic probabilities of all categories are averaged to obtain the contrastive learning loss L con .
[0122] S5 at least includes the following steps:
[0123] Assume that the source domain dataset with annotation information is and target domain dataset without annotation information
[0124] Among them, x represents the image itself, y = (b, c) represents the annotation information corresponding to the image, t represents the image from the target domain, s represents the image from the source domain, i represents the current image, and N s Represents the number of source domain images, N t represents the number of target domain images, and the annotation information includes the bounding box position b and the category label c.
[0125] The proposed UDA-DATR model is trained on the source domain and its performance is evaluated on the target domain. This setting aims to improve the detection performance of the model in the target domain by effectively utilizing the annotated data in the source domain, thereby alleviating the impact of domain shift.
[0126] Deformable-DETR is selected as the basic detector of the model. Deformable-DETR includes at least a backbone network, an encoder, and a decoder.
[0127] The source domain image is obtained by calculating the supervised loss L in supervised learning. sup To optimize the model, the loss is calculated on the source domain, as shown in formula (7):
[0128]
[0129] x s represents the source domain image, y s Represents the label value corresponding to the source domain image, represents the image bounding box regression loss, represents the giou loss of the source domain image, represents the category loss of the source domain image;
[0130] We further summarize the overall goal of UDA-DETR, which consists of three adversarial losses: A contrastive loss L con and a supervised detection loss composition;
[0131] Combining the above, we can construct the total loss function L UDA-DETR , as shown in formula (8):
[0132]
[0133] In order to demonstrate the superior performance of the UDA-DETR proposed in the present invention, comparative experiments were conducted with many advanced unsupervised domain adaptation methods on three domain adaptation benchmark datasets. The experimental results show that the method of the present invention is significantly superior to all previous advanced methods.
[0134] The present invention first studies the domain adaptation detection performance in different weather scenarios, taking the normal weather scene Cityscapes as the source domain and the foggy weather scene Foggy Cityscapes as the target domain. The experimental results are shown in Table 1. The UDA-DETR of the present invention achieves the best 46.0mAP. For some difficult-to-detect categories (such as rider), UDA-DETR still achieves the best AP of 50.3.
[0135] Secondly, the present invention studies the domain adaptation from Cityscapes to BDD100k-daytime, and the experimental results are shown in Table 2. Due to the large difference in feature distribution between the source domain and the target domain, most unsupervised domain adaptation methods have poor effects, but the UDA-DETR method of the present invention still achieves the best mAP of 30.2.
[0136] Finally, we study the domain adaptation from Sim10k to Cityscapes (Car), and the experimental results are shown in Table 3. The UDA-DETR method proposed in this paper significantly outperforms the previous state-of-the-art technology (SOTA) and achieves the best mAP of 55.1.
[0137] Also see Figure 2 , 3 4 is compared with the above. Figure 4 The darker the color, the more attention the model pays to the area;
[0138] The above experimental results prove that the UDA-DETR method proposed in this paper can show excellent performance in various domain adaptation scenarios.
[0139] Table 1. Experimental results of domain adaptation across different weather scenarios: Cityscapes to Foggy Cityscapes
[0140]
[0141] Table 2. Experimental results of cross-shot scene domain adaptation: Cityscapes to BDD100k-daytime
[0142]
[0143]
[0144] Table 3. Experimental results of domain adaptation from synthetic scenes to real scenes: Sim10kto Cityscapes results
[0145]
[0146] It will be apparent to those skilled in the art that the invention is not limited to the details of the exemplary embodiments described above and that the invention can be implemented in other specific forms without departing from the spirit or essential features of the invention. Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description, and it is intended that all variations falling within the meaning and scope of the equivalent elements of the claims be included in the invention. Any reference numeral in a claim should not be considered as limiting the claim to which it relates.
Claims
1. An efficient method for unsupervised domain adaptive object detection, characterized by: At least the following steps are included: S1: Design a domain query module. Under the constraint of adversarial loss, the model focuses on the image features of a specific domain, so that the detector can extract more domain-invariant features and achieve domain-adaptive object detection through adversarial training. S2: Design a category prototype alignment module, which focuses on aligning instance features of objects of the same category from different fields. In each model iteration, small batches of category prototype features are continuously updated, and these local category prototype features are stored to gradually generate category prototype features representing the entire data set. At the same time, through contrastive learning, the feature representations of the same category are made closer, while the feature representations of different categories are further apart, which effectively improves the model's ability to distinguish between objects in different fields. S3: Design a dataset-level category prototype alignment strategy to make full use of the information of the entire dataset to improve the global representation ability of features and the distinction between categories, and enhance the model's ability to obtain contextual information at the global feature level; S4: Build the UDA-DATR model based on the domain query module, category prototype alignment module and dataset-level category prototype alignment strategy; S5: The total loss function of the UDA-DETR model is optimized by comprehensive supervision detection loss, adversarial loss and contrast loss. The purpose of optimization is to improve the detection performance on the target domain and alleviate the domain shift problem. The adversarial loss includes the backbone network, encoder and decoder. S6: Object detection using the optimized UDA-DETR model.
2. The efficient method for unsupervised domain adaptive object detection according to claim 1, characterized in that: The designing of the domain query module in S1 at least comprises the following steps: Multi-scale feature alignment at the token-wise level is performed inside the encoder of the model; This module consists of a domain query and a cross-attention layer. The domain query is similar to the object query in the decoder part. It is a set of trainable vectors used to identify domain-specific features of objects. Inside the encoder, the feature tokens provide key-value pairs, and the domain query is input as a query vector into the cross-attention layer to update the domain query. The domain query tends to focus on feature tokens that are discriminative for domain classification, that is, those domain-specific features; use represents the domain query after being updated by the cross-attention layer in the i-th encoder layer; Indicates the domain query before update; Represents the feature token in the i-th layer encoder; The cross attention layer receives two inputs, query and key-value pair; In the attention mechanism, the query focuses on a specific key in the key-value pair according to the attention weight and maps itself to a linear combination of the corresponding values; Linear is a simple linear mapping layer used to perform this mapping; The optimization method of domain query is shown in formula (1): The entire domain query module is optimized through binary cross entropy loss, as shown in formula (2): represents the total loss of the domain query module, which is used to train the domain classifier; d represents the domain label, where the value 0 represents the source domain image and the value 1 represents the target domain image; Represents the domain prediction of the encoder domain classifier for the final domain query. The prediction value is in the range of [0,1], indicating the probability that the image belongs to the target domain; By minimizing the loss function The domain query is trained to highlight the features that are most critical for distinguishing the source domain from the target domain; this setting forces the detector to extract more domain-invariant features to minimize the adversarial loss. This enables domain adaptation.
3. The efficient method for unsupervised domain adaptive object detection according to claim 2, characterized in that: The design of the category prototype alignment module in S2 at least includes the following steps: Receive the object queries output from the decoder, and aggregate the corresponding semantic information of each object query through the decoder's cross-attention mechanism; For each object query, first determine its predicted category; Then, the object queries of the same category are merged and the center of the merged aggregated features is calculated to obtain the central feature representation of each category. This process is shown in formula (3): Where Q n represents the decoded object representation of the N initial object queries output by the decoder; C n represents the predicted category of the nth object query; c represents the category index; When C n = c, [C n =c] is 1, otherwise it is 0; In this way, I get the feature prototype representation F for each category c ; Then, the category prototype feature F c Input into the adversarial domain discriminator to distinguish whether the feature center is from the source domain or the target domain; The domain discriminator is trained by minimizing the binary cross entropy loss while making the category feature prototype F c Align in two domains to achieve the purpose of deceiving the discriminator. The optimization method is shown in formula (4): Where D d (F c ) represents the output of the domain discriminator, d represents the domain label, a value of 0 represents the source domain, and a value of 1 represents the target domain.
4. The efficient method for unsupervised domain adaptive object detection according to claim 3, characterized in that: The dataset-level category prototype alignment strategy includes at least the following steps: The small batch category prototype features extracted in each iterative training are stored, and the average value of the stored local category prototype features is calculated to model them as dataset-level category prototype features, as shown in formula (5): Represents the dataset-level category prototype feature representation of category c, n c represents the number of object queries of category c in the current iteration training, Then it represents the cumulative number of object queries of category c; During the model training process, the prototype features of each category are continuously updated so that they can reflect the overall features of the category including the source domain and the target domain; In this way, the model can learn more comprehensive and representative feature representations, thereby improving the performance of cross-domain object detection; In order to further enhance the inter-category distinguishability of the model, contrastive learning is introduced to bring the features of the same category closer and exclude the features of different categories. The calculation method is shown in formula (6): F Si and F Ti Represent the central features in the source domain and the target domain respectively, represents the global category prototype feature, and C represents the total number of categories; For each category, the central feature F Si and F Ti , and calculate them through the dot product with the global center feature F Ti The similarity between Then, the softmax function is used to calculate the relative probability of the similarity between the local prototype feature of each category and the global prototype feature. For the local prototype feature of each category, the logarithmic probability of its similarity with the global prototype feature is calculated, and the logarithmic probabilities of all categories are summed up. Finally, the logarithmic probabilities of all categories are averaged to obtain the contrastive learning loss L con .
5. The efficient method for unsupervised domain adaptive object detection according to claim 4, characterized in that: The S5 at least comprises the following steps: Assume that the source domain dataset with annotation information is and target domain dataset without annotation information Among them, x represents the image itself, y = (b, c) represents the annotation information corresponding to the image, t represents the image from the target domain, s represents the image from the source domain, i represents the current image, and N s Represents the number of source domain images, N t represents the number of target domain images, and the annotation information includes the bounding box position b and the category label c; The proposed UDA-DATR model is trained on the source domain and its performance is evaluated on the target domain. This setting aims to improve the detection performance of the model in the target domain by effectively utilizing the annotated data in the source domain, thereby alleviating the impact of domain shift. Deformable-DETR is selected as the basic detector of the model, and the Deformable-DETR includes at least a backbone network, an encoder and a decoder; The source domain image is obtained by calculating the supervised loss L in supervised learning. sup To optimize the model, the loss is calculated on the source domain, as shown in formula (7): x s represents the source domain image, y s Represents the label value corresponding to the source domain image, represents the image bounding box regression loss, represents the giou loss of the source domain image, represents the category loss of the source domain image; We further summarize the overall goal of UDA-DETR, which consists of three adversarial losses: A contrastive loss L con and a supervised detection loss composition; Combining the above items, construct the total loss function L UDA-DETR , as shown in formula (8):
Citation Information
Patent Citations
Self-adaptive cross-domain target detection method based on uncertainty guidance
CN113392933A
Training method based on image-instance alignment network and cross-domain target detection method
CN114693983A
Cross-domain target detection method based on comparative learning
CN116309466A
Method and system for detecting objects in images
EP4120132A1
Domain adaptation system and method therefor
JP2024110191A
Cited By
Cross-domain identification method for FeO content of sintered ore
CN120953706A
Photovoltaic power station detection method based on feature decoupling domain adaptation
CN121685926A