A Domain Adaptation Object Detection Method and Device Based on Contrastive Learning

Through the domain adaptive object detection method based on comparison learning, a domain adaptive object detection network is built, and the comparison learning module is used to align the characteristics of the source domain and the target domain, which solves the shortcomings of the existing methods in eliminating distribution differences and achieving high-precision and high recall object detection effect.

CN114972964BActive Publication Date: 2025-05-30INSTITUTE OF INFORMATION ENGINEERING CHINESE ACADEMY OF SCIENCES
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210397702.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-15
Publication Date
2025-05-30
Estimated Expiration
2042-04-15

AI Technical Summary

Technical Problem

The existing domain-adapted object detection methods have shortcomings in finding similarities between the source domain and the target domain while preserving them, which leads to biasing the detection results towards the source domain, and problems such as error detection and missed detection are encountered.

Method used

The domain adaptive object detection method based on contrast learning is adopted, and the domain adaptive object detection network consisting of a feature decoupling extraction module, an object detection module and a comparison learning module are constructed, and the annotated source domain image data set and annotated target domain image data set are trained to achieve the alignment of the source domain and the target domain features and the extraction of domain invariant features.

Benefits of technology

The consistency of domain invariant features is improved and the object detection effect on the target domain is improved. The results have high accuracy and high recall rate, indicating that the network has strong domain adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114972964B_ABST
    Figure CN114972964B_ABST
Patent Text Reader

Abstract

The present invention discloses a domain adaptation object detection method and device based on contrastive learning, belonging to the field of computer vision technology. By using the ResNet network in the deep neural network, in cooperation with contrastive learning and a feature decoupling module, it can match the features in the source domain and the target domain and decouple domain-invariant features therefrom when there are only labeled source domain image data and unlabeled target domain image data in the dataset, and perform object detection based on the domain-invariant features. During the training process, contrastive learning is used to achieve the alignment of the source domain and target domain features, improve the consistency of the decoupled features, obtain better domain-invariant features, and enhance the object detection effect on the target domain. The results have high precision and high recall rate, indicating that the network has strong domain adaptation ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision, and particularly relates to an object detection method and device capable of realizing domain adaptation. Background Art

[0002] Machine learning has been widely applied in many fields such as image recognition, object detection, etc. For common machine learning tasks, such as object detection, to achieve object detection in a specific environment, if only publicly available general datasets are used, there will be a problem of inconsistent sample distributions in the training set (source domain) and the test set (target domain), resulting in a poorly predicted model obtained through training; however, the cost of collecting and annotating a sufficiently large dataset for a specific task is extremely high. To solve such a contradiction, the academic community has proposed the concept of domain adaptation, the core of which is to find the similarity between the source domain and the target domain and utilize this similarity to apply the knowledge learned in the source domain to the target domain to complete the training of the model without annotating the target domain dataset. After applying the domain adaptation method, training for a specific task can be carried out using publicly available general datasets in combination with unannotated target domain datasets, greatly reducing the cost. For example, to train a street scene recognition model that can be used in a certain city, existing street scene datasets from other cities can be used in combination with unannotated street scene images of this city for training, saving the cost of annotating the street scene images of this city.

[0003] The main difficulty of the domain adaptation object detection method lies in correctly finding the similarities between the source domain and the target domain and preserving them, while at the same time eliminating the distribution differences between the source domain and the target domain. Since only data from the source domain is annotated during the training process, if the above two difficulties cannot be solved, it will lead to the detection results being biased towards the source domain, and problems such as false detection and missed detection will occur in the target domain. Existing domain adaptation object detection methods can be divided into three categories: one is the method based on domain distribution differences. Such methods usually start from the data distribution and measure the differences between domains through certain statistical rules, such as rules like maximum mean difference, covariance matrix difference, central moment difference, "earth mover's" distance, etc., and constrain the model to minimize the differences between the two domains as much as possible; the second is the method based on adversarial learning. The idea of this method comes from the generative adversarial network. The core is to train a pair of feature extractors and domain discriminators. The former attempts to extract invariant features from samples from different domains, and the latter attempts to determine which domain the features extracted by the former come from. After training is completed, the feature extractor can extract features that are both class-discriminative and domain-invariant, achieving the goal of domain adaptation; the third is the method based on reconstruction. A pair of cooperating encoders and decoders are used. The encoder is responsible for extracting domain-invariant features, and the decoder is responsible for reconstructing these features into their original forms. After training is completed, the features extracted by the encoder can be used for object detection.

[0004] In recent years, there have been more and more object detection tasks for specific scenarios, such as street scene recognition for different cities, road condition detection for different weather conditions, object detection for different imaging devices, etc. The proposal of such tasks has given great room for the development of domain adaptation object detection methods. Existing domain adaptation object detection methods usually rely on the aforementioned three methods and their combinations, but there are still problems such as missed detections and false detections, which reflects that the extraction of domain-invariant features in the source domain and the target domain by existing methods is not perfect enough, and it is necessary to further improve these methods. Summary of the Invention

[0005] In view of the situation where there are only source domain annotations but no target domain annotations in object detection tasks, the present invention proposes a domain adaptation object detection method and device based on contrastive learning.

[0006] The technical solution adopted by the present invention is as follows:

[0007] A domain adaptation object detection method based on contrastive learning, comprising the following steps:

[0008] Construct a domain adaptation object detection network composed of a feature decoupling extraction module, an object detection module, and a contrastive learning module; the feature decoupling extraction module uses ResNet-101 as the basic convolutional neural network structure, and includes a shallow feature extraction module, a shallow feature decoupling module, a deep feature extraction module, and a deep feature decoupling module;

[0009] For the input image data, the feature decoupling extraction module extracts the shallow feature map of the image through the shallow feature extraction module, and then processes the shallow feature map through the shallow feature decoupling module to obtain domain-related features and domain-unrelated features; then add the domain-unrelated features to the shallow feature map, extract the deep feature map through the deep feature extraction module, and finally process the deep feature map through the deep feature decoupling module to obtain deep domain-related features and deep domain-unrelated features; the object detection module locates and classifies the object according to the deep domain-unrelated features;

[0010] Train the domain adaptation object detection network using the labeled source domain image dataset and the unlabeled target domain image dataset; the feature decoupling extraction module calculates the domain-related features and domain-unrelated features according to the labeled source domain image dataset, and calculates the mutual information loss function between the shallow domain-related features and the shallow domain-unrelated features, the mutual information loss function between the deep domain-related features and the deep domain-unrelated features, and the reconstruction loss function of the deep feature map. Use the above two mutual information loss functions, reconstruction loss function and consistency loss function to train the shallow feature decoupling module and the deep feature decoupling module in the feature decoupling extraction module; use the classification loss function and the regression loss function to train the object detection module; the contrast learning module saves the features obtained by processing the source domain image dataset through the feature decoupling extraction module and the object detection module and the corresponding labels, as well as the features obtained by processing the target domain image dataset through the feature decoupling extraction module and the object detection module and the corresponding pseudo-labels, performs contrast learning on these two types of features, and uses the contrast loss function to optimize all modules included in the feature decoupling extraction module;

[0011] Input the target domain picture to be detected into the trained domain adaptation object detection network to locate and classify the objects in the target domain image.

[0012] Furthermore, ResNet-101 contains five layers. The first three layers form the shallow feature extraction module, and the last two layers form the deep feature extraction module; the shallow feature decoupling module consists of two branches, and each branch contains three consecutive 1*1 convolutional layers with activation layers.

[0013] Furthermore, the object detection module includes an RPN network, a regression module composed of two fully connected layers, and a classification module composed of two fully connected layers. The RPN network is responsible for screening out potential object candidate regions from the features and obtaining the features for classification and localization. The regression module is responsible for calculating the specific position of the object using the aforementioned features, and the classification module is responsible for classifying the object using the aforementioned features.

[0014] Furthermore, during the training process, the consistency loss is calculated using the deep feature map and the deep domain-unrelated features. The formula for this consistency loss is as follows:

[0015]

[0016] where L rc is the consistency loss, A b and A di are the autocorrelation matrices of the object features in the deep feature map and the deep domain-unrelated feature map respectively. The object features are intercepted from the corresponding feature map by the RPN network using the candidate regions.

[0017] Further, during the training process, the features output by the RPN network of the object detection module are combined with the corresponding labels and input into the contrast module to update the memory bank and calculate the contrast loss function.

[0018] Further, during the training process, for the unlabeled target domain image dataset, the gradients of the object detection module are locked, and the deep domain-independent features obtained by the feature decoupling extraction module are input into the object detection module to obtain prediction results. The features corresponding to the prediction results with a confidence higher than a preset value are selected, and the predicted labels are used as pseudo-labels.

[0019] Further, the contrast learning module includes a feature mapping branch composed of two fully connected layers and a memory bank. The feature mapping branch is used to reduce the dimensions of the features, labels, and pseudo-labels to be stored in the memory bank in advance; the memory bank is a sub-module for storing a certain number of feature-label pairs; when the memory bank reaches its set capacity limit, the earliest stored features are discarded and the newly stored features are saved.

[0020] Further, during each update of the parameters of each module in the feature decoupling extraction module during the training process, the contrast module calculates a contrast loss, and the formula of the contrast loss function is as follows:

[0021]

[0022]

[0023] Among them, L cont is the contrast loss, is the contrast loss for a feature i, N is the capacity of the memory bank, z i , z j , z k are the features numbered i, j, k in the memory bank, y i , y j are the labels of the features of i, j, refers to selecting the feature y i with the same label as y j , τ is the temperature parameter of the contrast loss function, and the base of the logarithm in log is the natural logarithm base e.

[0024] Further, the trained domain adaptation object detection network locates and classifies the objects in the target domain picture, including the following steps:

[0025] Input the target domain picture to be detected into the domain adaptation object detection network to obtain the classification result and the regression result;

[0026] Use the non-maximum suppression algorithm to filter the regression results obtained in the previous step, remove multiple detection boxes for the same object, and obtain the final result.

[0027] A domain adaptation object detection device based on contrast learning includes a domain adaptation object detection network, which is composed of a feature decoupling extraction module, an object detection module, and a contrast learning module. The feature decoupling extraction module uses ResNet-101 as the basic convolutional neural network structure, including a shallow feature extraction module, a shallow feature decoupling module, a deep feature extraction module, and a deep feature decoupling module; among them,

[0028] The shallow feature extraction module is responsible for extracting shallow information from the input image to obtain a shallow feature map;

[0029] The shallow feature decoupling module is responsible for decoupling the shallow feature map to obtain shallow domain-related features and domain-unrelated features;

[0030] The deep feature extraction module is responsible for combining the shallow feature map and the domain-unrelated features and extracting a deep feature map therefrom;

[0031] The deep feature decoupling module is responsible for decoupling deep domain-related features and deep domain-unrelated features for object detection from the deep feature map;

[0032] The object detection module is responsible for using the deep domain-unrelated features for object detection and outputting classification results and regression results;

[0033] The contrast learning module is responsible for assisting in aligning the features of the labeled source domain image data and the unlabeled target domain image data during the training phase, and optimizing all the modules included in the feature decoupling extraction module.

[0034] The present invention adopts the ResNet network in the deep neural network, combined with contrast learning and the feature decoupling module, which can match the features in the source domain and the target domain and decouple the domain-invariant features when only the source domain image data in the dataset is labeled and the target domain image data is not labeled, and perform object detection based on the domain-invariant features. In particular, during the training process, contrast learning is used to achieve the alignment of the source domain and target domain features, improve the consistency of the decoupled features, obtain better domain-invariant features, and enhance the object detection effect in the target domain. The results have high precision and high recall rate, indicating that the network has strong domain adaptation ability. Tests show that the present invention has been tested on the dataset pairs commonly used in the academic community to evaluate domain adaptation object detection, that is, using the cityscape dataset as the source domain and the foggy-cityscape dataset without labels as the target domain, and good results have been obtained. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] The present invention will be further described in detail below through specific embodiments and drawings.

[0036] Figure 1It is an architecture diagram of a domain adaptation object detection network according to an embodiment of the present invention;

[0037] Figure 2 It is a data flow diagram of the feature decoupling extraction module according to an embodiment of the present invention;

[0038] Figure 3 It is an example diagram of the domain adaptation object detection result according to an embodiment of the present invention. Detailed implementation manners

[0039] To make the above features and advantages of the present invention more obvious and understandable, specific embodiments are hereinafter given and described in detail in conjunction with the accompanying drawings as follows.

[0040] The domain adaptation object detection method based on contrast learning of the present invention is mainly divided into a training stage and a testing stage, which are specifically described as follows.

[0041] I. The steps of the training stage are as follows:

[0042] 1) Use ResNet-101 as the basic feature extractor, construct a feature extraction decoupling module, and then cascade an object detection module and a contrast learning module to construct a domain adaptation object detection network.

[0043] The structure of the domain adaptation object detection network is as Figure 1 shown. Among them, the feature extraction decoupling module is taken as a whole, and its construction is as follows: Use ResNet-101 as the basic feature extractor, Figure 2 where conv1, conv2, conv3, conv4, and conv5 are the five layers of ResNet-101, which are constructed by multiple residual blocks. These layers are separated from the middle. conv1, conv2, and conv3 are responsible for extracting the shallow information of the image to form a shallow feature extraction module; conv4 and conv5 form a deep feature extraction module. For the shallow feature map obtained by the shallow feature extraction module, it is sent to the shallow feature decoupling module for processing. Specifically, the shallow feature decoupling module consists of two branches, and each branch contains three consecutive 1*1 convolutional layers with activation layers. One of these two branches is used to output domain-related features, and the other outputs domain-unrelated features. Adding the domain-unrelated features obtained by the shallow feature decoupling module to the shallow feature map can obtain the shallow feature with enhanced domain-unrelated features. Inputting the above shallow features into the deep feature extraction module can obtain the deep feature map. Inputting the deep features into the deep feature decoupling module, the output domain-related features and deep domain-unrelated features are decoupled. The structure of the deep feature decoupling module is the same as that of the shallow feature decoupling module, but the parameters are different.

[0044] Inputting the above deep domain - independent features into the object detection module can obtain object localization and classification results. Specifically, the object detection module contains three sub - modules. One is the RPN network, which is responsible for screening out potential object candidate options from the features and obtaining features that can be used for classification and localization. The second is a regression module composed of two fully - connected layers, which is responsible for calculating the specific position of the object using the aforementioned features. The third is a classification module composed of two fully - connected layers, which is responsible for classifying the object using the aforementioned features.

[0045] The contrastive learning module only participates in the work during the training phase, and its structure is as Figure 1 shown. Specifically, the contrastive learning module contains a mapping branch composed of a two - layer fully - connected network and regularization, and a memory queue that can store "feature - label" pairs.

[0046] 2) Input the training data into the network, calculate the loss function in sequence, and adjust the network parameters.

[0047] In this step, the network parameters of multiple modules are adjusted through training. Specifically, first, when the labeled source - domain image data used as training data passes through the shallow - feature extraction module and the shallow - feature decoupling module in sequence, the mutual - information loss function is calculated using the output domain - related features and domain - independent features, so as to optimize the parameters of the shallow - feature decoupling module. Then, the domain - independent features are added to the existing shallow features and sent into the deep - feature extraction module and the deep - feature decoupling module in sequence. The mutual - information loss function is calculated again using the output deep domain - related features and deep domain - independent features, and combined with the deep - feature map to calculate the reconstruction loss function, so as to optimize the parameters of the deep - feature decoupling module. Then, the deep domain - independent feature information is sent into the object detection module. For the labeled source - domain image data, the object detection module is trained according to the labels, the regression loss and classification loss are calculated, and the corresponding module parameters are optimized. For the unlabeled target - domain image data, the parameters are fixed, and the localization and classification results output by the network are obtained according to the test process, and those with a high enough confidence are retained.

[0048] In the training process, the features output by the RPN network used in the previous step are sent into the memory bank of the contrastive module together with the corresponding labels, the memory bank is updated, and the contrastive loss function is calculated based on these features and labels, so as to optimize the parameters of all modules except the object detection module. The contrastive learning module is improved based on the contrastive learning method used in other fields, as Figure 1 shown. This module only works during the training phase, calculates the contrastive loss function, and is responsible for optimizing all feature - decoupling modules and feature - extraction modules.

[0049] During the training process, the consistency loss is calculated using the deep - feature map and the deep domain - independent features. The formula for this consistency loss is as follows:

[0050]

[0051] Among them, L rc is the consistency loss, A b and A di are the autocorrelation matrices of the object features in the deep feature map and the deep domain-invariant feature map respectively. The aforementioned object features are intercepted from the corresponding feature map by the RPN network using the candidate regions obtained in the previous steps.

[0052] The input of the contrastive learning module is the features and labels output by the RPN network. For the labeled source domain image dataset, the existing labeled labels are applied and stored in the memory bank; for the unlabeled target domain image dataset, the object detection module is set to the test mode (i.e., locking the gradients of the object detection module and keeping the parameters unchanged), the deep domain-invariant features are input into the object detection module to obtain the results predicted by the network, and the results with higher confidence are selected as labels and stored in the memory bank. Next, the above features are sequentially input into two fully connected layers in the contrastive module and regularized once to obtain the features after dimensionality reduction. In this example, the dimension of the features output by the RPN network is 2048 dimensions, and the features after dimensionality reduction are 128 dimensions. Next, the features obtained after dimensionality reduction and the corresponding labels obtained from the above process are stored in the memory bank. The features stored in the memory bank of the contrastive learning module are used for contrastive learning to align the features from the source domain and the target domain and improve the consistency between the above features.

[0053] The memory bank is a sub-module that can store a certain number of "feature-label" pairs. When the memory bank reaches its set capacity limit, the memory bank will discard the earliest stored features and save the newly stored features. In the example of this domain adaptation object detection method, the size of the memory bank is set to 4000 groups of feature pairs.

[0054] During each parameter update in the training process, the contrastive module calculates a contrastive loss, and the definition of this contrastive loss function is as follows:

[0055]

[0056]

[0057] Among them, N is the capacity of the memory bank, z i is the feature numbered i in the memory bank, y i is the label of this feature, refers to selecting the features with the same label as y i The τ refers to the temperature parameter of the contrastive loss function, and here 0.2 is selected, and the base of the log is the natural logarithm base e.

[0058] II. The steps in the test phase are as follows:

[0059] 1) Input the test images in the target domain into the trained domain adaptation object detection network. The detection results of the network are multiple possibly overlapping object bounding boxes with confidence scores and their labels.

[0060] 2) Select the results with a sufficiently high confidence score from the above results, and on this basis, use the local non-maximum suppression algorithm to remove the duplicate results for the same object. The confidence score used in this step 2) is 0.5. Figure 3 It is an example diagram of the domain adaptation object detection results.

[0061] The experimental tests are as follows:

[0062] (1) Test environment:

[0063] System environment: ubuntu18.04;

[0064] Hardware environment: Memory: 64GB, GPU: Nvidia TITAN XP, Hard disk: 2TB;

[0065] (2) Experimental data:

[0066] Training data:

[0067] ImageNet pre-trained ResNet-101 basic network.

[0068] Cityscapes dataset as the source domain (including 2975 labeled training images and 500 validation images, and 1575 unlabeled test images); Foggy-cityscapes dataset as the target domain (a foggy version synthetically generated based on the cityscapes dataset with the same number of images as the former).

[0069] Use the training set part of the above dataset for training until the model is stable. In particular, although the images in the source domain and the target domain are in one-to-one correspondence, this correspondence is not utilized during the training process.

[0070] Training optimization method: SGD

[0071] Test data: Foggy-cityscapes validation set (500 images)

[0072] Evaluation method: VOC object detection standard

[0073] (3) Experimental results:

[0074] To illustrate the effect of the present invention, the domain adaptation object detection network of the present invention with or without the contrastive learning module was trained using the same data set. The training was stopped when the stable effect of the model no longer improved. The Foggy-cityscapes validation set was used for testing, and the results were compared with those of existing mainstream domain adaptation object detection methods.

[0075] The test comparison results between the present invention and the existing mainstream prediction methods SCDA (see Zhu, Xinge and Pang, Jiangmiao and Yang, Ceyuan and Shi, Jianping and Lin, Dahua, “Adapting object detectors via selective cross-domain alignment” in CVPR 2019, pp. 687-696.) and ATF (see He, Zhenwei and Zhang, Lei, “Domain adaptive object detection via asymmetric tri-way faster-rcnn” in ECCV 2020, pp. 309-324.) are shown in Table 1 below, where mAP refers to the mean average precision.

[0076] Table 1. Comparison of test results between the present invention and existing prediction methods

[0077] Serial number Method mAP (%) 1 SCDA 33.8 2 ATF 38.7 3 The present invention (without using a contrastive learning module) 38.6 4 The present invention (using a contrastive learning module) 39.8

[0078] As can be clearly seen from Table 1, the domain adaptation object detection network of the present invention has a significant improvement in accuracy compared with the existing text detection methods SCDA and ATF, and the network model obtained by the method of adding the contrastive learning module for training has been further improved in accuracy.

[0079] Although the present invention has been disclosed above by way of examples, it is not intended to limit the present invention. Any appropriate modification or equivalent replacement of the technical solutions of the present invention by those of ordinary skill in the art shall be covered by the protection scope of the present invention. The protection scope of the present invention shall be determined by the claims.

Claims

1. A domain adaptation object detection method based on contrastive learning, characterized in that, it includes the following steps: Construct a domain adaptation object detection network composed of a feature decoupling extraction module, an object detection module, and a contrastive learning module; the feature decoupling extraction module uses ResNet-101 as the basic convolutional neural network structure, including a shallow feature extraction module, a shallow feature decoupling module, a deep feature extraction module, and a deep feature decoupling module; For the input image data, the feature decoupling extraction module extracts the shallow feature map of the image through the shallow feature extraction module, and then processes the shallow feature map through the shallow feature decoupling module to obtain domain-related features and domain-unrelated features; then add the domain-unrelated features to the shallow feature map, extract the deep feature map through the deep feature extraction module, and finally process the deep feature map through the deep feature decoupling module to obtain deep domain-related features and deep domain-unrelated features; The object detection module locates and classifies objects according to the deep domain-unrelated features; Use the labeled source domain image dataset and the unlabeled target domain image dataset to train the domain adaptation object detection network; The feature decoupling extraction module calculates the domain-related features and domain-unrelated features according to the labeled source domain image dataset, and calculates the mutual information loss function of the shallow domain-related features and shallow domain-unrelated features, the mutual information loss function of the deep domain-related features and deep domain-unrelated features, and the reconstruction loss function of the deep feature map. Use the above two mutual information loss functions, reconstruction loss function and consistency loss function to train the shallow feature decoupling module and deep feature decoupling module in the feature decoupling extraction module; use the classification loss function and regression loss function to train the object detection module; the contrastive learning module saves the features and corresponding labels obtained by processing the source domain image dataset through the feature decoupling extraction module and the object detection module, as well as the features and corresponding pseudo-labels obtained by processing the target domain image dataset through the feature decoupling extraction module and the object detection module, and performs contrastive learning on these two types of features, and uses the contrastive loss function to optimize all modules included in the feature decoupling extraction module; Input the target domain picture to be detected into the trained domain adaptation object detection network to locate and classify the objects in the target domain image.

2. The method according to claim 1, characterized in that, ResNet-101 contains five layers, the first three layers form the shallow feature extraction module, and the last two layers form the deep feature extraction module; the shallow feature decoupling module is composed of two branches, and each branch contains three consecutive 1*1 convolutional layers with activation layers.

3. The method according to claim 1, characterized in that, The object detection module includes an RPN network, a regression module composed of two fully connected layers, and a classification module composed of two fully connected layers. The RPN network is responsible for screening out potential object candidate regions from the features and obtaining features for classification and localization. The regression module is responsible for calculating the specific position of the object using the aforementioned features, and the classification module is responsible for classifying the object using the aforementioned features.

4. The method according to claim 3, characterized in that, During the training process, a consistency loss is calculated using deep feature maps and deep domain-invariant features. The formula for this consistency loss is as follows: Among them, L rc is the consistency loss, A b and A di are the autocorrelation matrices of the object features in the deep feature map and the deep domain-invariant feature map respectively, and the object features are intercepted from the corresponding feature map by the RPN network using the alternative regions.

5. The method according to claim 3, wherein, During the training process, the features output by the RPN network of the object detection module are combined with the corresponding labels and input into the contrast learning module to update the memory bank and calculate the contrast loss function.

6. The method according to claim 1, wherein, During the training process, for the unlabeled target domain image dataset, the gradients of the object detection module are locked, and the deep domain-invariant features obtained by the feature decoupling extraction module are input into the object detection module to obtain prediction results. The features corresponding to the prediction results with a confidence level higher than a preset value are selected, and the predicted labels are used as pseudo-labels.

7. The method according to claim 1, wherein, The contrast learning module includes a feature mapping branch composed of two fully connected layers and a memory bank. The feature mapping branch is used to reduce the dimension of the features, labels, and pseudo-labels to be stored in the memory bank first; the memory bank is a sub-module for storing a certain number of feature-label pairs; when the memory bank reaches its set capacity limit, the earliest stored features are discarded and the newly stored features are saved.

8. The method according to claim 1, wherein, During each update of the parameters of each module in the feature decoupling extraction module during the training process, the contrast learning module calculates a contrast loss. The formula for this contrast loss function is as follows: Among them, L cont is the contrastive loss, is the contrastive loss for a feature i, N is the capacity of the memory bank, z i , z j , z k are the features numbered i, j, k in the memory bank, y i , y j are the labels of the features of i, j, refers to selecting the feature y i with the same label as y j , τ is the temperature parameter of the contrastive loss function, and the base of log is the natural logarithm base e.

9. The method according to claim 1, wherein, The trained domain adaptation object detection network locates and classifies objects in the target domain picture, including the following steps: Input the target domain picture to be detected into the domain adaptation object detection network to obtain classification results and regression results; Use the local maximum suppression algorithm to filter the regression results obtained in the previous step, remove multiple detection boxes for the same object, and obtain the final results.

10. A domain adaptation object detection device based on contrast learning, wherein, It includes a domain adaptation object detection network, which is composed of a feature decoupling extraction module, an object detection module, and a contrast learning module; the feature decoupling extraction module uses ResNet-101 as the basic convolutional neural network structure, including a shallow feature extraction module, a shallow feature decoupling module, a deep feature extraction module, and a deep feature decoupling module; wherein, The shallow feature extraction module is responsible for extracting shallow information from the input image to obtain a shallow feature map; The shallow feature decoupling module is responsible for decoupling the shallow feature map to obtain shallow domain-related features and domain-invariant features; The deep feature extraction module is responsible for combining the shallow feature map and the domain-invariant features and extracting deep feature maps therefrom; The deep feature decoupling module is responsible for decoupling deep domain-related features and deep domain-invariant features for object detection from the deep feature map; The object detection module is responsible for using the deep domain-invariant features for object detection and outputting classification results and regression results; The contrastive learning module is responsible for assisting in aligning the features of the labeled source domain image data and the unlabeled target domain image data during the training phase, and optimizing all the modules included in the feature decoupling extraction module.

Citation Information

Patent Citations

  • Cross-domain target detection method based on multi-layer feature alignment

    CN110363122A

  • Method and device for image processing, and computer storage medium

    US20210034913A1