Domain adaptive DINO model target detection method based on feature fusion and alignment

By adopting feature fusion and alignment technology in the DINO model, the problem of limited performance of existing cross-domain detection methods is solved, and higher generalization and discrimination capabilities are achieved in cross-domain detection tasks.

CN120147761AActive Publication Date: 2025-06-1310TH RES INST OF CETC

Patent Information

Application Number
CN202510618107.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-06-13
Estimated Expiration
2045-05-14

AI Technical Summary

Technical Problem

The existing cross-domain detection method is based on the early detector architecture, with limited performance of native detectors, and the existing cross-domain adaptation strategy is incompatible with the DINO model detector, limiting the performance of the model in cross-domain detection tasks.

Method used

The domain adaptive DINO model object detection method based on feature fusion and alignment is adopted to improve feature representation capabilities and reduce the difference in feature distribution in different fields by differentiating feature fusion and image-level, pixel-level, and target-level feature alignment.

Benefits of technology

It significantly improves the performance of the DINO model in cross-domain detection tasks, enhances the cross-domain generalization and discrimination capabilities of the model, and reduces the dependence on labeled samples and the demand for computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147761A_ABST
    Figure CN120147761A_ABST
Patent Text Reader

Abstract

The invention discloses a domain adaptive DINO model target detection method based on feature fusion and alignment, and the method comprises the steps: sequentially inputting features outputted by a backbone network into a first feature fusion module and a first domain discriminator, and obtaining the image-level adversarial training loss of the backbone network; inputting the features output by the encoder into a second field discriminator to obtain encoder pixel-level adversarial training loss; sequentially inputting the features output by the decoder into a second feature fusion module and a third domain discriminator to obtain decoder target-level adversarial training loss; according to the backbone network image level adversarial training loss, the encoder pixel level adversarial training loss, the decoder target level adversarial training loss and the source domain detection loss, the DINO model learns cross-domain invariant features, and cross-domain target detection of target domain image data is completed. According to the method, the data distribution offset can be reduced, and the transferable feature representation is learned, so that the model also has good detection performance in the target domain.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of cross-domain object detection, and particularly to an object detection method for a domain adaptive DINO model based on feature fusion and alignment. Background Art

[0002] Object detection is one of the basic tasks in the field of computer vision, showing great application potential in fields such as autonomous driving, video surveillance, and industrial quality inspection. Early detection methods based on Convolutional Neural Networks (CNNs) achieved efficient object localization and classification through the anchor box mechanism and region proposal network. However, these methods rely on complex post-processing procedures, such as non-maximum suppression, and the model design is highly coupled, limiting flexibility and scalability in different scenarios. Detection Transformer (DETR) was proposed in 2020, introducing the Transformer architecture into the object detection task for the first time, modeling the object detection task as a set prediction problem, and eliminating the dependence on anchor boxes and post-processing of previous methods in an end-to-end manner. DETR made bold innovations in architecture and process, but problems such as slow convergence speed and insufficient detection performance restricted its practical applications. In response to the defects of DETR, researchers have made a series of improvements, such as Conditional DETR and Deformable DETR. Among them, the DETR with Improved deNOising anchor boxes (DINO model) significantly improved the convergence speed and detection accuracy through improvements such as contrastive denoising training, hybrid query selection, and gradient transfer optimization, making the DINO model the detector architecture with the best performance in the field of object detection.

[0003] The training of object detection models depends on large-scale high-quality labeled data, and the acquisition of such data is often time-consuming and laborious. Compared with detectors based on the CNN architecture, detectors based on the Transformer architecture, such as the DINO model, are more dependent on data. In addition, the generalization performance of detectors obtained through supervised training is also limited. When facing object data with domain shifts (such as different object shapes, complex backgrounds, and lighting changes), the detection performance of the model will drop sharply. To alleviate the limitation of domain shifts on the generalization performance of the model and avoid cumbersome data annotation work, more and more cross-domain learning methods have been proposed. These methods are based on domain adaptation techniques, using the knowledge of existing data and the correlation between existing data and target data to improve the generalization ability of the model to target data and reduce the dependence on labeled samples and the demand for computing resources in model training.

[0004] Object detection is more complex than image segmentation or classification tasks because a single input sample may correspond to multiple objects to be detected, and object localization and classification need to be completed simultaneously in a single model, which poses a huge challenge to cross-domain detection research but also inspires more ideas. Domain Adaptive Faster-RCNN is the first cross-domain detection method, which improves the model's cross-domain detection ability through image-level and instance-level adaptation. Based on this, cross-domain object detection research has become one of the hot issues in the field of object detection. However, existing cross-domain detection methods are based on early detector architectures such as Faster RCNN and YOLO, and the performance of their native detectors limits the cross-domain performance of the model, and cross-domain training methods based on the CNN architecture are also difficult to be directly applied to the Transformer architecture. Summary of the Invention

[0005] Aiming at the problems of limited performance of the native detector of the existing cross-domain detection model and the incompatibility of the existing cross-domain adaptation strategy with the DINO model detector, this application provides a domain adaptive DINO model object detection method based on feature fusion and alignment, which improves the feature representation ability through the method of differential feature fusion, and performs image-level, pixel-level and object-level feature alignment at different stages of the network, fully reducing the difference in feature distributions in different domains and improving the performance of the DINO model in cross-domain detection tasks.

[0006] This application discloses a domain adaptive DINO model object detection method based on feature fusion and alignment, which includes: Step 1: Input the source domain data and the target domain data into the backbone network of the DINO model for feature extraction; the network model includes a DINO model, a first feature fusion module, a first domain discriminator, a second domain discriminator, a second feature fusion module and a third domain discriminator; the DINO model includes a backbone network, an encoder, a decoder and a feed-forward network connected in sequence; the source domain data includes the source domain image and its corresponding target label, and the target label includes the bounding box coordinates and category information of the target; the target domain data only includes the target domain image; Step 2: Input the features output by the backbone network into the first feature fusion module and the first domain discriminator in sequence to obtain the backbone network image-level adversarial training loss; Step 3: Input the features output by the encoder into the second domain discriminator to obtain the encoder pixel-level adversarial training loss; Step 4: Input the features output by the decoder into the second feature fusion module and the third domain discriminator in sequence to obtain the decoder object-level adversarial training loss; Step 5: Optimize the network model according to the backbone network image-level adversarial training loss, the encoder pixel-level adversarial training loss, the decoder target-level adversarial training loss, and the source domain detection loss, so that the network model learns cross-domain invariant features and enhances the cross-domain generalization ability of the network model; Step 6: After the network model is trained, only load the DINO model weights to complete the cross-domain object detection of the target domain image data.

[0007] Further, the said Step 2 includes: Step 21: The first feature fusion module obtains image-level fusion features according to the multi-scale features extracted by the backbone network; Step 22: The first domain discriminator obtains an image-level domain prediction result according to the image-level fusion features; Step 23: Obtain the backbone network image-level adversarial training loss according to the image-level domain prediction result.

[0008] Further, the said Step 21 includes: The different scale features in the multi-scale features extracted by the backbone network are learned through the convolutional layer in the first feature fusion module, then the size of the feature map is unified through the global average pooling layer, and finally the image-level fusion features are obtained through the Concat operation layer; The first feature fusion module consists of a convolutional layer, a global average pooling layer and a fully connected layer; The said Step 22 includes: The first feature fusion module sequentially inputs the image-level fusion features into the gradient reversal layer and the fully connected layer in the first domain discriminator, and outputs the image-level domain prediction result; The first domain discriminator consists of a gradient reversal layer and a fully connected layer.

[0009] Further, the said Step 23 includes: Obtain the backbone network image-level adversarial training loss through the following formula: (1) where, is the backbone network image-level adversarial training loss, is the image-level domain prediction result of the sample and is the domain label of the sample . The domain label includes the source domain label and the target domain label, and their values are 0 and 1 respectively.

[0010] Further, the said Step 3 includes: Step 31: The encoder models the global context relationship through the self-attention mechanism. Its input is obtained by adding the one-dimensional sequence features converted from the multi-scale features output by the encoder and the position embedding. The input of each subsequent layer of the encoder is the output of the previous layer;​ Step 32: Input the output of the encoder into the gradient reversal layer and the fully connected layer in the second domain discriminator in sequence to obtain the pixel-level domain prediction result , is the output of the encoder, is the second domain discriminator; the structure of the second domain discriminator is the same as that of the first domain discriminator, and it consists of a gradient reversal layer and a fully connected layer; Step 33: Obtain the pixel-level adversarial training loss of the encoder according to the pixel-level domain prediction result.

[0011] Further, the said Step 33 includes: Obtain the pixel-level adversarial training loss of the encoder through the following formula: (2) wherein, is the pixel-level adversarial training loss of the encoder, represents the domain prediction result corresponding to the j-th pixel in the multi-scale feature map of the sample , is the domain label of the sample , and the domain label includes the source domain label and the target domain label, and their values are 0 and 1 respectively.

[0012] Further, the said Step 4 includes: Step 41: Input the output of the encoder into the decoder to optimize the query anchor box and the classification result. The input of each subsequent layer of the decoder is the output of the previous layer, and the output of each layer is denoted as , wherein , and , is the number of target queries, is the number of layers of the decoder, is the total number of layers of the decoder, C is the number of feature map channels, or the number of hidden layer channels, that is, the feature dimension of the sequence feature; Step 42: The feature output by the decoder is learned through a fully connected layer and an activation function, and then the target-level fusion feature is obtained through the Concat operation layer; the second feature fusion module consists of a fully connected layer and a Concat operation layer; Step 43: Input the target-level fusion feature into the gradient reversal layer and the fully connected layer in the third domain discriminator in sequence to obtain the domain prediction result , is the third domain discriminator; the structure of the third domain discriminator is the same as that of the first domain discriminator, and it consists of a gradient reversal layer and a fully connected layer; Step 44: Obtain the decoder target-level adversarial training loss according to the domain prediction result.

[0013] Further, the step 44 includes: Obtain the decoder target-level adversarial training loss through the following formula: (3) where, is the decoder target-level adversarial training loss, represents the domain prediction result corresponding to the j-th target query of the sample , is the domain label of the sample . The domain label includes the source domain label and the target domain label, and their values are 0 and 1 respectively, or their values are 1 and 0 respectively.

[0014] Further, the step 5 includes: Step 51: Obtain the adversarial training loss of the network model according to the backbone network image-level adversarial training loss, the encoder pixel-level adversarial training loss, and the decoder target-level adversarial training loss; Step 52: Obtain the optimal network model according to the adversarial training loss of the network model and the source domain detection loss.

[0015] Further, the step 51 includes: Obtain the adversarial training loss of the network model through the following formula: (4) where, is the adversarial training loss of the network model, is the backbone network image-level adversarial training loss, is the encoder pixel-level adversarial training loss, is the decoder target-level adversarial training loss; The step 52 includes: Complete learning by jointly optimizing the adversarial training loss and the source domain detection loss through the following formula: (5) where, represents the set of the first domain discriminator , the second domain discriminator and the third domain discriminator , represents the remaining part of the network model except , is the source domain detection loss, is a trade-off parameter used to balance the adversarial training loss and the source domain detection loss; the remaining part of the network includes a backbone network, an encoder, a decoder, a feed-forward network, a first feature fusion module, and a second feature fusion module.

[0016] Due to the adoption of the above technical solutions, the present application has the following advantages: (1) Design a feature fusion module and a global feature alignment strategy in the backbone network part of the DINO model, fuse the multi-scale features of the backbone network, and perform global coarse-grained alignment on the fused features to reduce the influence of factors such as lighting conditions and complex backgrounds, and improve the overall adaptability of the model to the target scene.

[0017] (2) Propose a local feature alignment strategy in the encoder part of the DINO model to perform local fine-grained alignment on the encoder features, alleviate the influence of factors such as scene layout on the cross-domain performance of the model, and reduce the pixel-level feature distribution difference.

[0018] (3) Propose a feature fusion module and a target-level alignment strategy in the decoder part of the DINO model to fuse the hierarchical features of the decoder and perform target-level feature alignment on the fused features to avoid the negative transfer phenomenon and ensure the cross-domain detection generalization and discriminative ability of the model.

[0019] (4) After the DINO model is trained, only the weights of the DINO model need to be loaded when performing the object detection task, which can improve the performance of the model in the cross-domain object detection task without increasing the computational amount. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments recorded in the embodiments of the present application, and those of ordinary skill in the art can also obtain other drawings based on these drawings.

[0021] Figure 1 is a schematic structural diagram corresponding to a domain adaptive DINO model object detection method based on feature fusion and alignment according to an embodiment of the present application; Figure 2 is a schematic structural diagram of the first feature fusion module according to an embodiment of the present application; Figure 3 is a schematic structural diagram of the first domain discriminator, or the second domain discriminator, or the third domain discriminator according to an embodiment of the present application; Figure 4 is a schematic structural diagram of the second feature fusion module according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] The present application will be further described in conjunction with the accompanying drawings and embodiments. The described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art shall fall within the scope of protection of the embodiments of the present application.

[0023] See Figure 1 , the present application provides an embodiment of a domain adaptive DINO model object detection method based on feature fusion and alignment. Based on the currently best-performing DINO model, it improves the feature representation ability through the method of differential feature fusion, and performs image-level, pixel-level, and object-level feature alignment at different stages of the network. The feature fusion and alignment strategies only play a role during the model training process. The trained model only needs to load the DINO model weights to perform cross-domain inference, which can improve the cross-domain detection performance of the model while not increasing the additional inference burden.

[0024] For the convenience of description, relevant knowledge of cross-domain object detection is introduced first: the source domain with labeled information , where represents the number of source domain samples, represents the source domain image, is the corresponding target label, including the bounding box coordinates and class information of the target. The target domain without labeled information , where represents the number of target domain samples, represents the target domain image. The label spaces of the source domain and the target domain are the same, but their data distributions are different. The purpose of the present application is to train a network model to reduce the data distribution shift and learn transferable feature representations, so that the model also has good detection performance in the target domain.

[0025] Secondly, the DINO model detection process is briefly introduced: The DINO model is an end-to-end object detection model based on the DETR architecture. As Figure 1 shown, specifically, it includes the following parts: (1) Backbone network: Generally, CNN or vision Transformer is used to extract the multi-scale feature representation of the input image , is the number of feature scales, and there is . After that, through operations such as flattening, embedding, and splicing, the multi-scale features are transformed into one-dimensional sequence features , the sequence length ; H and W respectively represent the height and width of the image; (2) Transformer encoder: It models the global context relationship through the self-attention mechanism. Its input is obtained by adding the above one-dimensional sequence feature plus position embedding and other information, and is expressed as , the input of each subsequent layer of the encoder is the output of the previous layer, and the final output of the encoder is denoted as ; (3) Transformer decoder: Continuously optimize the query anchor boxes and classification results. The input of each subsequent layer of the decoder is also the output of the previous layer, and the output of each layer is denoted as , where , is the total number of layers of the encoder, and there is , is the number of target queries; (4) The feed-forward network predicts the target class probabilities and bounding box positions through the outputs of all decoder layers; (5) Model optimization: The DINO model uses cross-entropy loss to predict the target class, and uses L1 loss and GIOU loss for bounding box regression. Combined, it is the source domain detection loss during the training of the DINO model, denoted as .

[0026] This application takes the DINO model as the main body and adds a backbone network feature fusion module (the first feature fusion module) and a global feature alignment strategy (corresponding to the first domain discriminator). Design the feature fusion module to ensure the model's detection ability for various targets in the target domain data, and further design the global feature alignment strategy to reduce the difference in image-level feature distributions, suppress domain shifts caused by factors such as image style, lighting conditions, and complex backgrounds, and improve the model's overall adaptability to the target scene; add an encoder local feature alignment strategy. Design the local feature alignment strategy (corresponding to the second domain discriminator) to perform local fine-grained alignment on the encoder features, alleviate the impact of factors such as scene layout on the model's cross-domain performance, and reduce the difference in pixel-level feature distributions; add a decoder feature fusion module (the second feature fusion module) and a target-level feature alignment strategy (corresponding to the third domain discriminator). Design the decoder feature fusion module to learn the rich target semantic information of the sequence features, refine the features to avoid negative transfer, and further design the target-level alignment strategy to ensure the model's cross-domain detection generalization and discrimination ability.

[0027] The embodiments of this application include the following technical solutions: Step 1: Input the source domain data and the target domain data into the backbone network of the DINO model for feature extraction; the network model includes the DINO model, the first feature fusion module, the first domain discriminator, the second domain discriminator, the second feature fusion module, and the third domain discriminator; the DINO model includes a backbone network, an encoder, a decoder, and a feed-forward network connected in sequence; the source domain data includes the source domain image and its corresponding target label, and the target label includes the bounding box coordinates and category information of the target; the target domain data only includes the target domain image; Step 2: Input the features output by the backbone network into the first feature fusion module and the first domain discriminator in sequence to obtain the backbone network image-level adversarial training loss; Step 3: Input the features output by the encoder into the second domain discriminator to obtain the encoder pixel-level adversarial training loss; Step 4: Input the features output by the decoder into the second feature fusion module and the third domain discriminator in sequence to obtain the decoder target-level adversarial training loss; Step 5: Optimize the network model according to the backbone network image-level adversarial training loss, the encoder pixel-level adversarial training loss, the decoder target-level adversarial training loss, and the source domain detection loss, so that the network model learns cross-domain invariant features and enhances the cross-domain generalization ability of the network model; Step 6: After the network model is trained, simply load the DINO model weights to complete the cross-domain object detection of the target domain image data.

[0028] This application is applicable to the situation where the target domain data lacks annotation information or there is a distribution shift between the target domain data and the training data. It can utilize the existing data knowledge and the correlation between the existing data and the target data to improve the discrimination and generalization ability of the model for the target data, and reduce the dependence of model training on labeled samples and the demand for computing resources.

[0029] Optionally, step 2 includes: Step 21: The first feature fusion module obtains the image-level fusion features according to the multi-scale features extracted by the backbone network; Step 22: The first domain discriminator obtains the image-level domain prediction result according to the image-level fusion features; Step 23: Obtain the backbone network image-level adversarial training loss according to the image-level domain prediction result.

[0030] Optionally, step 21 includes: The different-scale features in the multi-scale features extracted by the backbone network are learned through the convolutional layer in the first feature fusion module, then the size of the feature map is unified through the global average pooling layer, and finally the image-level fusion features are obtained through the Concat operation layer; the first feature fusion module consists of a convolutional layer, a global average pooling layer, and a fully connected layer; Specifically, the feature quality of the backbone network is the cornerstone of the detection performance of the entire model. It can effectively guide the learning process of the Transformer encoder and decoder. Learning the backbone network features invariant to the domain is a basic element of the model's generalization ability. Considering the design of multi-scale feature output in the backbone network of the DINO model, this application designs a feature fusion module and a global feature alignment strategy to ensure the detection ability of the model for various targets in the target domain data. The structure of the feature fusion module of the backbone network is as Figure 2As shown in the figure, fusing the multi-scale features of the backbone network can enhance the representational ability of the features, extract the image-level fusion features, and adapt to the subsequent global feature alignment strategy. The feature fusion module of the backbone network is composed of a convolutional layer, global average pooling, and Concat operation as a whole. The features of different scales are learned through a shared convolutional layer, and then the size of the feature map is unified through global average pooling operation , and finally the image-level fusion features are obtained through the Concat operation . Due to the reduction in the size and quantity of the feature maps, this fusion module also effectively reduces the computational complexity in the feature alignment process.

[0031] Step 22 includes: The first feature fusion module sequentially inputs the image-level fusion features into the gradient reversal layer and the fully connected layer in the first domain discriminator, and outputs the image-level domain prediction result; the first domain discriminator is composed of a gradient reversal layer and a fully connected layer.

[0032] Specifically, the feature alignment strategy refers to reducing the difference in feature distributions through adversarial training to improve the generalization of the model. In cross-domain adversarial training, the network usually includes a feature extractor (referring to the DINO model) and a domain discriminator. The domain discriminator is used to predict whether the input features belong to the source domain or the target domain. The feature extractor attempts to confuse the judgment of the domain discriminator and perform supervised learning on the source domain data. The model implicitly learns domain-invariant features during the confrontation between the two. The structure of the domain discriminator is as Figure 3 shown, and it is composed of a gradient reversal layer and a fully connected layer. The domain discriminator of the backbone network is denoted as , with the input being the backbone network fusion feature map , and the output being the image-level domain prediction result .

[0033] Optionally, step 23 includes: Obtain the backbone network image-level adversarial training loss through the following formula: (1) where is the backbone network image-level adversarial training loss, is the image-level domain prediction result of sample , is the domain label of sample . The domain label includes the source domain label and the target domain label, and their values are 0 and 1 respectively. Optionally, let the source domain label be 0 and the target domain label be 1, or let the source domain label be 1 and the target domain label be 0.

[0034] The above optimization can effectively reduce the difference in the feature distribution of the backbone network, but still cannot guarantee the local distribution alignment and the transferability of the Transformer serialized features. Therefore, aligning the Transformer serialized features is necessary to improve the cross-domain performance of the model. The Transformer encoder is a bridge connecting the backbone network and the decoder. To make up for the deficiency of the global alignment strategy of the backbone network, this application further designs an encoder local feature alignment strategy to alleviate the impact of factors such as scene layout on the cross-domain performance of the model and reduce the pixel-level feature distribution difference. Optionally, step 3 includes: Step 31: The encoder models the global context relationship through the self-attention mechanism. Its input is obtained by adding the one-dimensional sequence features transformed from the multi-scale features output by the encoder and the position embedding. The input of each subsequent layer of the encoder is the output of the previous layer; Step 32: The structure of the second domain discriminator used for encoder local feature alignment is also as Figure 3 shown. The output of the encoder is sequentially input into the gradient reversal layer and the fully connected layer in the second domain discriminator to obtain the pixel-level domain prediction result , is the output of the encoder, is the second domain discriminator; the structure of the second domain discriminator is the same as that of the first domain discriminator, and it consists of a gradient reversal layer and a fully connected layer; Step 33: According to the pixel-level domain prediction result, the encoder pixel-level adversarial training loss is obtained.

[0035] Optionally, step 33 includes: The encoder pixel-level adversarial training loss is obtained through the following formula: (2) where, is the encoder pixel-level adversarial training loss, represents the domain prediction result corresponding to the j-th pixel in the multi-scale feature map of sample , is the domain label of sample . The domain label includes the source domain label and the target domain label, and their values are 0 and 1 respectively, or their values are 1 and 0 respectively.

[0036] The prediction result of the DINO model depends on the output of each decoder layer of the Transformer. Since there is a lack of supervision information in the target domain, the decoder features will inevitably be biased towards the source domain data, thereby damaging the discrimination and generalization performance of the model on the target data. Therefore, it is necessary to perform cross-domain modeling on the decoder feature distribution. In order to avoid the negative transfer phenomenon and ensure the domain invariance of the model output information, the present application designs a decoder feature fusion module and a target-level feature alignment strategy.

[0037] Optionally, step 4 includes: Step 41: Input the output of the encoder into the decoder to optimize the query anchor boxes and classification results. The input of each subsequent layer of the decoder is the output of the previous layer. Denote the output of each layer as where and , is the number of target queries, is the number of decoder layers, is the total number of decoder layers, C is the number of feature map channels, or the number of hidden layer channels, that is, the feature dimension of the sequence features; Step 42: The feature output by the decoder is learned through a fully connected layer and an activation function, and then the target-level fusion feature is obtained through the Concat operation layer; The second feature fusion module consists of a fully connected layer and a Concat operation layer; The feature fusion module (second feature fusion module) of the Transformer decoder is as Figure 4 shown. By fusing the sequence features of the decoder stack, it learns rich target semantic information, while avoiding complex feature alignment and alleviating the negative transfer phenomenon; Step 43: Perform target-level feature alignment in the decoder part to reduce the instance-level distribution difference and suppress the influence of factors such as target shape, texture, and color. Compared with the previous two alignment strategies, the target-level alignment strategy not only improves the generalization of the model but also ensures its discriminability. The structure of the domain discriminator (third domain discriminator) in the decoder part is also as Figure 3 shown; The target-level fusion feature is sequentially input into the gradient reversal layer and the fully connected layer in the third domain discriminator to obtain the domain prediction result , is the third domain discriminator; The structure of the third domain discriminator is the same as that of the first domain discriminator, consisting of a gradient reversal layer and a fully connected layer; Step 44: Obtain the decoder target-level adversarial training loss according to the domain prediction result.

[0038] Optionally, step 44 includes: The decoder target-level adversarial training loss is obtained through the following formula: (3) where is the decoder target-level adversarial training loss, represents the domain prediction result corresponding to the j-th target query of sample , is the domain label of sample . The domain label includes the source domain label and the target domain label, and their values are 0 and 1 respectively, or their values are 1 and 0 respectively.

[0039] Optionally, step 5 includes: Step 51: Obtain the adversarial training loss of the network model according to the backbone network image-level adversarial training loss, the encoder pixel-level adversarial training loss, and the decoder target-level adversarial training loss; Step 52: Obtain the optimal network model according to the adversarial training loss of the network model and the source domain detection loss.

[0040] Optionally, step 51 includes: The adversarial training loss of the network model is obtained through the following formula. It is composed of the backbone network image-level adversarial training loss, the encoder pixel-level adversarial training loss, and the decoder target-level adversarial training loss, and can be expressed as: (4) where is the adversarial training loss of the network model, is the backbone network image-level adversarial training loss, is the encoder pixel-level adversarial training loss, is the decoder target-level adversarial training loss; Step 52 includes: The specific training process of the network is as follows: The source domain and target domain data are input into the DINO model network together, and feature extraction is performed through the backbone network, the encoder, and the decoder. The backbone network image-level adversarial training loss, the encoder pixel-level adversarial training loss, and the decoder target-level adversarial training loss can be obtained respectively according to formulas (1), (2), and (3) in the above three feature extraction stages, and the overall adversarial training loss is obtained according to formula (4). The source domain samples have annotation information, and the source domain detection loss can be obtained.

[0041] Finally, the learning is completed by jointly optimizing the adversarial training loss and the source domain detection loss, as shown in the following formula: (5) where represents the first domain discriminator , Second Domain Discriminator and the third domain discriminator A collection of Indicates that the network model The rest of the network, is the source domain detection loss, is a trade-off parameter used to balance adversarial training loss and source domain detection loss; the rest of the network includes the backbone network, encoder, decoder, feedforward network, first feature fusion module, and second feature fusion module. After the above learning process, the network model can fully reduce the differences in data distribution in different fields and improve the generalization and discrimination of the model in cross-domain detection tasks.

[0042] For ease of understanding, this application provides a more specific embodiment: refer to Figure 1 , the present application proposes an embodiment of a domain adaptive DINO model target detection method based on feature fusion and alignment, including: (1) Obtain source domain and target domain data: Source domain data needs to contain annotation information, while target domain data does not need annotation information.

[0043] (2) Model and environment: Figure 1 As shown, this application is based on the DINO model, which includes a backbone network, an encoder, a decoder, and a feature fusion and alignment strategy. In this example, the backbone network uses ResNet-50, and both the encoder and the decoder use a six-layer stack. In addition, this embodiment is implemented using the Docker+VScode+PyTorch coding and operating environment, and the model is trained on a single NVIDIA A800 GPU.

[0044] (3) Network training process: The source domain and target domain data are input into the model designed in this application together, feature extraction is performed through the backbone network, encoder and decoder, and loss calculation is performed according to formulas (1), (2), (3) and (4) in the above embodiment. By optimizing formula (5), the method proposed in this application can fully reduce the differences in data distribution in different fields and improve the generalization and discrimination of the model in cross-domain detection tasks.

[0045] (4) Reasoning and detection phase: The reasoning process of the network model is exactly the same as that of the DINO model, that is, the target domain image data is input into the network model for detection. After the training is completed, the network model only needs to load the DINO model weights. Therefore, the method proposed in this application does not bring additional computational overhead and memory usage to the model when used.

[0046] (5) Detection task: In this example, an object detection task under varying weather conditions is constructed based on the publicly available Streetscape datasets Cityscapes (clear) and FoggyCityscapes (foggy). Quantitative results of cross-domain object detection for the method proposed in this application are analyzed, as shown in Table 1, including Domain Adaptive Faster RCNN (DAF), Strong-Weak Distribution Alignment (SWDA), Graph-induced Prototype Alignment (GPA), Spatial Alignment YOLO (SA-YOLO), Sequence Feature Alignment (SFA), Domain Adaptive DETR (DA-DETR), and the method proposed in this application. It can be seen that this application achieves the best performance, with a significant improvement in the detection accuracy of almost all categories. The overall mAP is 57.8, and there is also a good performance improvement in difficult example categories such as trucks, buses, trains, motorcycles, and bicycles, fully demonstrating the superior cross-domain detection ability of the proposed method.

[0047] Table 1 Cityscapes Quantitative results of cross-domain detection for Foggy Cityscapes

[0048] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application and are not intended to limit them. Although this application has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent replacements can still be made to the specific implementation manners of this application. Any modification or equivalent replacement that does not depart from the spirit and scope of this application shall be covered by the protection scope of the claims of this application.

Claims

1. A domain-adaptive DINO model target detection method based on feature fusion and alignment, characterized in that: include: Step 1: Input the source domain data and the target domain data into the backbone network of the DINO model for feature extraction; the network model includes the DINO model, the first feature fusion module, the first domain discriminator, the second domain discriminator, the second feature fusion module and the third domain discriminator; the DINO model includes a backbone network, an encoder, a decoder and a feedforward network connected in sequence; the source domain data includes the source domain image and its corresponding target label, and the target label includes the frame coordinates and category information of the target; the target domain data only includes the target domain image; Step 2: Input the features output by the backbone network into the first feature fusion module and the first domain discriminator in sequence to obtain the backbone network image-level adversarial training loss; Step 3: Input the features output by the encoder into the second domain discriminator to obtain the encoder pixel-level adversarial training loss; Step 4: Input the features output by the decoder into the second feature fusion module and the third domain discriminator in turn to obtain the decoder target level adversarial training loss; Step 5: According to the backbone network image-level adversarial training loss, encoder pixel-level adversarial training loss, decoder target-level adversarial training loss, and source domain detection loss, the network model is optimized so that the network model learns cross-domain invariant features and enhances the cross-domain generalization ability of the network model; Step 6: After the network model is trained, you only need to load the DINO model weights to complete cross-domain object detection on the target domain image data.

2. The method according to claim 1, characterized in that The step 2 comprises: Step 21: The first feature fusion module obtains image-level fusion features based on the multi-scale features extracted by the backbone network; Step 22: The first domain discriminator obtains an image-level domain prediction result based on the image-level fusion features; Step 23: Based on the image-level domain prediction results, the backbone network image-level adversarial training loss is obtained.

3. The method according to claim 2, characterized in that The step 21 comprises: The different scale features in the multi-scale features extracted by the backbone network are learned through the convolution layer in the first feature fusion module, and then the feature map size is unified through the global average pooling layer, and finally the image-level fusion features are obtained through the Concat operation layer; the first feature fusion module consists of a convolution layer, a global average pooling layer and a fully connected layer; The step 22 comprises: The first feature fusion module sequentially inputs the image-level fusion features into the gradient reversal layer and the fully connected layer in the first domain discriminator, and outputs the image-level domain prediction result; the first domain discriminator is composed of the gradient reversal layer and the fully connected layer.

4. The method according to claim 2, characterized in that: The step 23 comprises: The backbone network image-level adversarial training loss is obtained by the following formula: (1) in, is the backbone network image-level adversarial training loss, It is a sample The image-level domain prediction results are: For sample The domain label includes the source domain label and the target domain label, and their values ​​are 0 and 1 respectively.

5. The method according to claim 1, characterized in that The step 3 comprises: Step 31: The encoder models the global contextual relationship through the self-attention mechanism. Its input is the one-dimensional sequence features converted from the multi-scale features output by the encoder plus the position embedding. The input of each subsequent layer of the encoder is the output of the previous layer. Step 32: Input the output of the encoder into the gradient reversal layer and the fully connected layer in the second domain discriminator in turn to obtain the pixel-level domain prediction result , is the output of the encoder, is the second domain discriminator; the structure of the second domain discriminator is the same as that of the first domain discriminator, and is composed of a gradient reversal layer and a fully connected layer; Step 33: Based on the pixel-level domain prediction results, the encoder pixel-level adversarial training loss is obtained.

6. The method according to claim 5, characterized in that The step 33 comprises: The encoder pixel-level adversarial training loss is obtained by the following formula: (2) in, is the encoder pixel-level adversarial training loss, Representation sample The domain prediction result corresponding to the j-th pixel in the multi-scale feature map is For sample The domain label includes the source domain label and the target domain label, and their values ​​are 0 and 1 respectively.

7. The method according to claim 1, characterized in that The step 4 comprises: Step 41: Input the output of the encoder to the decoder to optimize the query anchor box and classification results. The input of each subsequent layer of the decoder is the output of the previous layer. The output of each layer is recorded as ,in ,and , is the target query number, is the number of decoder layers, is the total number of decoder layers, C is the number of feature mapping channels, or the number of hidden layer channels, that is, the feature dimension of the sequence feature; Step 42: Features of decoder output After learning through the fully connected layer and activation function, the target-level fusion features are obtained through the Concat operation layer. ; The second feature fusion module consists of a fully connected layer and a Concat operation layer; Step 43: Fusion of target-level features Input the gradient reversal layer and the fully connected layer in the third domain discriminator in turn to obtain the domain prediction result , is the third domain discriminator; the structure of the third domain discriminator is the same as that of the first domain discriminator, and is composed of a gradient reversal layer and a fully connected layer; Step 44: According to the domain prediction results, the decoder target level adversarial training loss is obtained.

8. The method according to claim 7, characterized in that The step 44 comprises: The decoder target-level adversarial training loss is obtained by the following formula: (3) in, is the decoder target-level adversarial training loss, Representation sample The domain prediction result corresponding to the j-th target query, For sample The domain label includes the source domain label and the target domain label, and their values ​​are 0 and 1 respectively, or, their values ​​are 1 and 0 respectively.

9. The method according to claim 1, characterized in that: The step 5 comprises: Step 51: Obtain the adversarial training loss of the network model according to the backbone network image-level adversarial training loss, the encoder pixel-level adversarial training loss, and the decoder target-level adversarial training loss; Step 52: Obtain the optimal network model based on the adversarial training loss and source domain detection loss of the network model.

10. The method according to claim 9, characterized in that The step 51 comprises: The adversarial training loss of the network model is obtained by the following formula: (4) in, is the adversarial training loss of the network model, is the backbone network image-level adversarial training loss, is the encoder pixel-level adversarial training loss, is the decoder target-level adversarial training loss; The step 52 comprises: Learning is accomplished by jointly optimizing the adversarial training loss and the source domain detection loss through the following formula: (5) in, Represents the first domain discriminator , Second Domain Discriminator and the third domain discriminator A collection of Indicates that the network model The rest of the network, is the source domain detection loss, It is a trade-off parameter used to balance the adversarial training loss and the source domain detection loss; the rest of the network includes the backbone network, encoder, decoder, feedforward network, first feature fusion module, and second feature fusion module.

Citation Information

Patent Citations

  • Zero sample target detection method based on DETR and meta learning

    CN116958741A

  • Tiny target detection method and device, storage medium and electronic equipment

    CN117911736A

  • Automatic driving 3D target detection method

    CN118587680A

Cited By

  • Photovoltaic power station detection method based on feature decoupling domain adaptation

    CN121685926A