Transform-based domain adaptive target detection method and electronic equipment
Through the domain adaptive object detection method based on Transformer, the domain adversarial feature alignment module and the adaptive domain discriminator are used to solve the problem of weak migration capabilities of the object detection method in different scenarios, and efficient detection in complex environments is achieved.
Patent Information
- Application Number
- CN202510387635.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-04
AI Technical Summary
The existing object detection methods have weak migration and generalization capabilities in different scenarios, making it difficult to cope with large domain offsets, resulting in performance degradation in actual applications.
The domain adaptive object detection method based on Transformer is adopted, and the domain adversarial feature alignment module and adaptive domain discriminator are designed, and the domain adversarial token is optimized using the cross attention mechanism and gradient inversion layer to enhance the domain adaptability of the model.
Maintain good detection performance in a cross-domain environment, simplify the training process, improve the domain generalization ability of the model, and significantly improve the detection effect.
Smart Images

Figure CN120259634A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, and particularly relates to a Transformer-based domain adaptation object detection method and an electronic device. Background Art
[0002] Object detection is one of the most important basic tasks in the field of computer vision, and its purpose is to enable the model to have the ability to predict the categories and corresponding positions of objects in a given image; Object detection algorithms based on CNN (Convolutional Neural Network) can be roughly divided into two categories: two-stage and single-stage object detection; Most classic two-stage detectors consist of a region proposal network and a classification network, while single-stage object detection eliminates the region proposal network and post-processing steps, thus simplifying the network structure; In addition, detectors based on transformers model context information through an attention mechanism, further improving the detection performance.
[0003] Currently, object detection methods highly rely on manually labeled data and mainly focus on scenarios where the data distributions of training and testing are the same; Since object detection usually assumes that the training data and the testing data come from the same distribution, although many current object detection models have achieved excellent performance on benchmark datasets (the source domain and the target domain are the same), in the real world, due to various factors such as shooting devices, environmental conditions, different perspectives, and different data sources, there is a significant domain shift between the training data and the testing data, which will lead to a significant performance degradation of the object detection model. For example, an autonomous driving system trained with data in good visibility cannot work reliably in rainy or foggy scenarios; The currently commonly used object detection methods have problems of weak transfer ability and low generalization ability in different scenarios, severely restricting the practical application of existing algorithms in actual scenarios. Summary of the Invention
[0004] The purpose of the present invention is to provide a Transformer-based domain adaptation object detection method and an electronic device that can cope with large domain shifts, reduce the domain differences between the image level and the object level during the domain adaptation object detection process, and have a simpler training process and better detection effect.
[0005] Based on the above purpose, the present invention adopts the following technical solutions:
[0006] A Transformer-based domain adaptation object detection method includes the following steps:
[0007] S1, Obtain the source domain and the target domain: Obtain the image data with annotations as the domain for obtaining training data, which is called the source domain; Obtain the image data without annotations as the domain for obtaining test data, which is called the target domain;
[0008] S2. Establish a model: The model includes a backbone network, a Transformer encoder, and a Transformer decoder connected in sequence;
[0009] S3. Establish a domain adversarial feature alignment module: Import the source domain and target domain obtained in step S1 into the backbone network in step S2 to generate features, and design a domain adversarial feature alignment module in the encoder part. Use the learnable domain adversarial token as the query and the features generated by the encoder layer as the key-value pair; Pre-discriminate the final domain adversarial token through an adaptive domain discriminator, and align the features generated by the encoder by minimizing the pre-discrimination loss;
[0010] S4. Adaptive domain discriminator: Establish an adaptive domain discriminator with an adaptive gradient reversal layer as the core, which can combine domain adaptation to perform adversarial mining on hard examples; By mining hard examples that are difficult to align, apply a stronger gradient reversal to them, enabling the model to better handle hard examples and enhancing its ability to extract and adapt to complex sample features.
[0011] Preferably, the specific process of establishing the domain adversarial feature alignment module in step S3 includes:
[0012] S31. Image-level feature alignment: Introduce an adaptive domain discriminator to classify the domain of the features generated by the backbone network to achieve image-level feature alignment;
[0013] S32. Token-level feature alignment: Obtain the domain adversarial token through the token-level feature alignment module. Use the domain adversarial token as the query and the features generated by the encoder layer as the key-value pair, and perform cross-attention on the two to continuously optimize the domain adversarial token;
[0014] S33. Optimize the domain adversarial feature alignment module: Optimize the entire domain adversarial feature alignment module through binary cross-entropy loss.
[0015] Preferably, the specific process of the image-level feature alignment in step S31 includes:
[0016] The adaptive domain discriminator is a multi-layer perceptron with 3 hidden layers, and its function is to predict whether the input features belong to the source domain or the target domain; The binary cross-entropy loss of the adaptive domain discriminator is:
[0017] L img = [S img LogP + (1 - S img )(1 - logP)]
[0018] where S img represents the domain label of the sample. For source domain pictures, S img = 0; For target domain pictures, S img= 1; P ∈ (0, 1) represents the predicted value of the domain classifier for the sample.
[0019] Preferably, the specific process of token-level feature alignment in step S32 includes:
[0020] Randomly initialize a c-dimensional vector to have the same number of channels as the features generated by the backbone network to obtain a domain adversarial token; let q0 represent the initial domain adversarial token, q i represents the domain adversarial token before cross-attention, q i+1 represents the domain adversarial token after cross-attention with the feature token of the i-th encoder layer, and z i represents the features output by the i-th encoder layer; use the domain adversarial token as the query and the features generated by the encoder layer as the key-value pair, and perform cross-attention on the two to continuously optimize the domain adversarial token. The update method of the domain adversarial token is:
[0021] q i+1 = Linear(CA(z i , q i ))
[0022] where Linear represents the linear layer and CA represents the cross-attention mechanism, which is essentially a multi-head attention mechanism with 8 parallel attention heads.
[0023] Preferably, the specific process of optimizing the domain adversarial feature alignment module in step S33 includes:
[0024] Successively perform cross-attention operations on q i and the features z i output by 6 encoder layers to obtain the final domain adversarial token q 6 ; optimize the entire domain adversarial feature alignment module using binary cross-entropy loss:
[0025]
[0026] where, S token represents the domain label, with a value of 0 representing the source domain and a value of 1 representing the target domain, and q n ∈ (0, 1) represents the predicted value of the domain classifier for the final domain adversarial token.
[0027] Preferably, the specific process of the adaptive domain discriminator in step S4 includes:
[0028] The forward propagation of the traditional gradient reversal layer is:
[0029] G x (n) = n
[0030] where n is the input feature vector, G xExecution function representing the gradient reversal layer;
[0031] The backpropagation is as follows:
[0032]
[0033] where -x is a negative scalar, generally -x = -1, and E is the identity matrix.
[0034] Preferably, the adaptive gradient reversal layer replaces x in the backpropagation with x ad , x ad which is expressed as:
[0035]
[0036] where x0 = 1, ε is a very small positive number, and L d is the loss of the adaptive domain discriminator, and this loss value reflects the difficulty of domain discrimination for the training samples; α(t) is a threshold that changes dynamically with the training process, α0 is the initial threshold, t represents the training round, and T is the total number of training rounds;
[0037] By comparing the numerical size of L d with α(t) to determine whether the training samples are challenging. When L d < α(t), it means that the domain discriminator performs good domain discrimination on the sample features, that is, the encoder generates more domain-private features. At this time, a larger numerical gradient reversal is taken. As the training progresses, α(t) continuously decreases, so that more samples are determined to be challenging; β is the overflow threshold, and its role is to avoid generating too many gradients in the backpropagation.
[0038] An electronic device, including a memory and a processor, and a computer program is stored on the memory. It is characterized in that: when the processor executes the computer program, any step in the above-mentioned transformer-based domain adaptive object detection method is implemented.
[0039] The beneficial effects of the present invention include:
[0040] In the encoder part of the present invention, a domain adversarial feature alignment module is designed. The learnable domain adversarial token is used as the query, and the features generated by the encoder layer are used as the key-value pair. The query and the key-value pair are subjected to a cross-attention mechanism and the final domain adversarial token is subjected to domain discrimination and the loss is calculated to achieve the alignment of the features generated by the encoder; and an adaptive domain discriminator composed of an adaptive gradient reversal layer and a domain discriminator is designed, so that the model can perform a larger numerical gradient reversal on difficult examples that are difficult to align, further enhancing the domain generalization ability of the model.
[0041] Before each cross-attention operation, the domain adversarial feature alignment module of the present invention passes the features generated by the encoder through a gradient reversal layer, making the optimization objective of the domain adversarial token the same as that of the domain discriminator, while the optimization direction of the features generated by the encoder layer is opposite to the optimization objective of the domain discriminator; by minimizing the loss function the domain adversarial token is updated to make it focus on the features private to the source domain and the target domain, and at the same time prompts the encoder to extract more domain-common features, thereby achieving domain adaptation.
[0042] The adaptive domain discriminator of the present invention designs an adaptive gradient reversal layer based on the traditional gradient reversal layer, by replacing x in the backpropagation formula of the traditional gradient reversal layer with a new x ad to achieve; x ad is determined based on the loss value of the adaptive domain discriminator. When the loss of the adaptive domain discriminator is smaller, it means that the domain of the training samples is easier to be recognized, and its features are not the desired domain-invariant features; the adaptive domain discriminator can dynamically adjust the gradient reversal parameter according to the actual situation of the samples, so as to better achieve domain adaptation, mine more valuable domain-invariant features, and improve the transfer learning performance of the model under complex samples.
[0043] The method used in the present invention has been actually tested, and quantitative and qualitative analyses are carried out with existing mainstream algorithms; through actual comparison, it is verified that the present invention has strong generalization ability in a cross-domain environment, and can still maintain good detection performance when facing data with different distributions; traditional methods often participate in training annotation bounding boxes by collecting more training data, or use the teacher-student framework of knowledge distillation and data augmentation to reduce the impact of domain shift, which has the problems of high cost and time consumption of annotating training data, higher training complexity and consumption of computing resources. The domain adaptation object detection method based on transformer proposed by the present invention has a simple training process and better detection effect. Brief Description of the Drawings
[0044] Figure 1 is the structural diagram of the present invention;
[0045] Figure 2 is the domain adversarial feature alignment module. Detailed Embodiments
[0046] Embodiment 1
[0047] The following is a further explanatory description of the present invention in combination with specific embodiments. As Figure 1 shown, this embodiment is a domain adaptation object detection method based on transformer, including the following steps:
[0048] S1. Obtain the source domain and the target domain: Obtain the annotated image data as the domain for obtaining training data, which is called the source domain; and obtain the unannotated image data as the domain for obtaining test data, which is called the target domain.
[0049] S2. Build a model: The model established in this embodiment is based on Deformable DETR. The model includes a backbone network, a Transformer encoder, and a Transformer decoder connected in sequence; among which, the backbone network adopts ResNet-50.
[0050] S3. Build a domain adversarial feature alignment module: Import the source domain and the target domain obtained in step S1 into the backbone network in step S2 to generate features, and design a domain adversarial feature alignment module in the encoder part. Use the learnable domain adversarial token as the query, and the features generated by the encoder layer as the key-value pair; pre-discriminate the final domain adversarial token through an adaptive domain discriminator, and align the features generated by the encoder by minimizing the pre-discrimination loss, including the following steps:
[0051] S31. Image-level feature alignment: Introduce an adaptive domain discriminator to classify the domain of the features generated by the backbone network to achieve image-level feature alignment.
[0052] The adaptive domain discriminator is a multi-layer perceptron with 3 hidden layers, and its function is to predict whether the input features belong to the source domain or the target domain; the binary cross-entropy loss of the adaptive domain discriminator is:
[0053] L img =[S img logP+(1 - S img )(1 - logP)]
[0054] where S img represents the domain label of the sample. For source domain pictures, S img =0; for target domain pictures, S img =1; P∈(0,1) represents the predicted value of the domain classifier for the sample.
[0055] S32. Token-level feature alignment: Obtain the domain adversarial token through the token-level feature alignment module. Use the domain adversarial token as the query and the features generated by the encoder layer as the key-value pair, and perform cross-attention on them to continuously optimize the domain adversarial token.
[0056] As Figure 2 shown, randomly initialize a c-dimensional vector to have the same number of channels as the features generated by the backbone network to obtain the domain adversarial token; let q0 represent the initial domain adversarial token, q i represent the domain adversarial token before cross-attention, q i+1represents the domain adversarial token after cross-attention with the feature tokens of the i-th encoder layer, and z i represents the features output by the i-th encoder layer; taking the domain adversarial token as the query and the features generated by the encoder layer as the key-value pair, the two perform cross-attention to continuously optimize the domain adversarial token. The update method of the domain adversarial token is as follows:
[0057] q i+1 = Linear(CA(z i ,q i ))
[0058] where Linear represents the linear layer and CA represents the cross-attention mechanism, which is essentially a multi-head attention mechanism with 8 parallel attention heads.
[0059] S33. Optimize the domain adversarial feature alignment module: Optimize the entire domain adversarial feature alignment module through binary cross-entropy loss.
[0060] Successively perform cross-attention operations on q i and the features z i output by 6 encoder layers to obtain the final domain adversarial token q 6 ; optimize the entire domain adversarial feature alignment module using binary cross-entropy loss:
[0061]
[0062] where S token represents the domain label, with a value of 0 indicating the source domain and a value of 1 indicating the target domain, and q n ∈(0,1) represents the predicted value of the domain classifier for the final domain adversarial token.
[0063] S4. Adaptive domain discriminator: Establish an adaptive domain discriminator with an adaptive gradient reversal layer as the core, which can combine domain adaptation to perform adversarial mining on hard examples; by mining hard examples that are difficult to align, applying a stronger gradient reversal to them enables the model to better handle hard examples and enhances the model's transfer learning ability under challenging samples, further enhancing the model's domain adaptation ability.
[0064] Traditional domain discriminators usually cooperate with gradient reversal layers to achieve feature alignment. The traditional gradient reversal layer keeps the input feature vector unchanged during the forward propagation process, and when backpropagating to the base network during training, it reverses the gradient by multiplying by a negative scalar; the forward propagation of the traditional gradient reversal layer is:
[0065] G x (n) = n
[0066] where n is the input feature vector and G xExecution function representing the gradient reversal layer;
[0067] The backpropagation is as follows:
[0068]
[0069] Among them, -x is a negative scalar, generally -x = -1, and E is the identity matrix.
[0070] In this embodiment, a new type of adaptive domain discriminator is proposed, and its adaptive gradient reversal layer replaces x in the backpropagation with x ad to implement, x ad is expressed as:
[0071]
[0072] Among them, x0 = 1, ε is a very small positive number, and in this embodiment, it is taken as 10 -8 , L d is the loss of the adaptive domain discriminator, and this loss value reflects the difficulty of domain discrimination for training samples; α(t) is a threshold that changes dynamically with the training process, α0 is the initial threshold, t represents the number of training rounds, and T is the total number of training rounds. In this embodiment, it is taken as 50;
[0073] By comparing the numerical sizes of L d and α(t) to determine whether the training samples are challenging. When L d < α(t), it means that the domain discriminator performs good domain discrimination on the sample features, that is, the encoder generates more domain-private features. At this time, a larger numerical gradient reversal is taken. As the training progresses, α(t) continuously decreases, so that more samples are determined to be challenging; β is the overflow threshold, and its role is to avoid generating too many gradients in the backpropagation. In this embodiment, it is set as a fixed value of 30.
[0074] Embodiment 2
[0075] Based on the method provided in Embodiment 1, this embodiment sets up two groups of experiments, namely two groups of experiments with the same scene but different weather (the source domain is the CityScape dataset, and the target domain is the Foggy CityScape dataset), and different scenes (the source domain is the CityScape dataset, and the target domain is the BDD100k daytime dataset); the method in Embodiment 1 and several other common mainstream algorithms are respectively tested on the above standard datasets, and quantitative and qualitative analyses are carried out. The results are shown in the following table:
[0076] Table 1. Experimental results with the source domain being the CityScape dataset and the target domain being the Foggy CityScape dataset
[0077] Method Person Rider Car Truck Bus Train Motor Bike mAP DA FasterR-CNN 35.3 27.1 40.5 20.0 25.0 31.0 20.2 22.1 27.6 MeGA 49.2 39.0 52.4 34.5 37.7 49.0 46.9 25.4 42.8 TIA 52.1 38.1 49.7 37.7 34.8 46.3 48.6 31.1 42.3 Ours 49.3 55.4 67 25.0 50.0 19.4 36.4 47.7 43.8
[0078] Table 2. Experimental results with the source domain being the CityScape dataset and the target domain being the BDD100k daytime dataset
[0079] Method Person Rider Car Truck Bus Motor Bike mAP DA FasterR-CNN 28.8 25.4 44.1 17.9 16.1 13.9 22.4 24.1 ICR-CCR-SW 32.8 29.3 45.8 22.7 20.6 14.9 25.5 27.4 Deformable DETR 38.9 26.7 55.2 15.7 19.7 10.8 16.2 26.2 Ours 36.2 35.0 55.3 19.2 22.3 16.9 23.5 29.7
[0080] As shown in the results of the above two tables, the method in Embodiment 1 of the present invention has achieved the best results in all metrics, indicating that it has obvious advantages compared with the commonly used mainstream methods in scenarios with large domain offsets.
[0081] As described above, it is only a further explanatory description of the present invention in combination with specific embodiments. All the descriptions made do not represent a limitation on the protection scope of the present invention. Any changes or alternative solutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claimed rights.
Claims
1. A Transformer-based domain adaptation object detection method, comprising the following steps: S1. Obtain the source domain and the target domain: Obtain the annotated image data as the domain for obtaining training data, which is called the source domain; The unannotated image data is used as the domain for obtaining test data, which is called the target domain; S2. Establish a model: The model includes a backbone network, a Transformer encoder, and a Transformer decoder connected in sequence; S3. Establish a domain adversarial feature alignment module: Import the source domain and the target domain obtained in step S1 into the backbone network in step S2 to generate features, and design a domain adversarial feature alignment module in the encoder part. Use the learnable domain adversarial token as the query and the features generated by the encoder layer as the key-value pair; Pre-discriminate the final domain adversarial token through an adaptive domain discriminator, and align the features generated by the encoder by minimizing the pre-discrimination loss; S4. Adaptive domain discriminator: Establish an adaptive domain discriminator with an adaptive gradient reversal layer as the core, which can combine domain adaptation to perform adversarial mining on difficult examples; By mining difficult examples that are difficult to align, apply a stronger gradient reversal to them, so that the model can better handle difficult examples and enhance its ability to extract and adapt to the features of complex samples.
2. The transformer-based domain adaptation object detection method according to claim 1, wherein: The specific process of establishing the domain adversarial feature alignment module in step S3 includes: S31. Image-level feature alignment: Introduce an adaptive domain discriminator to classify the domain of the features generated by the backbone network to achieve image-level feature alignment; S32. Token-level feature alignment: Obtain the domain adversarial token through the token-level feature alignment module. Use the domain adversarial token as the query and the features generated by the encoder layer as the key-value pair, and perform cross-attention on the two to continuously optimize the domain adversarial token; S33. Optimize the domain adversarial feature alignment module: Optimize the entire domain adversarial feature alignment module through binary cross-entropy loss.
3. The transformer-based domain adaptive object detection method according to claim 2, wherein: The specific process of the image-level feature alignment in step S31 includes: The adaptive domain discriminator is a multi-layer perceptron with 3 hidden layers, and its function is to predict whether the input features belong to the source domain or the target domain; The binary cross-entropy loss of the adaptive domain discriminator is: L img = [S img logP + (1 - S img )(1 - logP)] Among them, S img represents the domain label of the sample. For source domain images, S img = 0; for target domain images, S img = 1; P ∈ (0, 1) represents the predicted value of the domain classifier for the sample.
4. The method for domain adaptive object detection based on transformer according to claim 3, wherein: The specific process of the token-level feature alignment in step S32 includes: Randomly initialize a c-dimensional vector to have the same number of channels as the features generated by the backbone network to obtain the domain adversarial token; let q0 denote the initial domain adversarial token, q i denote the domain adversarial token before cross-attention, q i+1 denote the domain adversarial token after cross-attention with the feature token of the i-th encoder layer, while z i denote the features output by the i-th encoder layer; use the domain adversarial token as the query and the features generated by the encoder layer as the key-value pair, and perform cross-attention between the two to continuously optimize the domain adversarial token. The update method of the domain adversarial token is as follows: q i+1 = Linear(CA(z i ,q i )) Among them, Linear represents the linear layer, CA represents the cross-attention mechanism, which is essentially a multi-head attention mechanism with 8 parallel attention heads.
5. The domain adaptation object detection method based on transformer according to claim 4, wherein: The specific process of optimizing the domain adversarial feature alignment module in step S33 includes: Successively take q i and perform cross-attention operations with the features z output by 6 encoder layers i to obtain the final domain adversarial token q 6 ; Optimize the entire domain adversarial feature alignment module using binary cross-entropy loss: Among them, S token represents the domain label, with a value of 0 indicating the source domain and a value of 1 indicating the target domain. q n ∈(0,1) represents the predicted value of the domain classifier for the final domain adversarial token.
6. The domain adaptation object detection method based on transformer according to claim 5, wherein: The specific process of the adaptive domain discriminator in step S4 includes: The forward propagation of the traditional gradient reversal layer is: G x (n) = n where n is the input feature vector, and G x represents the execution function of the gradient reversal layer; The backward propagation is: Among them, -x is a negative scalar, generally -x = -1, and E is the identity matrix.
7. The method for domain adaptive object detection based on transformer according to claim 6, wherein: The adaptive gradient reversal layer replaces x in backpropagation with x ad , x ad which is expressed as: where x0 = 1, ε is a very small positive number, L d is the loss of the adaptive domain discriminator, and this loss value reflects the difficulty of domain discrimination for training samples; α(t) is a threshold that dynamically changes with the training process, α0 is the initial threshold, t represents the training round, and T is the total number of training rounds; By comparing the magnitude of L d with that of α(t) to determine whether the training samples are challenging. When L d < α(t), it indicates that the domain discriminator performs well in domain discrimination on the sample features, that is, the encoder generates more domain-private features. At this time, a larger numerical gradient reversal is adopted. As the training progresses, α(t) continuously decreases, resulting in more samples being determined to be challenging; β is the overflow threshold, whose role is to avoid generating excessive gradients during backpropagation.
8. An electronic device, comprising a memory and a processor, wherein a computer program is stored on the memory, and characterized in that: When the processor executes the computer program, it implements any step in the Transformer-based domain adaptation object detection method according to any one of claims 1-7.
Citation Information
Cited By
Photovoltaic power station detection method based on feature decoupling domain adaptation
CN121685926A