Cross-domain remote sensing image target detection method based on multi-scale decoupling representation and reinforcement learning

Through the multi-scale decoupling of characterization module and the pseudo-label screening mechanism of reinforcement learning, the problems of data distribution differences and low pseudo-label quality in cross-domain remote sensing image object detection are solved, and more efficient cross-domain object detection performance is achieved.

CN120182580APending Publication Date: 2025-06-20CHINA UNIV OF MINING & TECH
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510335126.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The existing cross-domain remote sensing image object detection methods face problems such as large data distribution differences, difficulty in obtaining domain-invariant semantic features, and low pseudo-label quality, resulting in poor performance of the model in cross-domain detection.

Method used

Using a method based on multi-scale decoupled representation and reinforcement learning, the multi-scale domain invariant semantic features are obtained through the multi-scale decoupled representation module, and a category adaptive pseudo-label screening mechanism based on reinforcement learning is introduced to improve the quality of pseudo-labels.

Benefits of technology

It effectively reduces the impact of the distribution differences of remote sensing image data, improves the performance of cross-domain target detection, and improves the model's adaptability to the target domain.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005321500660000021
    Figure BDA0005321500660000021
  • Figure BDA0005321500660000044
    Figure BDA0005321500660000044
  • Figure BDA0005321500660000046
    Figure BDA0005321500660000046
Patent Text Reader

Abstract

The invention discloses a cross-domain remote sensing image target detection method based on multi-scale decoupling representation and reinforcement learning. The method comprises the following steps: step 1, acquiring a cross-domain remote sensing image target detection data set; 2, building a DINO remote sensing image cross-domain target detection model, obtaining an initial multi-scale feature map through a backbone network, obtaining multi-scale domain invariant semantic features by using a multi-scale decoupling representation module, and reducing domain deviation between a source domain and a target domain; 3, performing multiple rounds of training on the model built in the step 2; 4, multiple rounds of retraining are conducted on the trained model, on the basis that the training logic in the step 3 is kept, a category self-adaptive pseudo-label screening mechanism based on reinforcement learning is introduced, the model obtains high-quality pseudo-labels of a target domain, model training is guided, effective target domain information is learned, and adaptability to the target domain is improved; and 5, performing detection on the model obtained in the step 4 by using a target domain test set to improve the cross-domain target detection performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to a cross-domain remote sensing image target detection method based on multi-scale decoupled representation and reinforcement learning. Background Art

[0002] In recent years, with the continuous development of artificial intelligence technology, the acquisition of remote sensing images has become simpler, the quality of remote sensing image data has been continuously improved, and the remote sensing image target detection methods based on deep learning have shown good performance. However, these methods are usually data-driven and require a large amount of labeled data for training. Since remote sensing image data is updated quickly and the annotation is complex, the cost is relatively high. In addition, when the model trained on one remote sensing dataset is applied to another remote sensing dataset, the performance often decreases significantly, resulting in low reusability of the trained model. Essentially, this is because there are differences in the data distributions between the two remote sensing datasets, which may be caused by factors such as different atmospheric conditions, acquisition locations, and observation angles when collecting remote sensing images of different datasets. In addition, the background of remote sensing images is complex and the size varies greatly, which also makes the cross-domain detection task of remote sensing images difficult.

[0003] To solve the above problems, a series of cross-domain target detection methods for remote sensing images have been proposed. However, the existing cross-domain remote sensing image target detection methods still face some challenges as follows:

[0004] (1) Large difference in data distribution and difficulty in obtaining domain-invariant semantic features: Obtaining domain-invariant semantic features between the source domain and the target domain can improve the performance of cross-domain remote sensing image target detection. However, due to the large difference in the data distributions of remote sensing images, it is difficult to obtain domain-invariant semantic features.

[0005] (2) Low quality of pseudo-labels and containing noise: The self-training process generates pseudo-labels in the target domain, which can theoretically provide target domain annotations, reduce the annotation cost, enable the model to learn target domain information, and achieve cross-domain detection. However, an inappropriate pseudo-label selection strategy will result in pseudo-labels with low quality and containing noise, that is, mislabeled pseudo-labels, which makes the model learn confusing information and reduces the cross-domain adaptability of the model. Summary of the Invention

[0006] The problem to be solved by the present invention is to provide a cross-domain remote sensing image target detection method based on multi-scale decoupled representation and reinforcement learning, and propose corresponding solutions for the problems of large difference in data distribution, difficulty in obtaining domain-invariant semantic features, low quality of pseudo-labels and containing noise.

[0007] The present invention adopts the following technical solutions: A cross-domain remote sensing image target detection method based on multi-scale decoupled representation and reinforcement learning, comprising the following steps:

[0008] Step 1: Obtain a cross-domain remote sensing image target detection dataset and construct a training set;

[0009] Step 2: Build a DINO cross-domain target detection model for remote sensing images. Obtain an initial multi-scale feature map through the backbone network Resnet50, and obtain multi-scale domain-invariant semantic features through a multi-scale decoupled representation module to reduce the domain bias between the source domain and the target domain;

[0010] Step 3: Use the training set to perform multiple rounds of training on the DINO cross-domain target detection model for remote sensing images built in Step 2;

[0011] Step 4: Based on the results of multiple rounds of training in Step 3, retrain the DINO cross-domain target detection model for remote sensing images. Introduce a category-adaptive pseudo-label screening mechanism based on reinforcement learning to obtain high-quality pseudo-labels for the target domain, guide the model to perform multiple rounds of retraining, learn effective target domain information, improve the adaptability to the target domain, and obtain an optimized DINO cross-domain target detection model for remote sensing images;

[0012] Step 5: Use the target domain test set to perform detection on the optimized DINO cross-domain target detection model for remote sensing images obtained in Step 4.

[0013] Preferably, the method in Step 1 is as follows:

[0014] Step 101: Obtain the cross-domain remote sensing image target detection datasets xView and DOTA, and perform category screening;

[0015] Step 102: Crop the images and cut the original cross-domain remote sensing images into images of a set size;

[0016] Step 103: Construct a training set, use the cropped images for model training, and perform data augmentation during training, including: random rotation, flipping, and random photometric distortion.

[0017] Preferably, the process of building the DINO cross-domain target detection model for remote sensing images in Step 2 is as follows:

[0018] Step 201: Build a student model, which is improved based on the DINO framework and includes: a Resnet50 backbone network, a multi-scale decoupled representation module, a transformer encoder-decoder, and a detection head;

[0019] Step 202: Copy the student model as the teacher model. The parameter initialization is the same as that of the student model, and the parameter update is carried out by the student model using EMA. The update formula is as follows:

[0020]

[0021] Among them, t represents the current iteration round, and α represents the update weight. represents the parameters of the teacher model at the t-th iteration. represents the parameters of the teacher model at the (t - 1)-th iteration. represents the parameters of the student model at the t-th iteration.

[0022] Preferably, in step 201, the multi-scale decoupled representation module is constructed as follows:

[0023] Step 201.1: Input the source domain image x s and the target domain image x t into the backbone network Resnet50 of the student model respectively, and obtain feature maps of different scales, which are respectively denoted as

[0024] Among them, corresponds to the features after the initial convolutional layer and pooling layer of the source domain image. corresponds to the features of the subsequent layers of the source domain image; corresponds to the features after the initial convolutional layer and pooling layer of the target domain image. corresponds to the features of the subsequent layers of the target domain image;

[0025] Step 201.2: Input the source domain feature map and the target domain feature map into the semantic correlation feature extractor Apply a convolutional layer to and respectively, use a 1x1 convolutional kernel, with a stride of 1, and the output dimension is 256. Then connect a group normalization layer to extract semantic correlation features and reduce the feature dimension, and obtain the corresponding semantic correlation feature maps and

[0026] Input the source domain feature map and the target domain feature map into the domain correlation feature extractor Apply a convolutional layer to and respectively, use a 1x1 convolutional kernel, with a stride of 1, and the output dimension is 256. Then connect a group normalization layer to extract domain correlation features and reduce the feature dimension, and obtain the corresponding domain correlation feature maps and

[0027] Input the source domain feature map and the target domain feature map into the semantic correlation feature extractor Apply to and Apply a convolutional layer respectively, using a 1x1 convolutional kernel, a stride of 1, and an output dimension of 256. Then connect a group normalization layer to extract semantically related features and reduce the feature dimension, obtaining corresponding semantically related feature maps. and

[0028] Input the source domain feature map and the target domain feature map into the domain-related feature extractor respectively For and Apply a convolutional layer respectively, using a 1x1 convolutional kernel, a stride of 1, and an output dimension of 256. Then connect a group normalization layer to extract domain-related features and reduce the feature dimension, obtaining corresponding domain-related feature maps. and

[0029] Input the source domain feature map and the target domain feature map into the semantically related feature extractor respectively For and Apply a convolutional layer respectively, using a 1x1 convolutional kernel, a stride of 1, and an output dimension of 256. Then connect a group normalization layer to extract semantically related features and reduce the feature dimension, obtaining corresponding semantically related feature maps. and

[0030] Input the source domain feature map and the target domain feature map into the domain-related feature extractor respectively For and Apply a convolutional layer respectively, using a 1x1 convolutional kernel, a stride of 1, and an output dimension of 256. Then connect a group normalization layer to extract domain-related features and reduce the feature dimension, obtaining corresponding domain-related feature maps. and

[0031] Input the source domain feature map and the target domain feature map into the semantically related feature extractor respectively For and Apply a convolutional layer respectively, using a 3x3 convolutional kernel, a stride of 2, a padding of 1, and an output dimension of 256. Then connect a group normalization layer to extract semantically related features and reduce the feature dimension, obtaining corresponding semantically related feature maps. and

[0032] Input the source domain feature map and the target domain feature map into the domain-related feature extractor respectively Apply a convolutional layer to and respectively, using a 3x3 convolutional kernel, a stride of 2, a padding of 1, and an output dimension of 256. Then connect a group normalization layer to extract domain-related features and reduce the feature dimension, obtaining the corresponding domain-related feature maps and

[0033] Step 201.3: Through the multi-scale decoupled representation module, obtain the source domain multi-scale semantic-related feature map The target domain multi-scale semantic-related feature map The source domain multi-scale domain-related feature map The target domain multi-scale domain-related feature map

[0034] Preferably, in step 2, the process of obtaining multi-scale domain-invariant semantic features is as follows:

[0035] Step 211: Construct a domain discriminator network D, including several convolutional layers and RELU activation function layers;

[0036] Step 212: Input the source domain and target domain multi-scale domain-related feature maps The source domain and target domain multi-scale semantic-related feature maps into the domain discriminator network D respectively, and for each position in the feature map output by D, use the Softmax function for normalization processing to obtain the corresponding source domain and target domain domain discrimination probability values, and overall obtain the source domain and target domain multi-scale domain-related domain discrimination feature maps The source domain and target domain multi-scale semantic-related domain discrimination feature maps

[0037] Step 213: Set the source domain image feature map label to 0 and the target domain image feature map label to 1;

[0038] Step 214: Calculate the domain discrimination loss for the source domain and target domain multi-scale domain-related domain discrimination feature maps so that the domain discriminator network learns domain-related knowledge. The loss calculation is as follows:

[0039]

[0040] Among them, respectively represent and the values corresponding to the (u, v) position, u i、v i Respectively and Width and height; L CE (,) represents the cross entropy loss between prediction and label, represents the domain discrimination loss and the domain-related feature map of the i-th scale domain;

[0041] The mean domain discrimination loss of the overall domain-related feature map L dis-domain , calculated as follows:

[0042]

[0043] Step 215: Discriminate feature maps of multi-scale semantically related domains in the source domain and target domain Calculate the domain discrimination loss. The loss is calculated as follows:

[0044]

[0045] in, Respectively and The value corresponding to the (u,v) position, u i 、v i Respectively and Width and height; L CE (,) represents the cross entropy loss between prediction and label, represents the domain discrimination loss and the sum of the semantically relevant feature maps at the i-th scale;

[0046] The gradient reversal layer is introduced for adversarial training. The overall adversarial loss mean and adversarial optimization logic are as follows:

[0047]

[0048] Among them, L adv is the overall adversarial loss mean; for adversarial optimization logic, The purpose of is to deceive the discriminator, and the purpose of the domain discriminator D is to discriminate the domain source of the semantic feature. The adversarial process makes Gradually extract domain-invariant information from the source domain to the target domain, and obtain domain-invariant semantic features at multiple scales;

[0049] Step 216: Send the multi-scale domain-invariant semantic features of the source domain and target domain after adversarial training to the subsequent network structure. Since the domain discrimination network D has learned domain-related knowledge in step 214, It is necessary to dig deeper into the domain-independent components of the source domain and the target domain in order to deceive the domain discrimination network D and obtain domain-invariant semantic features;

[0050] The overall loss of the multi-scale decoupled representation module for obtaining multi-scale domain-invariant semantic features is denoted as L MDR , and the calculation is as follows:

[0051] L MDR = L dis-domain + L adv .

[0052] Preferably, in step 211, the domain discriminator network D has the following structure:

[0053] The first convolutional layer uses a 3x3 convolutional kernel, a stride of 1, padding of 1, and an output dimension of 256;

[0054] The first ReLU activation function layer;

[0055] The second convolutional layer uses a 3x3 convolutional kernel, a stride of 1, padding of 1, and an output dimension of 256;

[0056] The second ReLU activation function layer;

[0057] The third convolutional layer uses a 3x3 convolutional kernel, a stride of 1, padding of 1, and an output dimension of 256;

[0058] The third ReLU activation function layer;

[0059] The fourth convolutional layer uses a 3x3 convolutional kernel, a stride of 1, padding of 1, and an output dimension of 2.

[0060] Preferably, in step 3, the multi-round training method is as follows:

[0061] Step 301, calculate the supervision loss between the object detection results of the source domain images generated by the student model and the corresponding annotations

[0062] Step 302, calculate the loss L of the multi-scale decoupled representation module of the student model for obtaining multi-scale domain-invariant semantic features MDR ;

[0063] Step 303, use the calculated source domain supervision loss and the loss L of obtaining multi-scale domain-invariant semantic features MDR to train the student model and update the teacher model with the student model;

[0064] Step 304, repeat steps 301 to 303 until the end of multi-round training.

[0065] Preferably, in step 4, a category adaptive pseudo-label screening mechanism based on reinforcement learning is introduced to obtain high-quality pseudo-labels for the target domain and guide the model for multi-round retraining. The specific process is as follows:

[0066] Step 401: Set an initial screening threshold δ for each type of object in domain adaptation object detection. ij , δ ij represents the threshold for screening pseudo-labels of the j-th type of object in the i-th round of the initial round of multiple rounds of retraining, where j = 1, 2,..., N, and N is the number of object categories.

[0067] Step 402: Input the target domain image into the teacher model for prediction. Each image generates multiple prediction results, and each prediction result includes a confidence level and the corresponding coordinate box

[0068] wherein, represents that in the i-th round of prediction, the category of the l-th prediction result generated is the j-th category, and the corresponding confidence level is are respectively the x coordinate of the center point of the corresponding prediction box, the y coordinate of the center point, the width of the coordinate box, and the height of the coordinate box.

[0069] Step 403: Screen the pseudo-labels generated by the teacher model. For the multiple pseudo-labels predicted for each image, retain the predictions with a confidence level greater than the corresponding category threshold δ ij as reliable pseudo-labels, and use the coordinate box information as the annotation of the corresponding target domain image.

[0070] Step 404: Input the target domain image into the student model, and supervise the object detection results generated by the student model for the target domain image through the retained reliable pseudo-labels.

[0071] Step 405: Repeat steps 402 to 404 until the i-th round of training is completed, and adaptively adjust the threshold δ for screening pseudo-labels of each type of object in the next round by means of reinforcement learning i+1j ;

[0072] Step 406: Conduct multiple rounds of training according to the logic of steps 402 to 405 until the multiple rounds of training are completed.

[0073] Preferably, in step 405, adaptively adjust the threshold δ for screening pseudo-labels of each type of object in the next round by means of reinforcement learning i+1j , and the specific process is as follows:

[0074] Step 405.1: Take the teacher model as the agent, and take the setting of the threshold for screening pseudo-labels of each type of object as the strategy for the agent to select pseudo-labels.

[0075] Step 405.2: Use the agent obtained after the (i - 1)-th round of training to predict the target domain image, and obtain the prediction result mAP50 i-1j ;

[0076] Step 405.3: Use the agent obtained from the just-completed i-th round of training to predict the target-domain image, and obtain the prediction result mAP50 ij ;

[0077] Step 405.4: Calculate the reward of the j-th class object in the current i-th round according to the prediction results of the (i - 1)-th round and the i-th round ij , which is used to determine whether to update the pseudo-label selection strategy and how to update it:

[0078] reward ij = |mAP50 ij - mAP50 i-1j |

[0079] Determine the strategy for the agent to select pseudo-labels for the corresponding class according to the corresponding class reward, and obtain the threshold δ for screening pseudo-labels of the j-th class object in the (i + 1)-th round i+1j , and the formula is:

[0080]

[0081] Among them, α and β are hyperparameters. α is used to control the update amplitude. The larger α is, the greater the update amplitude; β determines the influence degree of new information on the update of the threshold strategy; IU ij is the increment during the policy iteration of the j-th class object in the i-th round; τ is a hyperparameter. When reward ij exceeds τ, the policy will be updated.

[0082] Preferably, in Step 4, the DINO remote sensing image cross-domain object detection model is retrained for multiple rounds, and the process is as follows:

[0083] Step 411: Determine the pseudo-label screening thresholds of various objects in the current round according to the method of Step 405 with the idea of reinforcement learning;

[0084] Step 412: Use the teacher model of the current round to generate predictions, and screen the predictions according to the pseudo-label screening thresholds determined in Step 411 to obtain reliable pseudo-labels

[0085] Step 413: Calculate the supervised loss between the object detection results of the source-domain images generated by the student model and the corresponding annotations

[0086] Step 414: Calculate the supervised loss between the object detection results of the target-domain images generated by the student model and the corresponding reliable pseudo-labels

[0087] Step 415: Calculate the loss L of the multi-scale decoupled representation module in the student model to obtain multi-scale domain-invariant semantic features MDR ;

[0088] Step 416: Use the loss L MDR to train the student model and update the teacher model with the student model;

[0089] Step 417: Repeat Steps 412 to 416 until the end of one round of training;

[0090] Step 418: Repeat Steps 411 to 417. First, adaptively adjust the threshold for screening pseudo-labels for each category of objects with the idea of reinforcement learning, and then perform the next round of model training. After at most several rounds of training, obtain the optimized DINO cross-domain object detection model for remote sensing images.

[0091] Compared with the prior art, the present invention adopts the above technical solutions and has the following technical effects:

[0092] 1. The cross-domain object detection method for remote sensing images of the present invention realizes feature decoupling at multiple scales through a multi-scale decoupled representation module, promotes the domain invariance of semantically related features with domain discrimination of domain-related features, reduces the influence of large differences in the data distribution of remote sensing images, obtains domain-invariant features, and improves the performance of cross-domain object detection, solving the problems of large differences in the data distribution of remote sensing images and difficulty in obtaining domain-invariant features in the prior art.

[0093] 2. The cross-domain object detection method for remote sensing images of the present invention, through a category-adaptive pseudo-label screening mechanism based on reinforcement learning, adaptively adjusts the strategy for screening pseudo-labels for each type of target with the idea of reinforcement learning according to the training process, obtains high-quality and low-noise pseudo-labels, enables the model to learn effective target domain information, and improves the performance of cross-domain object detection, solving the problem of partial errors and confusing information in the target domain pseudo-labels generated in the self-training process of the teacher-student model in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS

[0094] Figure 1 is the overall flow block diagram of the cross-domain object detection method for remote sensing images of the present invention;

[0095] Figure 2 is the overall network structure diagram of the cross-domain object detection method for remote sensing images of the present invention;

[0096] Figure 3 is the network structure diagram for obtaining multi-scale domain-related and semantically related features in the multi-scale decoupled representation module of the present invention;

[0097] Figure 4 is the network structure diagram for obtaining multi-scale domain-invariant semantic features in the multi-scale decoupled representation module of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0098] To make the objectives, technical solutions, and advantages of the present invention clearer, the following further elaborates on the technical solutions of the application in conjunction with the accompanying drawings. The described embodiments are only a part of the embodiments related to the present invention. All non-innovative embodiments made by other researchers in this field belong to the protection scope of the present invention. At the same time, for the step numbers in the embodiments of the present invention, they are only set for the convenience of elaboration and explanation, and no limitation is imposed on the order between steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0099] In one embodiment of the present invention, a cross-domain remote sensing image target detection method based on multi-scale decoupled representation and reinforcement learning is as Figure 1 shown, and includes the following steps:

[0100] Step 1: Obtain a cross-domain remote sensing image target detection dataset;

[0101] Step 2: Build a DINO remote sensing image cross-domain target detection model. Obtain an initial multi-scale feature map through the backbone network Resnet50, and obtain multi-scale domain-invariant semantic features through the multi-scale decoupled representation module to reduce the domain bias between the source domain and the target domain;

[0102] Step 3: Use the training set to perform multiple rounds of training on the DINO remote sensing image cross-domain target detection model built in Step 2;

[0103] Step 4: Based on the results of multiple rounds of training in Step 3, retrain the DINO remote sensing image cross-domain target detection model. Introduce a category adaptive pseudo-label screening mechanism based on reinforcement learning to obtain high-quality pseudo-labels in the target domain, guide the model to perform multiple rounds of retraining, learn effective target domain information, improve the adaptability to the target domain, and obtain an optimized DINO remote sensing image cross-domain target detection model;

[0104] Step 5: Use the target domain test set to perform detection on the optimized DINO remote sensing image cross-domain target detection model obtained in Step 4.

[0105] Specifically, the method of Step 1 is as follows:

[0106] (2.1) Obtain the cross-domain remote sensing image target detection datasets xView and DOTA.

[0107] In this embodiment, the xView dataset is obtained through WorldView-3 and contains more than one million instances of 60 categories; the DOTA dataset is obtained through Google Earth and includes 15 categories, covering 2806 images with 188,282 instances.

[0108] To meet the needs of domain adaptation object detection, three categories of airplanes, ships, and storage tanks that are common in two datasets were selected. When screening the categories, due to the fine-grained annotations in xView, the sub-categories were merged into the superior xView categories.

[0109] (2.2) Crop the image, cutting the original image into images of size 800*800;

[0110] (2.3) Use the cropped images for model training, and use data augmentations such as random rotation, flipping, and random photometric distortion during the training process.

[0111] Specifically, the overall network structure of the cross-domain remote sensing image object detection method involved in steps 2 to 4 is as Figure 2 shown.

[0112] The method in step 2 is as follows:

[0113] (3.1) Build a student model, which is improved based on the DINO framework and includes: a Resnet50 backbone network, a multi-scale decoupled representation module, a transformer encoder-decoder, and a detection head, as the student model;

[0114] (3.2) Copy the student model as the teacher model, with the parameter initialization being the same as that of the student model. Subsequent parameter updates are carried out by the student model using EMA, and the update formula is as follows:

[0115]

[0116] Among them, t represents the current iteration round, α represents the update weight, represents the parameters of the teacher model at the t-th iteration, represents the parameters of the teacher model at the (t - 1)-th iteration, represents the parameters of the student model at the t-th iteration.

[0117] Furthermore, in step (3.1), the multi-scale decoupled representation module, as Figure 3 and Figure 4 shown, is constructed as follows:

[0118] (4.1) Input the source domain image x s , the target domain image x t into the Resnet50 backbone network of the student model respectively, and obtain the original feature maps of different scales, which are respectively represented as

[0119] Among them, corresponds to the features after the initial convolutional layer and pooling layer of the source domain image, corresponds to the features of the subsequent layers of the source domain image; Corresponding to the features after the initial convolutional layer and pooling layer of the target domain image, Corresponding to the features of the subsequent layers of the target domain image.

[0120] (4.2) Input the source domain feature map Target domain feature map into the semantic-related feature extractor Specifically: For and First, apply a convolutional layer with a 1x1 convolutional kernel, a stride of 1, and an output dimension of 256, and then connect a group normalization layer, specifically divided into 32 groups, to extract semantic-related features and reduce the feature dimension, obtaining the corresponding semantic-related feature maps and

[0121] Input the source domain feature map Target domain feature map into the domain-related feature extractor Specifically: For and First, apply a convolutional layer with a 1x1 convolutional kernel, a stride of 1, and an output dimension of 256, and then connect a group normalization layer, specifically divided into 32 groups, to extract domain-related features and reduce the feature dimension, obtaining the corresponding domain-related feature maps and

[0122] Input the source domain feature map Target domain feature map into the semantic-related feature extractor Specifically: For and First, apply a convolutional layer with a 1x1 convolutional kernel, a stride of 1, and an output dimension of 256, and then connect a group normalization layer, specifically divided into 32 groups, to extract semantic-related features and reduce the feature dimension, obtaining the corresponding semantic-related feature maps and

[0123] Input the source domain feature map Target domain feature map into the domain-related feature extractor Specifically: For and First, apply a convolutional layer with a 1x1 convolutional kernel, a stride of 1, and an output dimension of 256, and then connect a group normalization layer, specifically divided into 32 groups, to extract domain-related features and reduce the feature dimension, obtaining the corresponding domain-related feature maps and

[0124] Input the source domain feature map and the target domain feature map into the semantic-related feature extractor respectively Specifically: For and First, apply a convolutional layer with a 1x1 convolutional kernel, a stride of 1, and an output dimension of 256, and then connect a group normalization layer, specifically divided into 32 groups, to extract semantic-related features and reduce the feature dimension, obtaining the corresponding semantic-related feature maps and

[0125] Input the source domain feature map and the target domain feature map into the domain-related feature extractor respectively Specifically: For and First, apply a convolutional layer with a 1x1 convolutional kernel, a stride of 1, and an output dimension of 256, and then connect a group normalization layer, specifically divided into 32 groups, to extract domain-related features and reduce the feature dimension, obtaining the corresponding domain-related feature maps and

[0126] Input the source domain feature map and the target domain feature map into the semantic-related feature extractor respectively Specifically: For and First, apply a convolutional layer with a 3x3 convolutional kernel, a stride of 2, a padding of 1, and an output dimension of 256, and then connect a group normalization layer, specifically divided into 32 groups, to extract semantic-related features and reduce the feature dimension, obtaining the corresponding semantic-related feature maps and

[0127] Input the source domain feature map and the target domain feature map into the domain-related feature extractor respectively Specifically: For and First, apply a convolutional layer with a 3x3 convolutional kernel, a stride of 2, a padding of 1, and an output dimension of 256, and then connect a group normalization layer, specifically divided into 32 groups, to extract domain-related features and reduce the feature dimension, obtaining the corresponding domain-related feature maps and

[0128] (4.3) Through the multi-scale decoupled representation module, obtain the source domain multi-scale semantic-related feature map and the target domain multi-scale semantic-related feature map Source domain multi-scale domain-related feature map Target domain multi-scale domain-related feature map In this embodiment, the dimensions of these feature maps are all 256.

[0129] Specifically, in step 2, the specific process of obtaining the multi-scale domain-invariant semantic features is as follows:

[0130] (5.1) First, construct a domain discriminator network D, and its specific structure is:

[0131] First is a convolutional layer, using a 3x3 convolutional kernel, with a stride of 1 and padding of 1, and the output dimension is 256; then is a RELU activation function layer; then is a convolutional layer, using a 3x3 convolutional kernel, with a stride of 1 and padding of 1, and the output dimension is 256; then is a RELU activation function layer; then is a convolutional layer, using a 3x3 convolutional kernel, with a stride of 1 and padding of 1, and the output dimension is 256; then is a RELU activation function layer; then is a convolutional layer, using a 3x3 convolutional kernel, with a stride of 1 and padding of 1, and the output dimension is 2.

[0132] (5.2) Input the source domain multi-scale domain-related feature maps into the domain discriminator network D respectively. For each position in the feature map output by D, in this embodiment, the Softmax function is used for normalization processing to obtain the corresponding source domain target domain domain discrimination probability value, that is, on the whole, the source domain multi-scale domain-related domain discrimination feature map is obtained respectively represent the values at the (u, v) positions of the corresponding domain discrimination feature map;

[0133] Input the target domain multi-scale domain-related feature maps into the domain discriminator network D respectively. For each position in the feature map output by D, in this embodiment, the Softmax function is used for normalization processing to obtain the corresponding source domain target domain domain discrimination probability value, that is, on the whole, the target domain multi-scale domain-related domain discrimination feature map is obtained respectively represent the values at the (u, v) positions of the corresponding domain discrimination feature map;

[0134] Input the source domain multi-scale semantic-related feature maps into the domain discriminator network D respectively. For each position in the feature map output by D, in this embodiment, the Softmax function is used for normalization processing to obtain the corresponding source domain target domain domain discrimination probability value, that is, on the whole, the source domain multi-scale semantic-related domain discrimination feature map is obtained respectively represent the values at the (u, v) positions of the corresponding domain discrimination feature map;

[0135] Input the multi-scale semantic-related feature maps of the target domain into the domain discriminator network D respectively. For each position in the feature map output by D, in this embodiment, the Softmax function is used for normalization processing to obtain the domain discrimination probability values of the source domain and the target domain corresponding to the output, that is, the multi-scale semantic-related domain discrimination feature map of the target domain is obtained as a whole respectively represent the values at the positions (u, v) of the corresponding domain discrimination feature map

[0136] (5.3) Set the label of the source domain image feature map to 0 and the label of the target domain image feature map to 1

[0137] (5.4) Calculate the domain discrimination loss for the multi-scale domain-related domain discrimination feature map of the source domain and the target domain so that the domain discriminator network learns domain-related knowledge. The loss calculation is as follows

[0138]

[0139] where u i and v i respectively represent and width and height; L CE (,) represents the cross-entropy loss between the prediction and the label represents the sum of the domain discrimination losses of the i-th scale domain-related feature map

[0140] The mean value of the domain discrimination loss of the overall domain-related feature map L dis-domain , the calculation is as follows

[0141]

[0142] (5.5) Calculate the domain discrimination loss for the multi-scale semantic-related domain discrimination feature map of the source domain and the target domain The loss calculation is as follows

[0143]

[0144] where u i and v i respectively represent and width and height; L CE (,) represents the cross-entropy loss between the prediction and the label represents the sum of the domain discrimination losses of the i-th scale semantic-related feature map

[0145] Introduce a gradient reversal layer for adversarial training. The overall adversarial loss mean value and adversarial optimization logic are as follows

[0146]

[0147] Among them, L adv is the overall adversarial loss mean; for adversarial optimization logic, The purpose of is to deceive the discriminator, and the purpose of the domain discriminator D is to discriminate the domain source of the semantic feature. The adversarial process makes Gradually extract domain-invariant information from the source domain to the target domain, and obtain domain-invariant semantic features at multiple scales.

[0148] (5.6) The multi-scale domain-invariant semantic features of the source domain and the target domain after adversarial training are sent to the subsequent network structure. Since the domain discrimination network D has learned domain-related knowledge in step (5.4), It is necessary to dig deeper into the domain-independent components of the source domain and the target domain in order to deceive the domain discrimination network D, so that the domain-invariant semantic features can be better obtained. The overall loss of the multi-scale decoupled representation module to obtain multi-scale domain-invariant semantic features is denoted as L MDR , the specific calculation is as follows:

[0149] L MDR =L dis-domain +L adv

[0150] Specifically, the method of step 3 is as follows:

[0151] (6.1) Calculate the supervision loss between the source domain image object detection results and the corresponding annotations generated by the student model

[0152] (6.2) Calculate the loss L of the multi-scale decoupled representation module in the student model to obtain multi-scale domain invariant semantic features MDR ;

[0153] (6.3) Use the two calculated losses to train the student model and use the student model to update the teacher model;

[0154] (6.4) Repeat (6.1) to (6.3) until multiple rounds of training are completed.

[0155] Specifically, in step 4, a category-adaptive pseudo-label screening mechanism based on reinforcement learning is introduced to enable the model to obtain high-quality pseudo-labels and guide the model to perform multiple rounds of retraining. The specific process is as follows:

[0156] (7.1) Set the initial screening threshold δ for each type of object in domain adaptation target detection ij , δ ij Indicates that the threshold for selecting pseudo labels for the j-th object in the i-th round of training is δ ij , j ranges from 1 to N, where N is the number of object categories.

[0157] (7.2) Input the target domain image into the teacher model for prediction. Multiple prediction results are generated for each image, and each prediction result contains a confidence level and the corresponding coordinate bounding box

[0158] Among them, it means that in the prediction of the i-th round, the category of the l-th ranked prediction result is the j-th category, and the corresponding confidence level is They are respectively the x coordinate of the center point of the corresponding prediction box, the y coordinate of the center point, the width of the coordinate bounding box, and the height of the coordinate bounding box

[0159] (7.3) Screen the pseudo-labels generated by the teacher model. For the multiple pseudo-labels predicted for each image, the confidence levels greater than the corresponding category threshold δ ij of the predictions are considered reliable pseudo-labels and are retained. The coordinate bounding box information is used as the annotation of the corresponding target domain image, and the predictions with a prediction confidence level less than or equal to the corresponding category threshold are discarded

[0160] (7.4) Input the target domain image into the student model, and use the retained reliable pseudo-labels to supervise the object detection results generated by the student model for the target domain image

[0161] (7.5) Repeat (7.2) to (7.4) until the i-th round of training is completed. According to the performance of the model, with the idea of reinforcement learning, adaptively adjust the threshold for screening pseudo-labels of each category of objects in the next round, that is, δ i+1j .

[0162] (7.6) Train each subsequent round according to the logic of (7.2) to (7.5). Since the thresholds for screening pseudo-labels of each category of objects in each round are continuously updated with the idea of reinforcement learning, fully considering the specific state during training, high-quality pseudo-labels can be obtained to guide the model training until the end of multiple rounds of training

[0163] Furthermore, in step (7.5), according to the performance of the model, with the idea of reinforcement learning, adaptively adjust the threshold for screening pseudo-labels of each category of objects in the next round. The specific process is as follows

[0164] (8.1) Regard the teacher model as an intelligent agent, and the setting of the threshold for screening pseudo-labels for each category of objects is regarded as the strategy for the intelligent agent to select pseudo-labels

[0165] (8.2) Use the intelligent agent obtained after the previous (i - 1)-th round of training to predict the target domain image, and obtain the prediction result mAP50 i-1j, that is, the average precision predicted by the model for the j-th type of object in the (i - 1)-th round when the intersection over union threshold is 0.5.

[0166] Among them, for the initial round of guiding the model to perform multiple rounds of retraining in step 4, mAP50 i-1j is the prediction result obtained by the agent predicting the target domain image after the last round of the multiple rounds of training in step 3; for non-initial rounds, mAP50 i-1j is the prediction result obtained by the agent predicting the target domain image after the (i - 1)-th round of training before.

[0167] (8.3) Use the agent obtained from the just-completed i-th round of training to predict the target domain image, and obtain the prediction result mAP50 ij , that is, the average precision predicted by the model for the j-th type of object in the i-th round when the intersection over union threshold is 0.5.

[0168] (8.4) According to the prediction results of the (i - 1)-th round and the i-th round, calculate the reward reward of the j-th type of object in the current i-th round ij , specifically mAP50 ij and mAP50 i-1j The absolute value of the difference between them:

[0169] reward ij = |mAP50 ij - mAP50 i-1j |

[0170] Among them, the reward reward ij is used to determine whether to update the pseudo-label selection strategy and how to update the pseudo-label selection strategy. τ is a hyperparameter. When the absolute value of the difference exceeds τ, the strategy will be updated.

[0171] According to the corresponding category reward, determine the strategy of the agent to select pseudo-labels for the corresponding category, and obtain the threshold δ for screening pseudo-labels of the j-th type of object in the (i + 1)-th round i+1j ,

[0172]

[0173] Among them, α is a hyperparameter used to control the amplitude of the update. The larger α is, the greater the amplitude of the update. β is also a hyperparameter, which determines the influence degree of new information on the update of the threshold strategy.

[0174] This weighted update mechanism allows the system to gradually integrate new feedback, while utilizing existing experience, avoiding fluctuations caused by excessive updates, and ensuring that the threshold update remains in an appropriate state.

[0175] IU ijThe increment during the policy iteration of the j-th type of object in the i-th round, IU ij The size of ij is obtained by successively feeding the corresponding reward value into the tanh function and the log function and calculating the square root, and the formula is:

[0176]

[0177] For IU ij Regarding the positive or negative of ij , if the difference between mAP50 ij and mAP50 i-1j is positive, then IU ij is positive; if the difference of mAP50 is negative, then IU ij is negative.

[0178] Finally, limit the threshold within a reasonable range:

[0179] δ i+1j = min(max(δ i+1j , a), b)

[0180] where a and b are two hyperparameters, which are used to limit the lower and upper bounds of the threshold for screening the pseudo-labels of various types of objects respectively.

[0181] Specifically, in step 4, the DINO remote sensing image cross-domain object detection model is retrained for multiple rounds, and the specific process is as follows:

[0182] (9.1) Determine the threshold for screening the pseudo-labels of various types of objects in the current round according to the method in steps (8.1) to (8.4) with the idea of reinforcement learning;

[0183] (9.2) Use the teacher model in the current round to generate predictions, and screen the predictions according to the threshold for screening the pseudo-labels of various types of objects determined in step (9.1) to obtain reliable pseudo-labels;

[0184] (9.3) Calculate the supervised loss between the object detection results of the source domain images generated by the student model and the corresponding annotations

[0185] (9.4) Calculate the supervised loss between the object detection results of the target domain images generated by the student model and the corresponding reliable pseudo-labels

[0186] (9.5) Calculate the loss L of the multi-scale decoupled representation module in the student model to obtain multi-scale domain-invariant semantic features MDR ;

[0187] (9.6) Use the three calculated losses to train the student model and update the teacher model with the student model;

[0188] (9.7) Repeat steps (9.2) to (9.6) until the end of one round of training;

[0189] (9.8) Repeat steps (9.1) to (9.7), first adaptively adjust the threshold for pseudo-label screening of each category of objects in the next round with the idea of reinforcement learning, and then perform the next round of model training until the end of multiple rounds of retraining.

[0190] In summary, the cross-domain remote sensing image target detection method of the present invention realizes feature decoupling at multiple scales through a multi-scale decoupled representation module, promotes the domain invariance of semantically related features with the domain discrimination of domain-related features, reduces the impact of large differences in the data distribution of remote sensing images, obtains domain-invariant features, and improves the performance of cross-domain target detection. At the same time, through the category-adaptive pseudo-label screening mechanism based on reinforcement learning, according to the training process, with the idea of reinforcement learning, adaptively adjust the strategy for screening pseudo-labels for each type of target, obtain high-quality and low-noise pseudo-labels, enable the model to learn effective target domain information, and improve the performance of cross-domain target detection.

[0191] It should be understood that although this specification is described according to the embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. The above is only the preferred embodiment of the present invention. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art. For those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A cross-domain remote sensing image target detection method based on multi-scale decoupled representation and reinforcement learning, characterized in that: The steps include: Step 1: Obtain a cross-domain remote sensing image target detection dataset and construct a training set; Step 2: Build a cross-domain object detection model for DINO remote sensing images, obtain the initial multi-scale feature map through the backbone network Resnet50, obtain multi-scale domain-invariant semantic features through the multi-scale decoupling representation module, and reduce the domain deviation between the source domain and the target domain; Step 3: Use the training set to perform multiple rounds of training on the DINO remote sensing image cross-domain object detection model built in step 2; Step 4: Based on the results of multiple rounds of training in step 3, the DINO remote sensing image cross-domain target detection model is retrained, and a category-adaptive pseudo-label screening mechanism based on reinforcement learning is introduced to obtain high-quality pseudo-labels in the target domain, guide the model to conduct multiple rounds of retraining, learn effective target domain information, improve adaptability to the target domain, and obtain the optimized DINO remote sensing image cross-domain target detection model; Step 5: Use the target domain test set to perform detection on the optimized DINO remote sensing image cross-domain target detection model obtained in step 4.

2. The cross-domain remote sensing image target detection method based on multi-scale decoupled representation and reinforcement learning according to claim 1 is characterized in that: The method for step 1 is as follows: Step 101, obtain the cross-domain remote sensing image target detection datasets xView and DOTA, and perform category screening; Step 102: cropping the image to cut the original cross-domain remote sensing image into an image of a set size; Step 103: construct a training set, use the cropped images for model training, and perform data enhancement during the training process, including random rotation, flipping, and random photometric distortion.

3. The cross-domain remote sensing image target detection method based on multi-scale decoupled representation and reinforcement learning according to claim 1 is characterized in that: Step 2 describes the construction of the DINO remote sensing image cross-domain object detection model. The process is as follows: Step 201, build a student model, based on the DINO framework for improvement, including: Resnet50 backbone network, multi-scale decoupled representation module, transformer encoder decoder, detection head; Step 202: copy the student model as the teacher model, initialize the parameters to be consistent with the student model, and update the parameters through the student model using EMA. The update formula is as follows: Among them, t represents the current iteration round, α represents the update weight, represents the parameters of the teacher model at the tth iteration, represents the parameters of the teacher model at the t-1th iteration, Represents the parameters of the student model at the tth iteration.

4. The cross-domain remote sensing image target detection method based on multi-scale decoupled representation and reinforcement learning according to claim 3 is characterized in that: In step 201, the multi-scale decoupling characterization module is constructed as follows: Step 201.1: transform the source domain image x s and the target domain image x t Input the student model backbone network Resnet50 respectively to obtain feature maps of different scales, which are represented as and in, Corresponding to the features after the initial convolutional layer and pooling layer of the source domain image, Features corresponding to subsequent layers of the source domain image; Corresponding to the features after the initial convolutional layer and pooling layer of the target domain image, Features corresponding to subsequent layers of the target domain image; Step 201.2: The source domain feature map and target domain feature map Input into the semantic related feature extractor respectively right and Apply a convolution layer respectively, use a 1x1 convolution kernel, a step size of 1, an output dimension of 256, and then a group normalization layer to extract semantically relevant features and reduce the feature dimension to obtain the corresponding semantically relevant feature map and The source domain feature map and target domain feature map Input to the domain-dependent feature extractor right and Apply a convolution layer respectively, use a 1x1 convolution kernel, a step size of 1, an output dimension of 256, and then a group normalization layer to extract domain-related features and reduce the feature dimension to obtain the corresponding domain-related feature map and The source domain feature map and target domain feature map Input into the semantic related feature extractor respectively right and Apply a convolution layer respectively, use a 1x1 convolution kernel, a step size of 1, an output dimension of 256, and then a group normalization layer to extract semantically relevant features and reduce the feature dimension to obtain the corresponding semantically relevant feature map and The source domain feature map and target domain feature map Input to the domain-dependent feature extractor right and Apply a convolution layer respectively, use a 1x1 convolution kernel, a step size of 1, an output dimension of 256, and then a group normalization layer to extract domain-related features and reduce the feature dimension to obtain the corresponding domain-related feature map and The source domain feature map and target domain feature map Input into the semantic related feature extractor respectively right and Apply a convolution layer respectively, use a 1x1 convolution kernel, a step size of 1, an output dimension of 256, and then a group normalization layer to extract semantically relevant features and reduce the feature dimension to obtain the corresponding semantically relevant feature map and The source domain feature map and target domain feature map Input to the domain-dependent feature extractor right and Apply a convolution layer respectively, use a 1x1 convolution kernel, a step size of 1, an output dimension of 256, and then a group normalization layer to extract domain-related features and reduce the feature dimension to obtain the corresponding domain-related feature map and The source domain feature map and target domain feature map Input into the semantic related feature extractor respectively right and Apply a convolution layer respectively, use a 3x3 convolution kernel, a step size of 2, a padding of 1, an output dimension of 256, and then a group normalization layer to extract semantically relevant features and reduce the feature dimension to obtain the corresponding semantically relevant feature map and The source domain feature map and target domain feature map Input to the domain-dependent feature extractor right and Apply a convolution layer respectively, use a 3x3 convolution kernel, a step size of 2, a padding of 1, an output dimension of 256, and then a group normalization layer to extract domain-related features and reduce the feature dimension to obtain the corresponding domain-related feature map and Step 201.3: After the multi-scale decoupling representation module, a multi-scale semantically related feature map of the source domain is obtained. Multi-scale semantically relevant feature maps in the target domain Source domain multi-scale domain related feature map Multi-scale domain-related feature maps in the target domain 5. The cross-domain remote sensing image target detection method based on multi-scale decoupled representation and reinforcement learning according to claim 4 is characterized in that: Step 2 is to obtain multi-scale domain invariant semantic features. The process is as follows: Step 211, construct a domain discriminator network D, including several convolutional layers and RELU activation function layers; Step 212: multi-scale domain related feature maps of source domain and target domain Multi-scale semantic correlation feature map of source domain and target domain They are input into the domain discriminator network D respectively, and the Softmax function is used to normalize each position in the feature map output by D to obtain the corresponding output source domain target domain domain discrimination probability value, and the source domain target domain multi-scale domain related domain discrimination feature map is obtained as a whole Multi-scale semantically related domain discriminant feature map of source domain and target domain Step 213, setting the source domain image label to 0 and the target domain image label to 1; Step 214: discriminate the source domain, target domain, multi-scale domain, and related domain feature maps Calculate the domain discrimination loss so that the domain discrimination network can learn domain-related knowledge. The loss is calculated as follows: in, Respectively and The value corresponding to the (u,v) position, u i 、v i Respectively and Width and height; L CE (,) represents the cross entropy loss between prediction and label, represents the domain discrimination loss and the domain-related feature map of the i-th scale domain; The mean domain discrimination loss of the overall domain-related feature map L dis-domain , calculated as follows: Step 215: Discriminate feature maps of multi-scale semantically related domains in the source domain and target domain Calculate the domain discrimination loss. The loss is calculated as follows: in, Respectively and The value corresponding to the (u,v) position, u i 、v i Respectively and Width and height; L CE (,) represents the cross entropy loss between prediction and label, represents the domain discrimination loss and the sum of the semantically relevant feature maps at the i-th scale; The gradient reversal layer is introduced for adversarial training. The overall adversarial loss mean and adversarial optimization logic are as follows: Among them, L adv is the overall adversarial loss mean; for adversarial optimization logic, The purpose of is to deceive the discriminator, and the purpose of the domain discriminator D is to discriminate the domain source of the semantic feature. The adversarial process makes Gradually extract domain-invariant information from the source domain to the target domain, and obtain domain-invariant semantic features at multiple scales; Step 216: Send the multi-scale domain-invariant semantic features of the source domain and target domain after adversarial training to the subsequent network structure. Since the domain discrimination network D has learned domain-related knowledge in step 214, It is necessary to dig deeper into the domain-independent components of the source domain and the target domain in order to deceive the domain discrimination network D and obtain domain-invariant semantic features; The overall loss of the multi-scale decoupled representation module to obtain multi-scale domain invariant semantic features is denoted as L MDR , calculated as follows: L MDR =L dis-domain +L adv .

6. The cross-domain remote sensing image target detection method based on multi-scale decoupled representation and reinforcement learning according to claim 5 is characterized in that: The domain discriminator network D in step 211 has the following structure: The first convolutional layer uses a 3x3 convolution kernel, a stride of 1, a padding of 1, and an output dimension of 256; The first RELU activation function layer; The second convolutional layer uses a 3x3 convolution kernel, a stride of 1, a padding of 1, and an output dimension of 256; The second RELU activation function layer; The third convolutional layer uses a 3x3 convolution kernel, a stride of 1, a padding of 1, and an output dimension of 256; The third RELU activation function layer; The fourth convolutional layer uses a 3x3 convolution kernel, a stride of 1, a padding of 1, and an output dimension of 2.

7. The cross-domain remote sensing image target detection method based on multi-scale decoupled representation and reinforcement learning according to claim 5 is characterized in that: The method for multiple rounds of training described in step 3 is as follows: Step 301: Calculate the supervision loss between the source domain image target detection result generated by the student model and the corresponding annotation Step 302: Calculate the loss L of the multi-scale decoupled representation module of the student model to obtain multi-scale domain invariant semantic features MDR ; Step 303: Use the calculated source domain supervision loss and the loss L for obtaining multi-scale domain invariant semantic features MDR Train the student model and use it to update the teacher model; Step 304: Repeat steps 301 to 303 until multiple rounds of training are completed.

8. The cross-domain remote sensing image target detection method based on multi-scale decoupled representation and reinforcement learning according to claim 7 is characterized in that: In step 4, a category-adaptive pseudo-label screening mechanism based on reinforcement learning is introduced to obtain high-quality pseudo-labels in the target domain and guide the model to perform multiple rounds of retraining. The process is as follows: Step 401: Set an initial screening threshold δ for each type of object in domain adaptation target detection ij , δ ij represents the threshold for filtering pseudo labels for the j-th class of objects in the i-th round of training in the initial round of multiple rounds of retraining, j = 1, 2, ..., N, where N is the number of object categories; Step 402: Input the target domain image into the teacher model for prediction. Each image generates multiple prediction results, each of which contains a confidence level. And the corresponding coordinate frame in, Indicates that in the prediction of the i-th round, the category of the prediction result of the lth position is the j-th category, and the corresponding confidence is They are the x-coordinate of the center point of the corresponding prediction box, the y-coordinate of the center point, the width of the coordinate box, and the height of the coordinate box; Step 403: Filter the pseudo labels generated by the teacher model, and retain the confidence level for multiple pseudo labels predicted for each image. Greater than the corresponding category threshold δ ij The prediction of is used as a reliable pseudo label, and the coordinate box information is used as the annotation of the corresponding target domain image; Step 404: input the target domain image into the student model, and supervise the target detection result generated by the student model for the target domain image through the retained reliable pseudo labels; Step 405: Repeat steps 402 to 404 until the i-th round of training is completed, and use the reinforcement learning method to adaptively adjust the threshold δ for the next round of pseudo-label screening of each category of objects. i+1j ; Step 406: Perform multiple rounds of training according to the logic of steps 402 to 405 until the multiple rounds of training are completed.

9. The cross-domain remote sensing image target detection method based on multi-scale decoupled representation and reinforcement learning according to claim 8 is characterized in that: Step 405 uses the reinforcement learning method to adaptively adjust the threshold δ for the next round of pseudo-label screening of each category of objects. i+1j , the process is as follows: Step 405.1, taking the teacher model as the intelligent agent, and setting the threshold for screening pseudo labels for each category of objects as the strategy for the intelligent agent to select pseudo labels; Step 405.2: Use the agent obtained after the i-1th round of training to predict the target domain image and obtain the prediction result mAP50 i-1j ; Step 405.3: Use the agent obtained from the i-th round of training to predict the target domain image and obtain the prediction result mAP50 ij ; Step 405.4: Calculate the reward for the j-th object in the current round i based on the prediction results of the i-1th round and the i-th round. ij , which is used to decide whether to update the pseudo label selection strategy and how to update the pseudo label selection strategy: reward ij =|mAP50 ij -mAP50 i-1j | According to the corresponding category reward, the agent determines the strategy for selecting pseudo labels for the corresponding category, and obtains the threshold δ for filtering pseudo labels for the j-th category object in the i+1th round i+1j , the formula is: Among them, α and β are hyperparameters. α is used to control the update amplitude. The larger α is, the larger the update amplitude is. β determines the impact of new information on the threshold strategy update; IU ij is the increment of the strategy iteration of the j-th object in the i-th round; τ is a hyperparameter. ij When τ is exceeded, the policy will be updated.

10. The cross-domain remote sensing image target detection method based on multi-scale decoupled representation and reinforcement learning according to claim 8, characterized in that: Step 4 described in the DINO remote sensing image cross-domain object detection model for multiple rounds of retraining, the process is as follows: Step 411: Determine the pseudo-label screening thresholds for various objects in the current round using the reinforcement learning method according to the method of step 405; Step 412: Use the teacher model of the current round to generate predictions, and filter the predictions according to the pseudo-label filtering threshold determined in step 411 to obtain reliable pseudo-labels. Step 413: Calculate the supervision loss between the source domain image target detection result generated by the student model and the corresponding annotation Step 414: Calculate the supervision loss between the target domain image object detection result generated by the student model and the corresponding reliable pseudo label Step 415: Calculate the loss L of the multi-scale decoupled representation module in the student model to obtain multi-scale domain invariant semantic features MDR ; Step 416: Use loss L MDR Train the student model and use it to update the teacher model; Step 417, repeating steps 412 to 416 until one round of training is completed; Step 418, repeating steps 411 to 417, first adaptively adjusting the threshold of the next round of pseudo-label screening for each category of objects with the idea of ​​reinforcement learning, and then performing the next round of model training. After multiple rounds of training, an optimized DINO remote sensing image cross-domain target detection model is obtained.

Citation Information

Cited By

  • Remote sensing target detection method and device based on multi-level knowledge distillation

    CN120997489A

  • A remote sensing target detection method and device based on multi-level knowledge distillation

    CN120997489B

  • Pseudo tag generation method based on conformal prediction and application thereof in semantic segmentation

    CN121053488A

  • Road defect target detection method based on domain adaptation

    CN121582690A