Semi-supervised remote sensing target detection method and system based on spatial resolution guidance, medium and equipment
By introducing a spatial resolution-guided semi-supervised target detection method into remote sensing images and utilizing a dynamic weight mechanism to stabilize the training process, the problems of large pseudo-label noise and model instability in remote sensing images are solved, thereby improving the accuracy and robustness of remote sensing image detection.
Patent Information
- Application Number
- CN202511089180.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-05
- Publication Date
- 2025-11-25
AI Technical Summary
Existing semi-supervised target detection algorithms suffer from problems such as high noise in pseudo-label generation and unstable model performance in the field of remote sensing images. In particular, it is difficult to effectively utilize the spatial resolution information of remote sensing images for training when there are remote sensing targets with high aspect ratio, small and dense size, arbitrary angles and large differences in spatial resolution.
By constructing a semi-supervised remote sensing target detection method guided by spatial resolution, this method utilizes the temporal distribution information of spatial resolution in remote sensing images and employs a dynamic weighting mechanism to guide the semi-supervised training process. This includes ground sampling distance standardization, online weighted update mechanism, category-level pseudo-label ground sampling distance typicality weighting, and image-level unsupervised loss weighting based on KL divergence, thereby reducing pseudo-label noise and stabilizing the training process.
It significantly improves the model's detection performance and robustness at low annotation rates, and enhances the accuracy and stability of remote sensing image detection, especially in remote sensing images with large-scale spatial resolution variations.
Smart Images

Figure CN121010889A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision and remote sensing image processing, and in particular to a semi-supervised remote sensing target detection method, system, medium, and device based on spatial resolution guidance. Background Technology
[0002] Key target identification and localization in remote sensing images is a core task in remote sensing image processing. Although deep learning has significantly improved the performance of supervised target detection algorithms in this field, it heavily relies on expensive and time-consuming manual annotation, especially for targets at arbitrary angles from overhead views, where annotation costs remain prohibitively high. To address the problem of limited labeled data restricting model performance, semi-supervised target detection algorithms have emerged. These algorithms can perform supervised training using limited labeled data while simultaneously using pseudo-label generation techniques combined with consistency training strategies to perform self-supervised training on unlabeled image data. This effectively utilizes the generalization model's detection capabilities across a large number of unlabeled images, achieving performance improvements without requiring further annotation input. However, existing semi-supervised target detection algorithms still face the following problems when applied to remote sensing image processing:
[0003] (1) Remote sensing targets generally have unique problems such as high aspect ratio, small and dense size, and arbitrary angle, which pose challenges to the direct application of semi-supervised methods in general fields.
[0004] (2) The spatial resolution distribution within remote sensing images is wide, and this difference leads to huge size differences among similar targets, affecting the generation of pseudo-labels.
[0005] (3) During semi-supervised training, the influence of labeled and unlabeled images on the teacher model changes over time, and the corresponding differences in the distribution of training data will also have a dynamic impact on the performance of the teacher model.
[0006] In summary, existing general-purpose semi-supervised target detection algorithms need targeted improvements based on the characteristics of remote sensing images, particularly in utilizing the spatial resolution information of these images to reduce noise from generated false labels. How to effectively leverage the spatial resolution distribution information of remote sensing images to guide the semi-supervised training process has become an increasingly pressing research topic. Summary of the Invention
[0007] To address the aforementioned problems, the present invention aims to provide a semi-supervised remote sensing target detection method, system, medium, and device based on spatial resolution guidance. This method utilizes the temporal distribution information of spatial resolution in remote sensing images to guide the semi-supervised training process using a dynamic weighting mechanism, thereby reducing the impact of false label noise and stabilizing the semi-supervised training process under low labeling rates, thus improving the performance of the final model.
[0008] To achieve the above objectives, in a first aspect, the technical solution adopted by this invention is as follows: a semi-supervised remote sensing target detection method based on spatial resolution guidance. This method is built on a teacher-student semi-supervised learning framework and includes: acquiring remote sensing image data containing labeled and unlabeled datasets, as well as corresponding ground sampling distance metadata; initializing teacher and student models, and initializing ground sampling distance distribution parameters for each target category to describe the labeled and unlabeled data; performing logarithmic transformation on the original ground sampling distance values of each input remote sensing image based on ground sampling distance standardization; maintaining the ground sampling distance distribution of each target category in the labeled and unlabeled datasets through an online weighted update mechanism; dynamically adjusting the pseudo-label weights of positive samples in unlabeled images based on a category-level pseudo-label ground sampling distance typicality weighting mechanism; dynamically adjusting the total loss contribution of unlabeled images according to the difference in ground sampling distance distribution between labeled and unlabeled data based on an image-level unsupervised loss weighting mechanism based on KL divergence; and integrating the weighting mechanism into the semi-supervised training process to train the final target detection model for remote sensing target detection.
[0009] Furthermore, the ground sampling distance is standardized as follows:
[0010] x = log2(g ori / s),
[0011] In the formula, g ori represents the original ground sampling distance value, s represents the scaling factor in data augmentation, and x represents the standardized ground sampling distance value.
[0012] Furthermore, the online weighted update mechanism includes:
[0013] Maintain cumulative weights and W for each category c. c Weighted mean μ c and the weighted sum of squared differences M c ;
[0014] When processing an n-valued array containing class c c When analyzing an image of an object, its standardized ground sampling distance is x. A time decay factor α is introduced, and the parameter update rule is as follows:
[0015] Update cumulative weights: Maintain the cumulative weight sum for the updated category c. Maintain the cumulative weight sum for category c before the update;
[0016] Calculate the difference δ from the old mean: This is the weighted mean before the update;
[0017] Update the weighted mean: This is the updated weighted mean;
[0018] Update the weighted sum of squared differences: This is the updated weighted sum of squared differences. This is the cumulative weighted squared difference value before the update;
[0019] Calculate the variance:
[0020] Furthermore, the category-level pseudo-label ground sampling distance typicality weighting mechanism is implemented through the following steps:
[0021] For a positive sample in an unlabeled image that is assigned a pseudo-label class c, calculate the typicality confidence score between the standardized ground sampling distance value x of the image and the ground sampling distance distribution parameter of class c learned from labeled data;
[0022] The original sample-level loss weights w old Multiplying the calculated ground sampling distance typicality confidence score by the pseudo-label weight yields the final category-level pseudo-label weight.
[0023] Furthermore, the image-level unsupervised loss weighting mechanism based on KL divergence is implemented through the following steps:
[0024] For each pseudo-label category c in the unlabeled image, calculate the KL divergence between its ground sampling distance distribution in labeled data and its ground sampling distance distribution in unlabeled data;
[0025] Log-normalize the KL divergence for each category;
[0026] For an unlabeled image containing n pseudo-labels, calculate its average normalized KL divergence, and calculate the image-level unsupervised loss weights based on the average normalized KL divergence using an inverse proportional function.
[0027] Furthermore, semi-supervised training includes:
[0028] A warm-up phase was set up, where only labeled data was used to train the teacher model;
[0029] After the warm-up phase, the teacher model generates pseudo-labels for the unlabeled images;
[0030] The typicality weights of the pseudo-label ground sampling distance are applied to the loss calculation of the R-CNN stage for unlabeled images;
[0031] The image-level KL divergence weights are applied to adjust the total unsupervised loss of the unlabeled image, and the total loss is calculated.
[0032] The student model parameters are updated via backpropagation, and the teacher model parameters are updated via exponential moving average.
[0033] Furthermore, the total loss is:
[0034]
[0035] in, Total loss; Standard monitoring loss; N u This represents the number of unlabeled images in the batch. These are the image-level unsupervised loss weights; The loss for the R-CNN stage of unlabeled images; For unlabeled images, the RPN stage loss is used. Let k be the unlabeled image.
[0036] Secondly, the technical solution adopted by this invention is as follows: a semi-supervised remote sensing target detection system based on spatial resolution guidance, under a teacher-student semi-supervised learning framework, comprising: a data acquisition module, which acquires remote sensing image data containing labeled and unlabeled datasets, and corresponding ground sampling distance metadata; an initialization module, which initializes the teacher model and the student model, and initializes ground sampling distance distribution parameters for each target category to describe the labeled and unlabeled data; and a GSD processing module, which performs logarithmic transformation on the original ground sampling distance values of each input remote sensing image based on ground sampling distance normalization processing. The system comprises: a processing module; an update module, which maintains the ground sampling distance distribution of each target category in both labeled and unlabeled datasets through an online weighted update mechanism; a weighting module, which dynamically adjusts the pseudo-label weights of positive samples in unlabeled images based on a category-level pseudo-label ground sampling distance typicality weighting mechanism; and an image-level unsupervised loss weighting mechanism based on KL divergence, which dynamically adjusts the total loss contribution of unlabeled images according to the difference in ground sampling distance distribution between labeled and unlabeled data. Finally, a detection module integrates the weighting mechanism into the semi-supervised training process to train the final target detection model for remote sensing target detection.
[0037] Thirdly, the technical solution adopted by the present invention is: a computer-readable storage medium for storing one or more programs, wherein the one or more programs include instructions, which, when executed by a computing device, cause the computing device to perform any of the methods described above.
[0038] Fourthly, the technical solution adopted by the present invention is: a computing device comprising: one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for performing any of the methods described above.
[0039] The present invention has the following advantages due to the adoption of the above technical solutions:
[0040] This invention effectively integrates spatial resolution information from remote sensing images into the entire training process of semi-supervised object detection. Through online modeling and utilization of the temporal distribution of Geometric Target Score (GSD), and a dual dynamic weighting mechanism at both the category and image levels, it significantly reduces pseudo-label noise caused by GSD distribution mismatch and data temporal sensitivity, stabilizing the training process. This improves the model's detection performance and robustness when processing remote sensing images with large spatial resolution variations, especially in low-labeling scenarios. This invention achieves effective guidance for the semi-supervised learning process with minimal additional computational overhead. Attached Figure Description
[0041] Figure 1 This is a flowchart of a semi-supervised remote sensing target detection method based on spatial resolution guidance in an embodiment of the present invention;
[0042] Figure 2 This is a schematic diagram of the module composition structure of the semi-supervised remote sensing target detection method based on spatial resolution guidance in an embodiment of the present invention. Detailed Implementation
[0043] Remote sensing target detection has significant applications in urban planning, resource monitoring, and disaster assessment. Semi-supervised learning methods have shown great potential in remote sensing image target detection due to their ability to effectively utilize large amounts of unlabeled data. However, the spatial resolution (GSD) of remote sensing images varies significantly, with target scale and sharpness characteristics changing drastically under different GSDs. This poses challenges to the quality of pseudo-labels and the stability of the model in semi-supervised learning. Existing semi-supervised target detection methods typically ignore the potential impact of GSD information on the learning process, resulting in limited performance on data with large GSD variations.
[0044] To address the aforementioned issues, this invention provides a semi-supervised remote sensing target detection method, system, medium, and device based on spatial resolution guidance. It introduces and utilizes ground sample distance (GSD) metadata of the image, and dynamically adjusts the weights of pseudo-labels and loss functions to guide the semi-supervised training process, thereby improving model performance and pseudo-label quality.
[0045] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the described embodiments of the present invention are within the scope of protection of the present invention.
[0046] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0047] In one embodiment of the present invention, a semi-supervised remote sensing target detection method based on spatial resolution guidance is provided. This method aims to improve the learning robustness and detection accuracy of the model on data with different spatial resolutions by dynamically adjusting the weights of pseudo-labels and the loss function using GSD metadata of the image. In this embodiment, the method is built upon a teacher-student semi-supervised learning framework, such as... Figure 1 As shown, the method includes the following steps:
[0048] 1) Obtain remote sensing image data containing labeled and unlabeled datasets, as well as the corresponding GSD metadata;
[0049] 2) Initialize the teacher and student models, and initialize the ground sampling distance distribution parameters for each target category to describe the labeled and unlabeled data;
[0050] 3) Based on the ground sampling distance standardization process, the original ground sampling distance values of each input remote sensing image are logarithmically transformed;
[0051] 4) By using an online weighted update mechanism, maintain the ground sampling distance distribution of each target category in the labeled and unlabeled datasets;
[0052] 5) Based on the class-level pseudo-label ground sampling distance typicality weighting mechanism, the weight of the pseudo-label of positive samples in the unlabeled image is dynamically adjusted; based on the KL divergence image-level unsupervised loss weighting mechanism, the total loss contribution of the unlabeled image is dynamically adjusted according to the difference in ground sampling distance distribution between labeled and unlabeled data.
[0053] 6) Integrate the weighting mechanism into the semi-supervised training process to train the final target detection model for remote sensing target detection.
[0054] In step 3) above, the ground sampling distance standardization process is as follows:
[0055] x = log2(g ori / s),
[0056] In the formula, g ori represents the original ground sampling distance value, s represents the scaling factor in data augmentation, and x represents the standardized ground sampling distance value.
[0057] Specifically, GSD data standardization processing: For each input image (labeled or unlabeled), obtain its original GSD value (g ori Considering the potential scaling transformation (scaling factor s) during data augmentation, the adjusted GSD value g is first calculated. adj =g ori / s. Then, a base-2 logarithmic transformation is performed on the adjusted GSD value to obtain the standardized GSD value x = log2(g adj This transformation can better reflect the relative characteristics of scale changes in remote sensing images.
[0058] In step 4) above, class-level GSD Gaussian distribution modeling and online weighted updates are performed. Specifically, for each target class c, its GSD distribution parameters in the labeled and unlabeled datasets are maintained respectively. It is assumed that the GSD distribution follows a Gaussian distribution. This invention employs a weighted update mechanism based on the Welford online algorithm and incorporating the concept of exponential moving average (EMA) to track the dynamic changes of the GSD distribution in real time. Specifically, three parameters are maintained for each class c: cumulative weight and W. c (Equivalent to sample counting with time decay), weighted mean μ c And the weighted sum of squared differences M used to calculate variance. c .
[0059] In this embodiment, when processing an n-type image containing category c... c When an image of an object is given, its normalized GSD value is α.
[0060] Introducing a time decay factor α∈(0,1), the online weighted update mechanism includes the following steps:
[0061] 4.1) Maintain cumulative weights and W for each category c. c Weighted mean μ c and the weighted sum of squared differences M c ;
[0062] 4.2) When processing an n-valued array containing class c c When analyzing an image of an object, its standardized ground sampling distance is x. A time decay factor α is introduced, and the parameter update rule is as follows:
[0063] 4.2.1) Update the cumulative weight sum: Maintain the cumulative weight sum for the updated category c. Maintain the cumulative weight sum for category c before the update;
[0064] 4.2.2) Calculate the difference δ between the mean and the old mean: This is the weighted mean before the update;
[0065] 4.2.3) Update the weighted mean: This is the updated weighted mean;
[0066] 4.2.4) Update the weighted sum of squared differences: This is the updated weighted sum of squared differences. This is the cumulative weighted squared difference value before the update;
[0067] 4.2.5) Calculate the variance:
[0068] For labeled data, its GSD distribution parameter (denoted as...) The GSD distribution parameter is updated using the true class labels from the start of training. For unlabeled data, the GSD distribution parameter (denoted as...) is... After the burn-in phase of semi-supervised learning, the pseudo-label categories generated by the teacher model are updated as unlabeled data enters the semi-supervised training process.
[0069] In step 5) above, the category-level pseudo-label GSD typicality weighting mechanism dynamically adjusts the weight of a single positive sample pseudo-label in the loss calculation based on the degree of matching between the GSD of the unlabeled image and the "typical" GSD of the category to which its pseudo-label belongs.
[0070] Specifically, the category-level pseudo-label ground sampling distance typicality weighting mechanism is implemented through the following steps:
[0071] 5.1.1) Typicality Confidence Calculation: For a positive sample in an unlabeled image assigned a pseudo-label class c, calculate the standardized ground sampling distance value x of the image and the ground sampling distance distribution parameter of class c learned from labeled data. and Typicality confidence scores between )
[0072]
[0073] To prevent the confidence score calculation from being too unstable due to excessively small variance, the variance learned from the labeled data is... Perform threshold processing:
[0074]
[0075] 5.1.2) Pseudo-label weight adjustment: Adjust the original sample-level loss weights w old (For example, obtained from teacher model confidence or other strategies) is multiplied by the calculated ground sampling distance typicality confidence score, and a minimum weight lower bound of 0.1 is set to obtain the final category-level pseudo-label weights:
[0076]
[0077] In this embodiment, this w cpl The weights are used to adjust the contribution of the pseudo-label of the positive sample in the unlabeled image to the classification loss and regression loss in the subsequent R-CNN stage.
[0078] In step 5) above, the image-level unsupervised loss weighting mechanism based on KL divergence dynamically adjusts the contribution of the entire unlabeled image to the overall unsupervised loss by measuring the difference between labeled and unlabeled data regarding the GSD distribution of each category.
[0079] Specifically, the image-level unsupervised loss weighting mechanism based on KL divergence is implemented through the following steps:
[0080] 5.2.1) Category-level KL divergence calculation: For each pseudo-labeled category c in the unlabeled image, calculate its ground sampling distance distribution in the labeled data. Ground sampling distance distribution in unlabeled data KL divergence between them;
[0081]
[0082] The variance was also thresholded.
[0083] 5.2.2) Image-level average normalized KL divergence calculation: Log-normalize the KL divergence for each category:
[0084] D norm (c)=ln(1+D kl (c));
[0085] 5.2.3) For a dataset containing n pseudo-labels (of category c1, ..., c...),n Calculate the average normalized KL divergence D of the unlabeled image I. avg (I) and the image-level unsupervised loss weights are calculated using an inverse proportional function based on the average normalized KL divergence.
[0086] The average normalized KL divergence is:
[0087]
[0088] The image-level unsupervised loss weights are:
[0089]
[0090] Here, γ is a hyperparameter controlling sensitivity. If the image has no effective pseudo-labels or is in the early stages of GSD distribution capture (e.g., before sufficient sample updates), the preset default weight w is used. default .
[0091] In this embodiment, this w unsup (I) The weights are used to multiply the total unsupervised loss of the unlabeled image I (including the RPN loss and the class-weighted R-CNN loss), that is:
[0092]
[0093] In step 6) above, a teacher-student model structure is adopted, and the detector uses Faster R-CNN (adapted to arbitrary orientation bounding boxes OBB). For unlabeled images, pseudo-labels are first generated by the teacher model. Then, the weight w of each positive sample pseudo-label is calculated using the aforementioned "category-level pseudo-label GSD typicality weighting mechanism". cpl Then, the R-CNN loss is adjusted. Next, the image-level weights w are calculated using the "image-level unsupervised loss weighting mechanism based on KL divergence". unsup And adjust the total unsupervised loss of the image.
[0094] Specifically, semi-supervised training includes the following steps:
[0095] 6.1) Set up a warm-up phase, using only labeled data to train the teacher model; among these phases, for labeled images, calculate the standard supervised loss.
[0096] 6.2) After the warm-up phase, the teacher model generates pseudo-labels for the unlabeled images.
[0097] 6.3) Apply the class-level pseudo-label ground sampling distance typicality weights to the R-CNN stage loss calculation for unlabeled images;
[0098] 6.4) Apply the image-level KL divergence weights to adjust the total unsupervised loss of the unlabeled image and calculate the total loss (i.e., the total training loss of the student model);
[0099] 6.5) Update the student model parameters through backpropagation and update the teacher model parameters through exponential moving average (EMA).
[0100] 6.6) Iterative update of GSD distribution parameters: After each batch of training, the corresponding class-level GSD distribution parameters are updated according to the labeled images and their true labels, unlabeled images and their pseudo labels, and their respective GSD values in the current batch, following an online weighted update mechanism.
[0101] In step 6.4) above, the total loss is:
[0102]
[0103] In the formula, Total loss; Standard monitoring loss; N u This represents the number of unlabeled images in the batch. These are the image-level unsupervised loss weights; The loss for the R-CNN stage of unlabeled images; For unlabeled images, the RPN stage loss is used. Let k be the unlabeled image.
[0104] Among them, the R-CNN classification loss for unlabeled image I and regression loss The formula for calculating the stage loss of R-CNN for unlabeled images is:
[0105]
[0106] In the formula, It is the set of positive sample proposals, c i It is the pseudo-label category of the i-th proposal, w neg It is the negative sample weight, l cls and l reg It is a single-sample loss.
[0107] The formula for calculating the total unsupervised loss of unlabeled images is as follows:
[0108]
[0109] In this embodiment, the target detection model can be adapted to any designed detector, whether it is a single-stage, dual-stage, or DETR-based detection head.
[0110] The following embodiment provides a further detailed description of the corresponding technical solution through an exemplary embodiment of the present invention. The spatial resolution-guided semi-supervised remote sensing target detection method in this embodiment includes the following steps:
[0111] Step 1: Initialization, the specific implementation process is as follows:
[0112] S101. Initialize the network parameters of the Teacher Model and Student Model. These two models typically have the same network structure, for example, using a Faster R-CNN OBB (Oriented Bounding Box) with a Feature Pyramid Network (FPN) as the base detector and a ResNet-50 as the backbone network. In the early stages of training, the parameters of the Student Model and Teacher Model can be randomly initialized, or the Teacher Model parameters can be initialized using an exponential moving average (EMA) of the Student Model parameters.
[0113] S102. For each target category c in the dataset (17 categories in the DOTA-v2.0 dataset excluding the "airport" category), initialize the GSD distribution parameters for both labeled and unlabeled data. Specifically, this includes:
[0114] a) Cumulative weights and The initial value is usually set to 0 or a very small value.
[0115] b) Weighted mean The initial value is usually set to 0.
[0116] c) Weighted sum of squared differences The initial value is usually set to 0.
[0117] S103, Set hyperparameters.
[0118] a) Set the time decay factor α = 0.005 to control the impact of historical data when updating GSD distribution parameters online.
[0119] b) Set the sensitivity parameter γ = 1.0 and the default weight w in the KL divergence weighting. default =0.5, used for adjusting the image-level unsupervised loss.
[0120] Step 2: Iterative Training. The specific implementation process for each training iteration (batch) is as follows:
[0121] S201, Data Loading and GSD Preprocessing, including the following steps:
[0122] a) Load a batch of labeled images and a batch of unlabeled images. Set the batch size to 3, with a 2:1 ratio of labeled to unlabeled images.
[0123] b) For each image, obtain its original GSD value g. ori This value is typically provided as image metadata.
[0124] c) Scaling factors s when acquiring or applying data augmentation. Common data augmentation methods include random flipping, rotation, and color dithering, among which scaling is a key factor affecting GSD.
[0125] d) Calculate the standardized GSD value x = log2(g ori Using a logarithmic scale can make the GSD distribution closer to a Gaussian distribution, which is convenient for subsequent statistics.
[0126] S202, Labeled Image Processing.
[0127] a) The student model performs forward propagation on labeled images and calculates supervised loss based on the real labels. This loss typically includes classification loss and bounding box regression loss. For OBB detectors, the regression loss may be based on the rotated box parameters.
[0128] b) Based on the standardized GSD value x of the labeled image and the true object category it contains, update the labeled GSD distribution parameters of the corresponding category using an online weighted update formula. The mean of the GSD distribution for labeled data of this category can then be calculated using these parameters. and variance Note that after the preheating phase, this step needs to be completed in step S203.
[0129] S203, Unlabeled Image Processing (started after the warm-up phase). The warm-up phase refers to the first 10k iterations of training where unsupervised learning is not performed; only labeled data is used to train the teacher (or student) model to achieve a certain basic performance. This includes the following steps:
[0130] a) The teacher model performs forward propagation on the unlabeled image to generate pseudo labels, including predicted object categories and bounding box coordinates (and orientation angles). We set a confidence threshold of 0.5 to filter out low-quality pseudo labels.
[0131] b) Category-level pseudo-label GSD typicality weighting:
[0132] For each positive sample pseudo-label that passes the confidence threshold:
[0133] (1) Obtain its predicted category c and the normalized GSD value x of the current unlabeled image.
[0134] (2) Using the GSD distribution parameters of the corresponding category c of the labeled data (i.e. and by Calculated variance ), calculate the typicality of the pseudo-label's GSD value x in the GSD distribution of labels in this category;
[0135]
[0136] (3) The student model performs forward propagation on the unlabeled image (and its weighted pseudo-labels). The computation includes... Weighted R-CNN stage loss and RPN stage loss RPN loss typically doesn't directly use class-level weights because it's primarily responsible for proposing regions. R-CNN loss, on the other hand, multiplies the loss for each pseudo-label (classification and regression) by its corresponding w. cpl .
[0137] c) After the warm-up phase, a 5000-step warm-up phase is introduced to capture the GSD distribution of the unlabeled data. Following this, an image-level unsupervised loss, KLD weighted, is introduced: For the current unlabeled image I:
[0138] (1) Collect the set C of all high-quality pseudo-labels in the image. I .
[0139] (2) For C I For each category c, obtain its labeled data GSD distribution. GSD distribution of unlabeled data
[0140] (3) Calculate the KL divergence between these distributions:
[0141]
[0142] (4) Log-normalize and average the KL divergence for each category, then aggregate the results and obtain the final weights through an inverse proportional mapping function:
[0143]
[0144] d) Calculate the final weighted unsupervised loss for the unlabeled image:
[0145]
[0146] e) Based on the normalized GSD value x of the current unlabeled image and the categories of all its high-quality pseudo-labels, update the unlabeled GSD distribution parameters of the corresponding categories using an online weighted update formula.
[0147] f) Based on the standardized GSD value x of the labeled image and the true object category it contains, update the labeled GSD distribution parameters of the corresponding category using an online weighted update formula.
[0148] S203, Total Loss Calculation and Parameter Update, including the following steps:
[0149] a) Calculate the total loss for the current batch.
[0150] b) Using the backpropagation algorithm, based on Update the network parameters of the student model. The optimizer can use SGD with momentum, with an initial learning rate of 0.01 (for Faster RCNN OBB), momentum of 0.9, and weight decay of 0.0001. The learning rate will decay at 120k and 160k iterations during training.
[0151] c) The EMA (Exponential Moving Average) method is used to update the parameters of the teacher model using the updated parameters from the student model. Where θ T ←βθ T +(1-β)θ S , where θ T ,θ S These are the parameters for the teacher and student models, respectively, and β = 0.999 is the EMA decay rate.
[0152] Step 3, Repeat and Terminate. Repeat steps S201-S204 in Step 2 until the preset number of training iterations of 180k is reached.
[0153] Experimental validation: To verify the effectiveness of the proposed spatial resolution-guided semi-supervised remote sensing target detection method, experiments were conducted on the DOTA-v2.0 dataset. DOTA-v2.0 contains 11,268 high-resolution aerial images and 1.79 million instances. 1,882 images with GSD metadata (from the training and validation sets) were selected. After excluding the "airport" category due to insufficient data, experiments were conducted on the remaining 17 categories. The original high-resolution images were cropped into 1024×1024 pixel sub-images, resulting in 29,739 sub-images (denoted as trainval-gsd). The model was evaluated on the test-dev set, and the results were submitted to the official server for evaluation to obtain AP in COCO format. 50 and AP75 Metrics. 1%, 2%, 5%, and 10% of the images in the trainval-gsd dataset are randomly selected as labeled data, with the remainder as unlabeled data. To ensure that at least 5% of the images in rare categories are labeled, constraints were imposed during the sampling process. Five random samplings were performed at each proportion, resulting in five sets of labeled and unlabeled data. The mean and standard deviation of the results for each set are reported. The detector and backbone network use the Faster R-CNN OBB version, FPN structure, and ResNet-50 backbone. The method of this invention is integrated into both the classic Soft Teacher framework and the higher-performance MixTeacher framework.
[0154] Table 1 Performance Comparison of Different Semi-Supervised Target Detection Algorithms
[0155]
[0156]
[0157] Table 2 Ablation experiments of each weighted mechanism (1% labeling rate)
[0158]
[0159] As shown in Tables 1 and 2, the results after applying the method of this invention, compared with their baselines Soft Teacher and MixTeacher, show that AP 50 The performance metrics improved between +1.82 and +5.12. For example, at a 1% label ratio, Mix SRGT achieved 31.51 AP. 50 Compared to MixTeacher (27.08 AP) 50 AP increased by 4.43 50 At a 10% label ratio, Mix SRGT achieved 46.97 AP. 50 It outperforms other advanced methods of the same period, such as MCL (43.08 AP). 50 ).
[0160] For the Soft Teacher baseline (21.18 AP) 50 Using category-wise pseudo-label weighting (CLW) alone increases the AP to 22.79. 50 (+1.61); Improved to 22.63 AP using only image-level KLD weighting (KIW). 50 (+1.45); the combined result is 23.00 AP. 50 (+1.82).
[0161] For the MixTeacher baseline (27.08 AP) 50 ): CLW increased to 29.38 AP 50 (+2.30); KIW increased to 29.70 AP 50 (+2.62); the combination of the two reaches 31.51 AP. 50 (+4.43). The results show that both CLW and KIW are effective, and the combination of the two is the most effective.
[0162] In one embodiment of the present invention, a semi-supervised remote sensing target detection system based on spatial resolution guidance is provided, comprising, within a teacher-student semi-supervised learning framework:
[0163] The data acquisition module acquires remote sensing image data containing labeled and unlabeled datasets, as well as corresponding ground sampling distance metadata.
[0164] The initialization module initializes the teacher and student models, and initializes the ground sampling distance distribution parameters for each target category to describe the labeled and unlabeled data.
[0165] The GSD processing module performs logarithmic transformation on the original ground sampling distance values of each input remote sensing image, based on ground sampling distance normalization processing.
[0166] The update module maintains the ground sampling distance distribution of each target category in the labeled and unlabeled datasets through an online weighted update mechanism.
[0167] The weighting module dynamically adjusts the pseudo-label weights of positive samples in unlabeled images based on a class-level pseudo-label ground sampling distance typicality weighting mechanism; and the image-level unsupervised loss weighting mechanism based on KL divergence dynamically adjusts the total loss contribution of unlabeled images according to the difference in ground sampling distance distribution between labeled and unlabeled data.
[0168] The detection module integrates a weighting mechanism into the semi-supervised training process to train the final target detection model for remote sensing target detection.
[0169] In the above embodiments, the ground sampling distance standardization process is as follows:
[0170] x = log2(g ori / s),
[0171] In the formula, g ori represents the original ground sampling distance value, s represents the scaling factor in data augmentation, and x represents the standardized ground sampling distance value.
[0172] In the above embodiments, the online weighted update mechanism includes:
[0173] Maintain cumulative weights and W for each category c. c Weighted mean μ c and the weighted sum of squared differences M c ;
[0174] When processing an n-valued array containing class c c When analyzing an image of an object, its standardized ground sampling distance is x. A time decay factor α is introduced, and the parameter update rule is as follows:
[0175] Update cumulative weights: Maintain the cumulative weight sum for the updated category c. Maintain the cumulative weight sum for category c before the update;
[0176] Calculate the difference δ from the old mean: This is the weighted mean before the update;
[0177] Update the weighted mean: This is the updated weighted mean;
[0178] Update the weighted sum of squared differences: This is the updated weighted sum of squared differences. This is the cumulative weighted squared difference value before the update;
[0179] Calculate the variance:
[0180] In the above embodiments, the category-level pseudo-label ground sampling distance typicality weighting mechanism is implemented through the following steps: For a positive sample in an unlabeled image that is assigned a pseudo-label category c, calculate the typicality confidence score between the standardized ground sampling distance value x of the image and the ground sampling distance distribution parameter of category c learned from labeled data;
[0181] The original sample-level loss weights w old Multiplying the calculated ground sampling distance typicality confidence score by the pseudo-label weight yields the final category-level pseudo-label weight.
[0182] In the above embodiments, the image-level unsupervised loss weighting mechanism based on KL divergence is implemented through the following steps:
[0183] For each pseudo-label category c in the unlabeled image, calculate the KL divergence between its ground sampling distance distribution in labeled data and its ground sampling distance distribution in unlabeled data;
[0184] Log-normalize the KL divergence for each category;
[0185] For an unlabeled image containing n pseudo-labels, calculate its average normalized KL divergence, and calculate the image-level unsupervised loss weights based on the average normalized KL divergence using an inverse proportional function.
[0186] In the above embodiments, semi-supervised training includes:
[0187] A warm-up phase was set up, where only labeled data was used to train the teacher model;
[0188] After the warm-up phase, the teacher model generates pseudo-labels for the unlabeled images;
[0189] The typicality weights of the pseudo-label ground sampling distance are applied to the loss calculation of the R-CNN stage for unlabeled images;
[0190] The image-level KL divergence weights are applied to adjust the total unsupervised loss of the unlabeled image, and the total loss is calculated.
[0191] The student model parameters are updated via backpropagation, and the teacher model parameters are updated via exponential moving average.
[0192] In the above embodiments, the total loss is:
[0193]
[0194] in, Total loss; Standard monitoring loss; N u This represents the number of unlabeled images in the batch. These are the image-level unsupervised loss weights; The loss for the R-CNN stage of unlabeled images; For unlabeled images, the RPN stage loss is used. Let k be the unlabeled image.
[0195] The system provided in this embodiment is used to execute the above-described method embodiments. For specific processes and details, please refer to the above embodiments, which will not be repeated here.
[0196] In one embodiment of the present invention, a computing device is provided. This computing device can be a terminal and may include a processor, a communication interface, memory, a display screen, and an input device. The processor, communication interface, and memory communicate with each other via a communication bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. When the computer programs are executed by the processor, they implement the methods described in the above embodiments. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface is used for wired or wireless communication with external terminals. Wireless communication can be achieved through Wi-Fi, a management network, NFC (Near Field Communication), or other technologies. The display screen can be a liquid crystal display (LCD) or an e-ink display. The input device can be a touch layer covering the display screen, or buttons, a trackball, or a touchpad mounted on the casing of the computing device, or an external keyboard, touchpad, or mouse. The processor can call logical instructions stored in the memory.
[0197] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and sold or used as independent products, and can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0198] In one embodiment of the present invention, a computer program product is provided, the computer program product including a computer program stored on a non-transitory computer-readable storage medium, the computer program including program instructions, and when the program instructions are executed by a computer, the computer is able to perform the methods provided in the above-described method embodiments.
[0199] In one embodiment of the present invention, a non-transitory computer-readable storage medium is provided, which stores server instructions that cause a computer to perform the methods provided in the above embodiments.
[0200] The computer-readable storage medium provided in the above embodiments has a similar implementation principle and technical effect to the above method embodiments, and will not be described again here.
[0201] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0202] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0203] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0204] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A semi-supervised remote sensing target detection method based on spatial resolution guidance, characterized in that, This method is built upon a teacher-student semi-supervised learning framework, and includes: Acquire remote sensing image data containing labeled and unlabeled datasets, along with corresponding ground sampling distance metadata; Initialize the teacher and student models, and initialize the ground sampling distance distribution parameters for each target category to describe the labeled and unlabeled data; Based on the ground sampling distance standardization process, the original ground sampling distance values of each input remote sensing image are logarithmically transformed. A weighted online update mechanism is used to maintain the ground sampling distance distribution of each target category in both labeled and unlabeled datasets. Based on the class-level pseudo-label ground sampling distance typicality weighting mechanism, the pseudo-label weights of positive samples in unlabeled images are dynamically adjusted; based on the KL divergence image-level unsupervised loss weighting mechanism, the total loss contribution of unlabeled images is dynamically adjusted according to the difference in ground sampling distance distribution between labeled and unlabeled data. The weighting mechanism is integrated into the semi-supervised training process to train the final target detection model for remote sensing target detection.
2. The semi-supervised remote sensing target detection method based on spatial resolution guidance as described in claim 1, characterized in that, The ground sampling distance is standardized as follows: x=log2(g ori / s), In the formula, g ori represents the original ground sampling distance value, s represents the scaling factor in data augmentation, and x represents the standardized ground sampling distance value.
3. The semi-supervised remote sensing target detection method based on spatial resolution guidance as described in claim 1, characterized in that, Online weighted update mechanisms include: Maintain cumulative weights and W for each category c. c Weighted mean μ c and the weighted sum of squared differences M c ; When processing an n-valued array containing class c c When analyzing an image of an object, its standardized ground sampling distance is x. A time decay factor α is introduced, and the parameter update rule is as follows: Update cumulative weights: Maintain the cumulative weight sum for the updated category c. Maintain the cumulative weight sum for category c before the update; Calculate the difference δ from the old mean: This is the weighted mean before the update; Update the weighted mean: This is the updated weighted mean; Update the weighted sum of squared differences: This is the updated weighted sum of squared differences. This is the cumulative weighted squared difference value before the update; Calculate the variance:
4. The semi-supervised remote sensing target detection method based on spatial resolution guidance as described in claim 1, characterized in that, The category-level pseudo-label ground sampling distance typicality weighting mechanism is implemented through the following steps: For a positive sample in an unlabeled image that is assigned a pseudo-label class c, calculate the typicality confidence score between the standardized ground sampling distance value x of the image and the ground sampling distance distribution parameter of class c learned from labeled data; The original sample-level loss weights w old Multiplying the calculated ground sampling distance typicality confidence score by the pseudo-label weight yields the final category-level pseudo-label weight.
5. The semi-supervised remote sensing target detection method based on spatial resolution guidance as described in claim 1, characterized in that, The image-level unsupervised loss weighting mechanism based on KL divergence is implemented through the following steps: For each pseudo-label category c in the unlabeled image, calculate the KL divergence between its ground sampling distance distribution in labeled data and its ground sampling distance distribution in unlabeled data; Log-normalize the KL divergence for each category; For an unlabeled image containing n pseudo-labels, calculate its average normalized KL divergence, and calculate the image-level unsupervised loss weights based on the average normalized KL divergence using an inverse proportional function.
6. The semi-supervised remote sensing target detection method based on spatial resolution guidance as described in claim 1, characterized in that, Semi-supervised training includes: A warm-up phase was set up, where only labeled data was used to train the teacher model; After the warm-up phase, the teacher model generates pseudo-labels for the unlabeled images; The typicality weights of the pseudo-label ground sampling distance are applied to the loss calculation of the R-CNN stage for unlabeled images; The image-level KL divergence weights are applied to adjust the total unsupervised loss of the unlabeled image, and the total loss is calculated. The student model parameters are updated via backpropagation, and the teacher model parameters are updated via exponential moving average.
7. The semi-supervised remote sensing target detection method based on spatial resolution guidance as described in claim 6, characterized in that, The total loss is: in, Total loss; Standard monitoring loss; N u This represents the number of unlabeled images in the batch. These are the image-level unsupervised loss weights; The loss for the R-CNN stage of unlabeled images; For unlabeled images, the RPN stage loss is used. Let k be the unlabeled image.
8. A semi-supervised remote sensing target detection system based on spatial resolution guidance, characterized in that, Within the teacher-student semi-supervised learning framework, this includes: The data acquisition module acquires remote sensing image data containing labeled and unlabeled datasets, as well as corresponding ground sampling distance metadata. The initialization module initializes the teacher and student models, and initializes the ground sampling distance distribution parameters for each target category to describe the labeled and unlabeled data. The GSD processing module performs logarithmic transformation on the original ground sampling distance values of each input remote sensing image, based on ground sampling distance normalization processing. The update module maintains the ground sampling distance distribution of each target category in the labeled and unlabeled datasets through an online weighted update mechanism. The weighting module dynamically adjusts the pseudo-label weights of positive samples in unlabeled images based on a class-level pseudo-label ground sampling distance typicality weighting mechanism; and the image-level unsupervised loss weighting mechanism based on KL divergence dynamically adjusts the total loss contribution of unlabeled images according to the difference in ground sampling distance distribution between labeled and unlabeled data. The detection module integrates a weighting mechanism into the semi-supervised training process to train the final target detection model for remote sensing target detection.
9. A computer-readable storage medium for storing one or more programs, characterized in that, The one or more programs include instructions that, when executed by a computing device, cause the computing device to perform any of the methods described in claims 1 to 7.
10. A computing device, characterized in that, include: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing any of the methods described in claims 1 to 7.
Citation Information
Cited By
Small sample defect detection method based on class awareness prototype migration
CN121616602A
A small sample defect detection method based on a similar perception prototype migration
CN121616602B