A method for detecting a directed target of a remote sensing image based on weakly supervised learning

By employing multi-angle rotation data augmentation and pseudo-label mining, a directed target detection model for remote sensing images is constructed. This solves the problem of high manual annotation costs in directed target detection of remote sensing images and enables efficient detection of targets in any direction in remote sensing images.

CN116758421BActive Publication Date: 2025-11-04CHONGQING UNIV OF POSTS & TELECOMM SPACE COMM RES INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310714600.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-15
Publication Date
2025-11-04
Estimated Expiration
2043-06-15

AI Technical Summary

Technical Problem

Existing methods for directional target detection in remote sensing images rely on large-scale directional bounding box annotations, which result in high manual annotation costs and limited detection performance. How to extract the orientation and angle information of targets under weak supervision remains an urgent problem to be solved.

Method used

Rotated images of remote sensing images are generated by multi-angle rotation data augmentation. Directed pseudo-labels are mined using a teacher network, and iterative training is performed by combining horizontal bounding boxes and directed bounding box loss functions to construct a directed target detection model for remote sensing images. This reduces the dependence on fully supervised datasets and reduces annotation costs.

Benefits of technology

It effectively detects targets in any direction in remote sensing images, reduces the difficulty of data annotation and labor costs, and improves the prediction accuracy and efficiency of target detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116758421B_ABST
    Figure CN116758421B_ABST
Patent Text Reader

Abstract

The present application belongs to the field of intelligent analysis of remote sensing images, and particularly relates to a remote sensing image directed target detection method based on weak supervision learning, which comprises obtaining remote sensing image data to be detected, and pre-processing the remote sensing image data to be detected; inputting the pre-processed remote sensing image data to be detected into a trained remote sensing image directed target detection model for detection processing to obtain a target remote sensing image; the present application only needs to use a small amount of labeled horizontal frame targets, adopts a weak supervision method to mine directed targets, and realizes remote sensing image target detection. The weak supervision data set used by the present application not only reduces the data labeling difficulty and reduces the number of required labeling frames, but also effectively reduces the dependence of the existing remote sensing image directed target detection method on a full supervision data set, and reduces the artificial labeling cost.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of intelligent analysis of remote sensing images, and particularly relates to a weakly supervised remote sensing image oriented object detection method based on oriented object mining. BACKGROUND

[0002] Target detection is one of the basic tasks of high-resolution optical remote sensing image intelligent interpretation, and can be widely applied to fields such as city planning, agricultural monitoring, disaster assessment, and military reconnaissance. Compared with natural images taken from a ground level perspective, remote sensing images are usually imaged from a bird's eye view, and the ground object targets in the images are usually distributed in any direction. In order to effectively detect the direction information of the target, researchers introduce a large-scale oriented bounding box annotation dataset for deep learning network training to achieve accurate detection of targets in any direction in remote sensing images. However, compared with the general horizontal bounding box annotation method, annotating oriented bounding boxes is more time-consuming and labor-intensive. In addition, due to the large imaging scene of remote sensing images, the complex ground targets, and the rich texture information, the cost of manual annotation is further increased.

[0003] Although great progress has been made in the research of oriented target detection of remote sensing images based on deep learning in recent years, most of the existing methods need to be fully supervised on a large-scale dataset annotated with oriented bounding boxes, and the detection performance of the network is severely dependent on the accuracy of the dataset annotation and the number of samples. In order to reduce the cost of data annotation, the method of remote sensing image target detection based on weakly supervised learning has become a research trend in recent years. However, due to the lack of direction supervision information, the detection ability of the existing weakly supervised method for oriented targets in the image is still limited, and how to mine the direction angle information of the target to be detected from the weakly supervised training data is still a problem to be solved. SUMMARY

[0004] To solve the problems existing in the prior art, the present application proposes a remote sensing image oriented target detection method based on weakly supervised learning. The method comprises:

[0005] Obtaining remote sensing image data to be detected, and pre-processing the remote sensing image data to be detected;

[0006] Inputting the pre-processed remote sensing image data to be detected into a trained remote sensing image oriented target detection model for detection processing to obtain a target remote sensing image;

[0007] The training process of the remote sensing image oriented target detection model comprises:

[0008] Obtaining an original remote sensing dataset, and pre-processing the original remote sensing dataset; the original remote sensing dataset comprises training remote sensing images with real horizontal bounding box labels;

[0009] The training remote sensing image is subjected to rotation data enhancement to generate multi-angle rotation images of the training remote sensing image;

[0010] The multi-angle rotation images of the training remote sensing image are input into a teacher network for directional target mining to predict directional pseudo labels of the training remote sensing image;

[0011] The training remote sensing image, the corresponding real horizontal bounding box label and the directional pseudo label are input into a student network to predict directional prediction boxes of the training remote sensing image;

[0012] The directional prediction boxes of the training remote sensing image and the real horizontal bounding box label of the training remote sensing image are used to construct a horizontal bounding box loss;

[0013] The directional prediction boxes of the training remote sensing image and the directional pseudo label of the training remote sensing image are used to construct a directional bounding box loss;

[0014] The parameters of the student network are adjusted through iterative training of the horizontal bounding box loss and the directional bounding box loss, the parameters of the teacher network are updated by exponential moving average according to the parameters of the student network, and when the loss function converges, a trained remote sensing image directional target detection model is obtained.

[0015] The present application has the following advantages:

[0016] The present application uses only part of the horizontal frame label as a weakly supervised training data set, continuously mines unannotated directional targets from the input image through a directional target mining network based on multi-angle rotation transformation, and adds the mined directional targets as pseudo labels to subsequent training, so that the network can finally detect any direction of the target of interest in the remote sensing image. Compared with the fully supervised data set using complete directional bounding box annotation, the weakly supervised data set used in the present application not only reduces the difficulty of data annotation and reduces the number of annotation frames required, but also effectively reduces the dependence of the existing remote sensing image directional target detection method on the fully supervised data set, and reduces the cost of manual annotation. At the same time, the position of the candidate frame is refined by random jitter, and the position prediction accuracy of the refined candidate frame is judged by the refined regression method, which effectively improves the positioning accuracy of the generated effective pseudo label, thereby further improving the prediction effect of the overall target detection. BRIEF DESCRIPTION OF DRAWINGS

[0017] Figure 1 A weakly supervised learning-based remote sensing image directional target detection method flowchart of the present application;

[0018] Figure 2 A weakly supervised learning-based directional target detection framework diagram of the present application;

[0019] Figure 3 A schematic diagram of a remote sensing image directed target detection framework for the teacher network and the student network of the present application;

[0020] Figure 4 A flowchart of a remote sensing image directed target detection model training process for the present application;

[0021] Figure 5 A schematic diagram of a remote sensing image directed target rotation transformation for the present application;

[0022] Figure 6 A real horizontal frame labeling and directed pseudo label example diagram for the present application;

[0023] Figure 7 A flowchart of pseudo label screening threshold calculation in the dynamic pseudo label threshold filtering method for the present application;

[0024] Figure 8 A flowchart of screening out candidate pseudo labels with unreliable positioning for the present application. DETAILED DESCRIPTION

[0025] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0026] Figure 1 A weakly supervised learning remote sensing image directed target detection method flowchart for the embodiments of the present application, as shown in Figure 1 , the method comprises:

[0027] 101, obtaining remote sensing image data to be measured, and pre-processing the remote sensing image data to be measured;

[0028] 102, inputting the pre-processed remote sensing image data to be measured into the trained remote sensing image directed target detection model for detection processing, to obtain a target remote sensing image.

[0029] As Figure 2As shown, the remote sensing image oriented target detection model provided by the present application mainly consists of two parts: a teacher network for mining oriented targets and a student network for training an oriented target detector. In each iteration training, the teacher network is used to mine the directional information of unannotated targets in the training image, and provides the student network with oriented pseudo labels; the student network takes the original training image, the horizontal bounding box labeled target corresponding to the original training image and the oriented pseudo labels generated by the teacher network as training samples, respectively calculates the horizontal bounding box loss and the oriented bounding box loss, and trains an oriented target detector for remote sensing images. The teacher network and the student network use the same remote sensing image oriented target detection network framework, and the network parameters of the teacher network are updated by exponential moving average according to the network parameters of the student network.

[0030] In some embodiments of the present application, the present application uses Oriented RCNN as the remote sensing image oriented target detection framework of the teacher network and the student network, as shown in Figure 3 As shown, the framework includes: 1) a backbone network based on ResNet-FPN, used to extract feature maps of the input image; 2) an Oriented RPN network, used to generate oriented region candidate boxes; 3) a Rotated RoIAlign operation, used to extract region features of a fixed size and use these features as inputs of the Oriented R-CNN Head detection head; 4) an Oriented R-CNN Head, used to perform foreground classification and spatial position regression on the oriented candidate boxes.

[0031] Figure 4 The training flowchart of the remote sensing image oriented target detection model of the embodiments of the present application is shown in Figure 4 As shown, the training flowchart includes:

[0032] 201, obtaining an original remote sensing dataset and preprocessing the original remote sensing dataset; the original remote sensing dataset includes training remote sensing images with real horizontal bounding box labels;

[0033] In the embodiments of the present application, the preprocessing of the original remote sensing dataset can specifically include extracting subgraphs of 1024x1024 resolution size on all original remote sensing images in a sliding window manner with a step size of 824 pixels; selecting all horizontal and vertical targets with angles of 0°(±5°) and 90°(±5°) in the training set as real labels, and labeling them using horizontal bounding boxes; of course, in actual operation, the step size, the resolution size and the corresponding rotation angle can be appropriately adjusted.

[0034] 202, performing multi-angle rotation data augmentation on the training remote sensing images to generate multi-angle rotation images of the training remote sensing images;

[0035] In this embodiment of the invention, by performing multi-angle rotation transformation on the training remote sensing image, the oriented target in the training remote sensing image is brought as close as possible to the horizontal and vertical directions.

[0036] Because this invention only labels horizontal and / or vertical targets in the training image with true horizontal bounding boxes, it lacks the ability to detect other oriented targets in the image. Therefore, as Figure 5 As shown, this invention transforms unlabeled directed targets into horizontal or vertical targets by performing multi-angle rotation transformations on training remote sensing images, thereby enabling their detection. The corresponding rotation angle can be approximated as the target's orientation angle in the original image, generating corresponding directed bounding boxes. These directed bounding boxes are used as pseudo-labels for training the directed target detection model, which adds angular information to the training data, enabling the detection of targets in any orientation. The true horizontal bounding box annotations and the final generated directed pseudo-labels of this invention are shown below. Figure 6 As shown.

[0037] 203. Input the multi-angle rotated image of the training remote sensing image into the teacher network for directed target mining, and predict the directed pseudo-labels of the training remote sensing image.

[0038] In this embodiment of the invention, the multi-angle rotated image of the training remote sensing image is input into the teacher network for directed target mining, and the predicted directed pseudo-labels of the training remote sensing image include:

[0039] Multi-angle rotational augmented images of the training remote sensing images are input into the teacher network for inference detection, and directed prediction boxes are obtained for each rotational augmented image;

[0040] The directed prediction boxes of the multi-angle rotation-enhanced image are projected onto the corresponding training remote sensing image through inverse rotation transformation;

[0041] The directed predicted bounding boxes and the true horizontal bounding box labels corresponding to all angles of the training remote sensing image are stitched together. The non-maximum suppression algorithm is used to remove the duplicate directed predicted bounding boxes, and the directed predicted bounding boxes at the true horizontal bounding box labels are removed from the results of the non-maximum suppression algorithm.

[0042] The confidence score thresholds for directed prediction boxes with different categories are calculated using a dynamic pseudo-label threshold filtering method.

[0043] The directed prediction boxes are filtered using the confidence score threshold to remove low-quality directed prediction boxes.

[0044] The reserved directional prediction box is sent into a directional pseudo-label position correction module as a candidate pseudo-label, and candidate pseudo-labels with unreliable positions are screened out, and finally reserved directional prediction boxes are used as directional pseudo-labels of corresponding training remote sensing images.

[0045] In the embodiment of the present application, as shown in Figure 7 The confidence score threshold of each directional prediction box is calculated by using the dynamic pseudo-label threshold filtering method, which includes:

[0046] In the preheating stage, the teacher network is not used to mine directional pseudo-labels, and only the original horizontal box label is used to train the student network. After one complete traversal of the entire training set, the directional target is mined as a pseudo-label to join the training of the student network;

[0047] The confidence score distribution of directional prediction boxes of different categories is counted by performing inference detection on all training remote sensing images in the original remote sensing data set, and the mean of the confidence scores in the front of each category is selected as the confidence score threshold of the category in the next round of training. For example, the mean of the top fifteen percent confidence scores of each category can be selected as the confidence score threshold of the category in the next round of training;

[0048] The student network is trained using the real horizontal bounding box label and the real-time generated directional pseudo-label until the training remote sensing images in the entire data set are iterated once.

[0049] The above steps are repeated until the confidence score threshold of all categories reaches the set maximum threshold upper limit.

[0050] In the preferred embodiment of the present application, the confidence score threshold is used to screen directional prediction boxes, which includes retaining the directional prediction box when the confidence score of the directional prediction box is greater than or equal to the set confidence score threshold; otherwise, the directional prediction box is discarded.

[0051] In the embodiment of the present application, the directional pseudo-label position correction module is used to screen out candidate pseudo-labels with unreliable positions, as shown in Figure 8 which specifically includes:

[0052] The bounding box jitter, refinement regression, regression offset calculation and threshold screening will be described in detail as follows:

[0053] Based on the jitter coefficient, a group of candidate boxes after directional bounding box jitter are generated around the candidate pseudo-label;

[0054] The group of jittered candidate boxes are sent into the detection head of the teacher network for refinement regression of the position of the candidate box to generate a refined candidate box; the refined candidate box includes:

[0055]

[0056] wherein, represents the i-th refined candidate box, i∈{1,2,...,K}, K represents the number of candidate boxes; jitter(·ε) represents random jittering of the candidate box with jittering coefficient ε; refine represents a refinement regression operation.

[0057] According to the center point position of the refined candidate box and the length and width size of the boundary box, a position regression offset between the refined candidate boxes is calculated.

[0058] The position regression offset between the refined candidate boxes comprises:

[0059]

[0060] wherein, represents the position regression offset between the refined candidate boxes, δ center , δ shape respectively are the sum of standard deviations of the center point horizontal and vertical coordinates and the length and width of the boundary box of all candidate boxes after K times jittering; δ angle is the standard deviation of the corresponding angle of all candidate boxes; w, h respectively represent the width and length of the boundary box of the candidate pseudo label p, and K represents the number of candidate boxes.

[0061] According to the position regression offset, a candidate pseudo label with unreliable position prediction is screened out, and the finally reserved candidate pseudo label will be used as the final directional pseudo label.

[0062] In the embodiment of the present application, according to the position regression offset, a candidate pseudo label with unreliable position prediction is screened out, which includes that when the position regression offset of the candidate pseudo label is greater than or equal to the set regression offset threshold, the candidate pseudo label is reserved; otherwise, the candidate pseudo label is discarded.

[0063] 204, input the training remote sensing image and its corresponding real horizontal boundary box label and directional pseudo label into the student network to predict the directional prediction box of the training remote sensing image;

[0064] In the embodiment of the present application, the training remote sensing image and the corresponding real horizontal boundary box label and directional pseudo label are input into the student network, and the student network can predict the directional prediction box of the training remote sensing image, and can also compare the directional prediction box with the real horizontal boundary box label and the directional pseudo label respectively to optimize the student network model.

[0065] 205, construct a horizontal boundary box loss by using the directional prediction box of the training remote sensing image and the real horizontal boundary box label of the training remote sensing image; the horizontal boundary box loss is represented as:

[0066] L hbb =L clsh +L regh

[0067]

[0068]

[0069] wherein, L hbb represents a horizontal bounding box loss; L clsh represents a horizontal bounding box classification loss; L regh represents a horizontal bounding box regression loss; N cls is the number of to-be-classified oriented prediction boxes in the training remote sensing image; p i is the predicted class of the i-th to-be-classified oriented prediction box; is the class of the i-th to-be-classified oriented prediction box corresponding to the real horizontal bounding box label; N reg is the number of to-be-regressed oriented prediction boxes in the training remote sensing image; is the real coordinate offset between the i-th to-be-regressed oriented prediction box and the corresponding real horizontal bounding box label; t i is the predicted coordinate offset between the i-th to-be-regressed oriented prediction box and the prediction box whose coordinates are regressed; smoothL1 represents a Smooth L1 function.

[0070] 206, using the oriented prediction box of the training remote sensing image and the oriented pseudo label of the training remote sensing image, a directed bounding box loss is constructed; the directed bounding box loss is represented as:

[0071] L obb =L clso +L rego

[0072]

[0073]

[0074] wherein, L obb represents a directed bounding box loss; L clso represents a directed bounding box classification loss; L rego represents a directed bounding box regression loss; N cls is the number of to-be-classified oriented prediction boxes in the training remote sensing image; p i is the predicted class of the i-th to-be-classified oriented prediction box; is the class of the i-th to-be-classified oriented prediction box corresponding to the oriented pseudo label; N reg is the number of to-be-regressed oriented prediction boxes in the training remote sensing image; is the real coordinate offset between the i-th to-be-regressed oriented prediction box and the corresponding oriented pseudo label; t i is the predicted coordinate offset between the i-th to-be-regressed oriented prediction box and the prediction box whose coordinates are regressed; and smoothL1 represents a Smooth L1 function.

[0075] 207、Through iterative training of the joint horizontal bounding box loss and the oriented bounding box loss, the parameters of the teacher network and the student network are adjusted, and when the loss function reaches convergence, a trained remote sensing image oriented target detection model is obtained.

[0076] In the embodiments of the present application, the network parameters are updated in a back propagation manner by continuously calculating the loss function, and the iteration is continuously updated, so as to improve the detection accuracy of the model. When the loss function is minimized or converges, it indicates that the model training is completed, and the model can be used to test the to-be-tested remote sensing image data to obtain a target remote sensing image.

[0077] Those skilled in the art can understand that all or part of the steps in the above-mentioned embodiments can be completed by a program instructing related hardware, and the program can be stored in a computer readable storage medium, which can include ROM, RAM, magnetic disk or optical disk, etc.

[0078] Although the embodiments of the present application have been shown and described, it can be understood by those skilled in the art that various changes, modifications, replacements and variations can be made to the embodiments without departing from the principles and spirits of the present application, and the scope of the present application is defined by the appended claims and their equivalents.

Claims

1. A method for directional target detection in remote sensing images based on weakly supervised learning, characterized in that, include: Acquire the remote sensing image data to be tested and preprocess the remote sensing image data to be tested; The preprocessed remote sensing image data to be tested is input into the trained remote sensing image directed target detection model for detection processing to obtain the target remote sensing image; The training process for the directional target detection model in remote sensing images includes: Obtain the original remote sensing dataset and preprocess it; the original remote sensing dataset includes training remote sensing images with true horizontal bounding box labels. The training remote sensing image is subjected to multi-angle rotation data augmentation to generate a multi-angle rotated image of the training remote sensing image; The multi-angle rotated images of the training remote sensing images are input into the teacher network for directed target mining, and directed pseudo-labels of the training remote sensing images are predicted, including: Multi-angle rotational augmented images of the training remote sensing images are input into the teacher network for inference detection, and directed prediction boxes are obtained for each rotational augmented image; The directed prediction boxes of the multi-angle rotation-enhanced image are projected onto the corresponding training remote sensing image through inverse rotation transformation; The directed predicted bounding boxes and the true horizontal bounding box labels corresponding to all angles of the training remote sensing image are stitched together. The non-maximum suppression algorithm is used to remove the duplicate directed predicted bounding boxes, and the directed predicted bounding boxes at the true horizontal bounding box labels are removed from the results of the non-maximum suppression algorithm. The confidence score thresholds for directed prediction boxes with different categories are calculated using a dynamic pseudo-label threshold filtering method. The directed prediction boxes are filtered using the confidence score threshold to remove low-quality directed prediction boxes. The retained directed prediction boxes are sent as candidate pseudo-labels to the directed pseudo-label position correction module to filter out candidate pseudo-labels with unreliable positioning, and the finally retained directed prediction boxes are used as the directed pseudo-labels of the corresponding training remote sensing images. The training remote sensing image and its corresponding ground truth horizontal bounding box label and directed pseudo label are input into the student network to predict the directed prediction box of the training remote sensing image. A horizontal bounding box loss is constructed using the directed predicted bounding boxes of the training remote sensing images and the true horizontal bounding box labels of the training remote sensing images. A directed bounding box loss is constructed using directed predicted boxes and directed pseudo-labels of training remote sensing images. The parameters of the student network are adjusted by iterative training using a combination of horizontal and directed bounding box losses. The parameters of the teacher network are then updated using an exponential moving average based on the parameters of the student network. When the loss function converges, a well-trained directed target detection model for remote sensing images is obtained.

2. The method for directional target detection in remote sensing images based on weakly supervised learning according to claim 1, characterized in that, The method of calculating the confidence score thresholds for different directed prediction box categories using dynamic pseudo-label threshold filtering includes: Infer all training remote sensing images in the original remote sensing dataset, statistically analyze the distribution of confidence scores for different categories of directed prediction boxes, and select the average of the top confidence scores for each category as the confidence score threshold for that category in the next round of training. The student network was trained using real horizontal bounding box labels and real-time generated directed pseudo-labels until all training remote sensing images in the entire dataset had been iterated once. Repeat the above steps until the confidence score thresholds for all categories reach the set maximum threshold limit.

3. The method for directional target detection in remote sensing images based on weakly supervised learning according to claim 1, characterized in that, The directed prediction box is filtered using the confidence score threshold. If the confidence score of the directed prediction box is greater than or equal to the set confidence score threshold, the directed prediction box is retained; otherwise, the directed prediction box is discarded.

4. The method for directional target detection in remote sensing images based on weakly supervised learning according to claim 1, characterized in that, The directed pseudo-tag position correction module is used to filter out candidate pseudo-tags with unreliable positioning, including: A set of candidate boxes, after being jittered by the directed bounding box, is generated around the candidate pseudo-label based on the jitter coefficient. A set of jittered candidate boxes is fed into the detection head of the teacher network, and the position of the candidate boxes is refined by regression to generate refined candidate boxes. Calculate the positional regression offset between the refined candidate boxes based on the center point position of the refined candidate boxes and the length and width dimensions of the bounding boxes; Based on the location regression offset, candidate pseudo-labels with unreliable location predictions are filtered out, and the final retained candidate pseudo-labels will be used as the final directed pseudo-labels.

5. The method for directional target detection in remote sensing images based on weakly supervised learning according to claim 4, characterized in that, The refined candidate boxes include: in, This represents the i-th refined candidate box. , Indicates the number of candidate boxes; Indicates the jitter coefficient for the candidate box. Perform random shaking; This indicates a more refined regression operation.

6. The method for directional target detection in remote sensing images based on weakly supervised learning according to claim 4, characterized in that, The positional regression offset between the refined candidate boxes includes: in, This represents the positional regression offset between refined candidate boxes. , They were respectively through The sum of the standard deviations of the horizontal and vertical coordinates of the center points of all candidate boxes and the length and width of the bounding boxes after each shake; The standard deviation of the angles corresponding to all candidate boxes; , These represent candidate pseudo-labels. The width and length of the bounding box, This indicates the number of candidate boxes.

7. The method for directional target detection in remote sensing images based on weakly supervised learning according to claim 4, characterized in that, Based on the location regression offset, candidate pseudo-labels with unreliable location predictions are filtered out. If the location regression offset of a candidate pseudo-label is greater than or equal to a set regression offset threshold, the candidate pseudo-label is retained; otherwise, the candidate pseudo-label is discarded.

8. The method for directional target detection in remote sensing images based on weakly supervised learning according to claim 1, characterized in that, The horizontal bounding box loss is expressed as: in, Indicates the loss of the horizontal bounding box; This represents the classification loss for the horizontal bounding box; This represents the regression loss of the horizontal bounding box; To train the number of directed prediction boxes to be classified in remote sensing images; The predicted category is the i-th directed bounding box to be classified. Let be the category of the true horizontal bounding box label corresponding to the i-th directed predicted bounding box to be classified; To train the number of directed prediction boxes to be regressed in remote sensing images; is the true coordinate offset between the i-th directed prediction box to be regressed and the corresponding true horizontal bounding box label; The predicted coordinate offset between the i-th directed prediction box to be regressed and the prediction box of its coordinate regression; This represents the Smooth L1 function.

9. The method for directional target detection in remote sensing images based on weakly supervised learning according to claim 1, characterized in that, The directed bounding box loss is expressed as: in, This represents the directed bounding box loss; This represents the directed bounding box classification loss; This represents the directed bounding box regression loss; To train the number of directed prediction boxes to be classified in remote sensing images; The predicted category is the i-th directed bounding box to be classified. Let i be the category of the directed pseudo-label corresponding to the i-th directed prediction box to be classified; To train the number of directed prediction boxes to be regressed in remote sensing images; Let be the true coordinate offset between the i-th directed prediction box to be regressed and the corresponding directed pseudo-label; The predicted coordinate offset between the i-th directed prediction box to be regressed and the prediction box of its coordinate regression; This represents the Smooth L1 function.

Citation Information

Patent Citations

  • Remote sensing image target detection method and system based on semi-supervised iterative learning

    CN113688665A

  • RGB image semi-supervised target detection method based on double-pseudo-label optimization learning

    CN115393687A