Weakly Supervised Object Detection Method for Remote Sensing Images Based on Weight Redistribution and Task Alignment
By introducing high-quality positive instance mining, positive instance weight reassignment and task-aligned border regression losses, the misclassification and branch inconsistency in weakly supervised remote sensing image object detection is solved, and the accuracy and consistency of the detection is improved.
Patent Information
- Application Number
- CN202411636045.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-15
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2044-11-15
AI Technical Summary
In the existing weakly supervised remote sensing image object detection, there are problems such as misclassification of neighborhood instances, focusing only on the most prominent areas of the target, and inconsistent classification branches and regression branches.
By introducing high-quality positive instance mining, positive instance weight reassignment and task-aligned border regression loss, the model is optimized to solve the misclassification problem, encourage the model to focus on the overall target, and improve the coupling between classification branches and regression branches.
It effectively improves the misclassification of neighborhood instances, improves the overall effect of weakly supervised object detection in remote sensing images, and improves the accuracy and consistency of detection.
Smart Images

Figure CN119540788B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and further relates to the technology of weakly supervised object detection in remote sensing images. Specifically, it is a weakly supervised object detection method for remote sensing images based on weight redistribution and task alignment, which can be used in military strategy, urban planning, and the construction of geographic information systems, etc. Background Art
[0002] Weakly supervised object detection uses image-level class labels to complete the localization and classification of high-value targets. Compared with the instance-level labels used in fully supervised object detection, weakly supervised object detection saves a large amount of annotation costs. This technology is widely used in military, civilian, environmental monitoring, agriculture and other fields. Existing weakly supervised object detection is achieved through multi-instance learning. Specifically, the input image is regarded as a set of object instances, and the image-level label is used to supervise the instance classifier in the multi-instance learning framework. With the development of weakly supervised object detection, the iterative optimization of the instance classifier using instance-level pseudo-labels (including positive instances and negative instances) has attracted much attention from researchers. First, seed instances are mined as positive instances using the prediction scores. Secondly, during the label propagation process, when there is a large spatial similarity between the neighboring instances and the seed instances, positive labels are assigned to the neighboring instances, otherwise negative labels are assigned. Accordingly, the instance-level pseudo-labels supervise the instance classifier.
[0003] In the prior art document WSDDN [H. Bilen and A. Vedaldi, “Weakly supervised deep detection networks,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), Jun. 2016, pp. 2846–2854.], WSDDN first combines multi-instance learning and deep learning to complete weakly supervised object detection. First, the selective search algorithm is used to generate object candidate boxes for the input image. Second, the object candidate boxes and the input image are input into the backbone network, region of interest pooling, and two fully connected layers to obtain the feature vectors of each object candidate box. Then, the feature vectors of the object candidate boxes are input into the classification branch and the detection branch in the multi-instance detection network. Finally, the prediction scores of the classification branch and the detection branch are multiplied element-wise and summed along the instance dimension to obtain the image-level prediction score. The binary cross-entropy loss is calculated between the image-level prediction score and the image-level class label for backpropagation. The prior art document OICR [P. Tang, X. Wang, X. Bai, and W. Liu, “Multiple instance detection network with online instance classifier refinement,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., Jul. 2017, pp. 2843–2851.], on the basis of WSDDN, adds multiple parallel instance classification optimization branches. The supervision signal of each branch generates instance-level pseudo-labels using the prediction scores of the previous branch. Specifically, the instance with the highest prediction score is selected as the seed instance. Second, during the label propagation process, when the neighboring instances and the seed instance have a large spatial relationship, positive labels are assigned to the neighboring instances, and negative labels are assigned otherwise. Accordingly, the generated instance-level pseudo-labels supervise the instance classification optimization branches.
[0004] The main problems existing in current weakly supervised remote sensing image object detection include the following: 1) The misclassification problem of neighboring instances; existing weakly supervised object detection only relies on spatial similarity to assign pseudo-labels to neighboring instances during the label propagation process, resulting in spatially adjacent objects being wrongly assigned pseudo-labels, thus causing the neighboring instances to be misclassified. 2) The problem of locating the most prominent region of the target; the weakly supervised object detection model assigns high loss weights to the instances covering the most prominent region of the target, resulting in the model only focusing on the most prominent region of the target. 3) The incompatibility problem between the classification task and the regression task; in weakly supervised object detection, the classification task and the regression task use shared features, and there is no coupling between them, resulting in the fact that some object candidate boxes with high classification scores are not necessarily the most accurate in regression. Summary of the Invention
[0005] Aiming at the deficiencies of the existing technology, the present invention aims to propose a weakly supervised object detection method for remote sensing images based on weight redistribution and task alignment, so as to solve the problems of misclassification of neighboring objects, only focusing on the most prominent region of the object, and inconsistency between the classification branch and the regression branch in the existing detection scheme. By introducing high-quality positive instances mining (HPIM) into the existing online instance classifier optimization model OICR, considering the spatial similarity and feature similarity between instances during the label propagation process, the problem of misclassification of neighboring instances is avoided; at the same time, by calculating the spatial relationship matrix between positive instances, the instances that cover the entire object are screened out, and through positive instances weights re-assigning (PIWR), the weakly supervised remote sensing image object detection model is encouraged to focus on the whole object; finally, considering the instability of the weakly supervised object detection model in the initial stage, the task-aligned bounding box regression loss (TABRL) is introduced. By designing a task-aligned modulation factor, the model can dynamically adjust the weight according to the number of iterations, and then combine the modulation factor with the difficulty level of the instance as the weight of the regression branch, improving the consistency between the classification branch and the regression branch. The present invention can effectively overcome the defects existing in the existing technology and improve the weakly supervised object detection effect of remote sensing images.
[0006] The specific steps for the present invention to achieve the above object include the following:
[0007] (1) Given a remote sensing image where H, W, and B respectively represent the width, height, and number of channels of the image; use the selective search algorithm to generate M target candidate boxes for the remote sensing image I, forming a target candidate box set PR = {pr1, pr2,... pr m , pr M}, where pr m represents the m-th target candidate box;
[0008] (2) Input the remote sensing image I and the target candidate box set PR into the online instance classifier optimization model OICR, which includes a backbone network, a multi-instance detection network, an instance classification optimization branch, and a task-aligned bounding box regression branch; obtain the feature vectors of all target candidate boxes through the backbone network, region of interest pooling, and two fully connected layers and send them into the multi-instance detection network, where f m represents the feature vector of the m-th target candidate box;
[0009] (3) The feature vector of the target candidate box The two parallel branches in the input multi-instance detection network respectively obtain two score matrices in class c after passing through the two branches Obtain the class confidence score matrix of the target candidate box in class c according to the following formula
[0010]
[0011] where ⊙ represents the Hadamard product, and c is any one of all classes C;
[0012] Obtain the image-level class prediction score φ according to the following formula c :
[0013]
[0014] where represents the class confidence score of the m-th target candidate box in class c;
[0015] (4) Add K parallel instance classification optimization branches to the baseline weakly supervised object detection model. By inputting the target candidate box feature vector into the k-th instance classification optimization branch, obtain the class confidence score of the target candidate box where C + 1 represents the background class;
[0016] (5) Use the prediction score of the (k - 1)-th branch as the input for high-quality positive instance mining. By sorting in descending order, select the top p% of the instances as the initial seed instance set for the c-th class in the k-th branch Secondly, perform non-maximum suppression operation on the initial seed instances to obtain the seed instance set to achieve seed instance mining;
[0017] (6) Mine the initial neighborhood instances through spatial similarity. Set the spatial similarity determination threshold ζ. When the spatial similarity IoU between the j-th seed instance and other instances is greater than the threshold ζ, determine this instance as an initial neighborhood instance, otherwise determine it as a background instance; obtain the initial neighborhood instance set of the seed instance
[0018] (7) Eliminate misclassified initial neighborhood instances through feature similarity between instances. Set the feature similarity threshold δ. When the feature similarity between the j-th seed instance and the initial neighborhood instance is less than δ, determine this initial neighborhood instance as a background instance, and obtain the neighborhood instance set of the j-th seed instance Obtain the neighborhood instance set of all seed instances to get the instance-level supervision information y of the k-th instance classification optimization branch k ;
[0019] (8) Assign loss weights to positive instances that completely cover the target through the positive instance weight redistribution strategy. First, construct a subset of positive instances SN for the seed instance and its neighborhood instances, then calculate the spatial relationship between the instances in SN to construct a spatial distance matrix Then sum G along the row direction to obtain a spatial distance vector Find the maximum value in V to mine the instance r that completely covers the target z , and at the same time find the instance r in the subset of positive instances SN that has the highest predicted class confidence score p , in the subset of positive instances SN, swap the loss weights of instance r z and r p to achieve positive instance weight redistribution;
[0020] (9) Construct the loss function of the k-th instance classification optimization branch
[0021]
[0022] Among them, represents the loss weight of the target candidate box m in the k-th branch, represents the element in the c-th row and m-th column of y k ; when is equal to 1, it means that the target candidate box m contains the target of class c; when it is equal to 0, it means that the target candidate box m does not contain the target of class c; represents the predicted class confidence score of the target candidate box m on class c in the k-th branch;
[0023] (10) The task-aligned bounding box regression branch performs bounding box regression on positive instances, and at the same time considers the difficulty of positive instances and the number of model iterations to increase the coupling between the model classification task and the regression task, as follows:
[0024] (10.1) Design a modulation factor μ that changes dynamically with model iteration to ensure the stability of training:
[0025]
[0026] Among them, t represents the current iteration number, T represents the total number of model iterations; σ represents a hyperparameter used to control the growth rate;
[0027] (10.2) The task-aligned bounding box regression loss of the k-th instance classification optimization branch It is defined as follows:
[0028]
[0029]
[0030] The number of elements in; Indicates the predicted class confidence score of the l-th positive instance corresponding to the class in the k-th instance classification optimization branch; and g l respectively represent the target offset and predicted offset of the l-th positive instance;
[0031] (11) Construct the overall loss L of the weakly supervised remote sensing image object detection model ALL :
[0032]
[0033] where L WSDDN represents the loss function of the baseline weakly supervised object detection network, and K represents the number of instance classification optimization branches;
[0034] (12) Use L ALL to train the entire weakly supervised remote sensing image object detection model, input the remote sensing image to be measured and its object candidate boxes into the trained model, and obtain the classes and positions of the objects of interest in the image, so as to realize weakly supervised object detection of remote sensing images.
[0035] Compared with the prior art, the present invention has the following advantages:
[0036] First, since the present invention introduces high-quality positive instance mining, during the label propagation process, by considering the feature similarity between the seed instance and the neighborhood instances, the misclassified neighborhood instances are removed, so as to alleviate the misclassification problem of neighborhood instances and effectively improve the mining effect.
[0037] Second, the present invention proposes and designs a positive instance weight redistribution strategy, by calculating the spatial relationship matrix between positive instances to mine the instances that cover the entire target, and assigning large loss weights to the instances that cover the entire target, encouraging the model to focus on the overall target.
[0038] Third, since the present invention introduces the task alignment bounding box regression loss, in the bounding box regression branch, the instability at the initial stage of model iteration and the difficulty degree of positive instances are considered simultaneously, so as to significantly improve the coupling between the regression branch and the classification branch. Brief Description of the Drawings
[0039] Figure 1 Is the overall implementation flowchart of the method of the present invention;
[0040] Figure 2Schematic diagram of the application process for target detection using the present invention;
[0041] Figure 3 Partial detection result diagrams obtained using the present invention on the DIOR dataset. Specific embodiments
[0042] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0043] Embodiment 1: Referring to Figure 1 , the present invention proposes a weakly supervised target detection method for remote sensing images based on weight redistribution and task alignment; specifically, it includes the following steps:
[0044] Step 1. Given a remote sensing image where H, W, and B represent the width, height, and number of channels of the image respectively; use the selective search algorithm to generate M target candidate boxes for the remote sensing image I, forming a target candidate box set PR = {pr1, pr2,... pr m ,... pr M}, where pr m represents the m-th target candidate box;
[0045] Step 2. Input the remote sensing image I and the target candidate box set PR into the online instance classifier optimization model OICR, which includes a backbone network, a multi-instance detection network, an instance classification optimization branch, and a task alignment bounding box regression branch; obtain the feature vectors of all target candidate boxes through the backbone network, region of interest pooling, and two fully connected layers and send them into the multi-instance detection network, where f m represents the feature vector of the m-th target candidate box;
[0046] Step 3. The target candidate box feature vector is input into two parallel branches in the multi-instance detection network. In this embodiment, for these two parallel branches, both are composed of a fully connected layer and a Softmax classifier; after passing through the two branches respectively, two score matrices in category c are obtained Obtain the category confidence score matrix of the target candidate box in category c according to the following formula
[0047]
[0048] where ⊙ represents the Hadamard product, and c is any one of all categories C;
[0049] Obtain the image-level category prediction score φ according to the following formula c :
[0050]
[0051] Among them, represents the class confidence score of the m-th target candidate box in class c;
[0052] Step 4. Add K parallel instance classification optimization branches to the baseline weakly supervised object detection model. By inputting the target candidate box feature vector into the k-th instance classification optimization branch, the class confidence score of the target candidate box is obtained where C + 1 represents the background class;
[0053] Step 5. Use the class confidence score of the (k - 1)-th branch as the input for high-quality positive instance mining. By sorting in descending order, select the top p% of the instances as the initial seed instance set for the c-th class in the k-th branch Then perform non-maximum suppression on it to remove small and redundant initial seed instances, and obtain the seed instance set to achieve seed instance mining;
[0054] Step 6. Mine initial neighborhood instances through spatial similarity. Set the spatial similarity determination threshold ζ. When the spatial similarity IoU between the j-th seed instance and other instances is greater than the threshold ζ, determine this instance as an initial neighborhood instance, otherwise determine it as a background instance; obtain the initial neighborhood instance set of the seed instance
[0055] Step 7. Remove misclassified initial neighborhood instances through feature similarity between instances. Set the feature similarity threshold δ. When the feature similarity between the j-th seed instance and the initial neighborhood instance is less than δ, determine this initial neighborhood instance as a background instance, and obtain the neighborhood instance set of the j-th seed instance Obtain the neighborhood instance sets of all seed instances to get the instance-level supervision information y k .
[0056] The feature similarity between the j-th seed instance and the initial neighborhood instance ini i is obtained according to the following formula:
[0057]
[0058] Among them, is and its initial neighborhood instance The feature similarity between, sim(·,·) represents the dot product operation, represents the feature vector of represents the initial neighborhood instance ini i the feature vector of; when the initial neighborhood instance ini i is recognized as a neighborhood instance, otherwise it is recognized as a background instance.
[0059] Step 8. Assign loss weights to the positive instances that cover the target completely through the positive instance weight redistribution strategy. First, construct a positive instance subset SN for the seed instance and its neighborhood instances, then calculate the spatial relationship between the instances in SN, and construct a spatial distance matrix Then sum G along the row direction to obtain a spatial distance vector Find the maximum value in V, and mine the instance r z that covers the target completely, and at the same time find the instance r p in the positive instance subset SN that has the highest predicted class confidence score. In the positive instance subset SN, swap the loss weights of the instances r z and r p to achieve positive instance weight redistribution.
[0060] The method of assigning loss weights to the positive instances that cover the target completely through the positive instance weight redistribution strategy is specifically implemented as follows:
[0061] (8.1) For the jth seed instance and its neighborhood instances construct a positive instance subset SN = {r1,...r n ,...,r N}, where r n represents the nth positive instance in the positive instance subset;
[0062] (8.2) By calculating the spatial relationship between the instances in the positive instance subset SN, construct a spatial distance matrix The spatial relationship between the instances is established according to the following formula:
[0063]
[0064] where, g n,n′ ∈G is the spatial distance between the positive instances r n and r n' , ∩ and ∪ respectively represent the intersection and union of two target candidate boxes, and r n′ represents the n'th positive instance in the positive instance subset, and n'≠n.
[0065] (8.3) For the spatial distance matrix Sum along the row direction to obtain the spatial distance vector
[0066]
[0067] (8.4) Find the maximum value in the spatial distance vector V and mine the complete instance r covering the target z , and its index is obtained as follows:
[0068] z = argmaxV,
[0069] Find the instance r in the subset SN of positive real examples with the highest predicted class confidence score p , and its index is obtained as follows:
[0070]
[0071] (8.5) In the subset SN of positive real examples, exchange the loss weights between instance r z and r p to achieve reallocation of positive instance weights, as follows:
[0072]
[0073] where, respectively represent the predicted class confidence scores of class c for instance r p and r z in the (k - 1)-th instance classification optimization branch;
[0074] Step 9. Construct the loss function of the k-th instance classification optimization branch
[0075]
[0076] where, represents the loss weight of the target candidate box m in the k-th branch, represents the element in the c-th row and m-th column of y k ; when is equal to 1, it means that the target candidate box m contains the target of class c; is equal to 0, it means that the target candidate box m does not contain the target of class c; represents the predicted class confidence score of class c for the target candidate box m in the k-th branch;
[0077] Step 10. The task-aligned bounding box regression branch performs bounding box regression on positive instances, considering the difficulty of positive instances and the number of model iterations, and increasing the coupling between the model classification task and the regression task, as follows:
[0078] (10.1) Design a modulation factor μ that changes dynamically with model iteration to ensure the stability of training:
[0079]
[0080] Where t represents the current number of iterations, T represents the total number of model iterations; σ represents the hyperparameter used to control the growth rate;
[0081] (10.2) Align the task of the k-th instance classification optimization branch with the bounding box regression loss The definition is as follows:
[0082]
[0083] Among them, P k represents the set of positive instances of the kth instance classification optimization branch, |P k | represents the number of elements in the positive instance set; represents the predicted category confidence score of the corresponding category of the lth positive instance in the kth instance classification optimization branch; and g l denote the target offset and predicted offset of the lth positive instance respectively;
[0084] Step 11. Construct the overall loss L of the weakly supervised remote sensing image target detection model ALL :
[0085]
[0086] Among them, L WSDDN represents the loss function of the baseline weakly supervised object detection network, K represents the number of instance classification optimization branches;
[0087] The loss function L of the baseline weakly supervised object detection network is WSDDN , as follows:
[0088]
[0089] Among them, C represents the number of target categories, y c =1 or 0 means that the input image contains or does not contain the target of the cth category.
[0090] Step 12. Use L ALL The entire weakly supervised remote sensing image target detection model is trained, and the remote sensing image to be tested and its target candidate box are input into the trained model to obtain the category and position of the target of interest in the image, thereby realizing weakly supervised target detection in remote sensing images.
[0091] Embodiment 2: The overall implementation steps of the method proposed in this embodiment are the same as those in Embodiment 1. Figure 2, which will be described separately from the aspects of model construction, high-quality positive instance mining, positive instance weight redistribution, and task-aligned bounding box regression loss in the present invention, and the implementation process of the present invention will be further described in detail:
[0092] Step 1: Construction of the online optimized model OICR of the instance classifier
[0093] First, as Figure 2 shown, given a remote sensing image where H, W, and B represent the width, height, and number of channels of the image respectively. Using the selective search algorithm [J.R. Uijlings, K.E. van de Sande, T. Gevers, and A.W. Smeulders. Selective search for object recognition. IJCV, 104(2):154–171, 2013.1, 3, 6] to generate M target candidate boxes PR = {pr1,... pr m ,... pr M} for the remote sensing image I, where pr m represents the m-th target candidate box.
[0094] Second, input the remote sensing image I and the set of target candidate boxes PR into the backbone network, region of interest pooling, and two fully connected layers to obtain the feature vectors of all target candidate boxes where f m represents the feature vector of the m-th target candidate box.
[0095] Third, input the target candidate box feature vectors into two parallel branches in the multi-instance detection network, and each branch consists of a fully connected layer and a Softmax classifier. When all the feature vectors pass through the two branches respectively, two score matrices in class c are obtained Thus far, the class confidence score matrix of all target candidate boxes in class c is obtained in the following way:
[0096]
[0097] where ⊙ represents the Hadamard product. The image-level class prediction score φ c is obtained in the following way:
[0098]
[0099] where, represents the class confidence score of the m-th target candidate box in class c. Thus far, the loss function L WSDDN of the baseline weakly supervised object detection network is expressed as follows:
[0100]
[0101] Among them, C represents the number of target categories, and y c = 1 or 0 indicates whether the target of the c-th category is included in the input image.
[0102] Fourth, as Figure 2 shown in the lower right corner, K parallel instance classification optimization branches are added to the baseline weakly supervised object detection model. The target candidate box feature vector is input into the k-th instance classification optimization branch to obtain the class confidence score of the target candidate box Among them, the (C + 1)-th dimension represents the background category.
[0103] Fifth, the supervision signal y of the k-th instance classification optimization branch k , is generated through (see Step 2). The loss function of the k-th instance classification optimization branch is as follows:
[0104]
[0105] Among them, represents the loss weight of the target candidate box m in the k-th branch (see Step 3). represents k the element in the c-th row and m-th column of y. When is equal to 1, it means that the target candidate box m contains the target of category c. Conversely, when is equal to 0, it means that the target candidate box m does not contain the target of category c. represents the predicted class confidence score of the target candidate box m for category c in the k-th branch.
[0106] Step 2: High-quality positive instances mining (HPIM)
[0107] Instance-level pseudo-labels include positive instances and negative instances. Positive instances include seed instances and neighborhood instances, and negative instances include background instances. The HPIM module first mines seed instances. After the seed instances are determined, during the label propagation process, neighborhood instances and background instances are determined based on the spatial similarity and feature similarity between the seed instances and other instances.
[0108] First, mine seed instances through . Among the categories c existing in the image, the target candidate boxes are sorted in descending order according to the scores. The top p percent of the instances are selected as the initial seed instance set for the c-th category in the k-th branch Secondly, non-maximum suppression operation is performed on the initial seed instances [J.Hosang, R.Benenson, and B.Schiele, “Learning non-maximum suppression,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., Jul. 2017, pp. 4507–4515.] to remove small and redundant initial seed instances, and a set of seed instances is obtained.
[0109] Second, neighborhood instances are mined through spatial similarity and feature similarity. First, when the spatial similarity between the j-th seed instance and other instances is large (i.e., IoU > 0.5), the instance is determined to be an initial neighborhood instance; otherwise, it is determined to be a background instance. Accordingly, a set of initial neighborhood instances of the seed instance is obtained. The seed instance and its initial neighborhood instance The feature similarity between them is expressed as follows:
[0110]
[0111] where sim(·,·) represents the dot product operation, represents the feature vector of the seed instance , and represents the feature vector of the initial neighborhood instance ini i . When , the initial neighborhood instance ini i is recognized as a neighborhood instance; otherwise, the initial neighborhood instance is recognized as a background instance. Accordingly, a set of neighborhood instances of the seed instance is obtained. The above operations are performed for all seed instances, and all target candidate boxes are divided into seed instances, neighborhood instances, and background instances. Seed instances and neighborhood instances form positive samples, and background instances form negative samples. Positive samples and negative samples constitute the supervision information y k of the k-th instance classification optimization branch.
[0112] Step 3: Positive instances weights re-assigning (PIWR)
[0113] First, for the seed instance and its neighborhood instance , a positive instance subset SN = {r1,...r n ,...,r N}, where r n represents the nth positive instance in the subset.
[0114] Second, by calculating the spatial relationships between the instances in the positive instance subset SN, construct a spatial distance matrix The spatial distance g n between the positive instances r n' and r n,n′ ∈G is expressed as follows:
[0115]
[0116] Third, the spatial distance vector is obtained by summing the spatial distance matrix along the row direction, and the expression is as follows:
[0117]
[0118] Fourth, by finding the maximum value in the spatial distance vector V, mine the instance r z that covers the target more completely. Its index is obtained in the following way:
[0119] z = argmaxV<8>
[0120] Fifth, find the instance r p in the positive instance subset SN that has the highest predicted class confidence score. Its index is obtained in the following way:
[0121]
[0122] Sixth, in the positive instance subset SN, exchange the loss weights between the instance r z that covers the target completely and the instance r p that has the highest class confidence score, to encourage the model to focus on the overall target. Specifically as follows:
[0123]
[0124] Among them, respectively represent the predicted class confidence scores of the instances r p and r z for class c in the (k - 1)th instance classification optimization branch.
[0125] Step 4: Task-aligned bounding box regression loss (TABRL)
[0126] First, as Figure 2As shown, a bounding box regression branch is added to each instance classification regression branch, and each bounding box regression branch contains a fully connected layer.
[0127] Second, the weakly supervised target detection model is unstable in the early stage of iteration. A modulation factor μ is designed that changes dynamically with the model iteration. Its definition is as follows:
[0128]
[0129] Among them, t represents the current number of iterations, T represents the total number of model iterations; σ represents a hyperparameter used to control the growth rate.
[0130] Third, the task-aligned bounding box regression loss combines the predicted class confidence score and the modulation factor to increase the coupling between the model classification task and the regression task. The definition is as follows:
[0131]
[0132] Among them, P k represents the set of positive instances of the kth instance classification optimization branch, |P k | represents the number of elements in the positive instance set. represents the predicted category confidence score of the corresponding category of the lth positive instance in the kth instance classification optimization branch, Used to measure the difficulty of the instance; the modulation factor μ is used to ensure the stability of training. and g l denote the target offset and predicted offset of the lth positive instance respectively;
[0133] Step 5: Total model loss
[0134] The overall loss L of the weakly supervised remote sensing image target detection model proposed in this paper is ALL The definition is as follows:
[0135]
[0136] Where K represents the number of instance classification optimization branches. ALL The entire weakly supervised remote sensing image target detection model is trained. In the inference stage, the remote sensing image to be tested and its target candidate box are input into the trained model to obtain the category and location of the target of interest in the image.
[0137] The effect of the present invention will be further described below in conjunction with experiments.
[0138] 1. Experimental conditions:
[0139] Hardware configuration for the implementation of the present invention: Experiments were conducted on a workstation with an E5-2650V4 CPU (2.2 GHz, 12x2 cores), 512 GB of memory, and 8 NVIDIA RTX Titan graphics cards. The software platform configuration: Ubuntu16.04, Python3.7, Pytorch1.7.
[0140] 2. Experimental content:
[0141] The detection performance of the present invention on the DIOR dataset is compared with the following 10 popular algorithms in the prior art: WSDDN [H. Bilen, A. Vedaldi, Weakly supervised deep detection networks, in: Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2016, pp. 2846–2854], OICR [P. Tang, X. Wang, X. Bai, W. Liu, Multiple instance detection network with online instance classifier refinement, in: Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2017, pp. 3059–295 3067], PCL [P. Tang, X. Wang, S. Bai, W. Shen, X. Bai, W. Liu, A. L. Yuille, PCL: proposal cluster learning for weakly supervised object detection, IEEE Trans. Pattern Anal. Mach. Intell. 42(1)(2020)176–191], MELM [F. Wan, P. Wei, J. Jiao, Z. Han, and Q. Ye, “Min-entropy latent model for weakly supervised object detection,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., Jun. 2018, pp. 1297–1306.], DCL [X. Yao, X. Feng, J. Han, G. Cheng, L. Guo, Automatic weakly supervised object detection from high spatial resolution remote sensing images via dynamic curriculum learning, IEEE Trans. Geosci. Remote Sens. 59(1)(2021)675–685], MIST [Z. Ren, Z. Yu, X. Yang, M.-Y. Liu, Y. J. Lee, A. G. Schwing, and J.Kautz, "Instance-aware, context-focused, and memory-efficient weakly supervised object detection," in Proc. IEEE / CVF Conf. Comput. Vis. Pattern Recognit., Jun. 2020, pp. 10598–10607.], PCIR [X. Feng, J. Han, X. Yao, G. Cheng, Progressive contextual instance refinement for weakly supervised object detection in remote sensing images, IEEE Trans. Geosci. Remote Sens. 58(11)(2020)8002–8012], TCA [X. Feng, J. Han, X. Yao, and G. Cheng, "Tcanet: Triple context-aware network for weakly supervised object detection in remote sensing images," IEEE Trans. Geosci. Remote Sens., vol. 59, no. 8, pp. 6946–6955, Oct. 2021.], MIG [B. Wang, Y. Zhao, X. Li, Multiple instance graph learning for weakly supervised remote sensing object detection, IEEE Trans. Geosci. Remote Sens. 60(2022)1–12], SPG [G. Cheng, X. Xie, W. Chen, X. Feng, X. Yao, and J. Han, "Self guided proposal generation for weakly supervised object detection," IEEE Trans. Geosci. Remote Sens., vol. 60, 2022, Art. no. 5625311.]. mAP and Corloc represent the mean average precision and the localization precision respectively.
[0142] 3. Simulation results and analysis:
[0143] On the DIOR dataset, the present invention was compared with 10 popular algorithms in the prior art, and the comparison results regarding the mean average precision (mAP) and correct localization (CorLoc) are shown in Table 1 as follows:
[0144] Table 1
[0145] Method mAP CorLoc WSDDN 13.3 32.4 OICR 16.5 34.8 PCL 18.2 41.5 MELM 18.7 43.3 DCL 20.2 42.2 MIST 22.2 43.6 PCIR 24.9 46.1 TCA 25.8 48.4 MIG 25.1 46.8 SPG 25.8 48.3 The present invention 28.9 53.9
[0146] The above simulation analysis proves the correctness and effectiveness of the method proposed by the present invention. As can be seen from Table 1, the mean average precision (mAP) of the present invention is 15.6%, 12.4%, 10.7%, 10.2%, 8.7%, 6.7%, 4.0%, 3.1%, 3.8%, 3.1% higher than that of WSDDN, OICR, PCL, EELM, DCL, MIST, PCIR, TCA, MIG, SPG respectively, and the correct localization (CorLoc) is 21.5%, 19.1%, 12.4%, 10.6%, 11.7%, 10.3%, 7.8%, 5.5%, 7.1%, 5.6% higher than that of WSDDN, OICR, PCL, EELM, DCL, MIST, PCIR, TCA, MIG, SPG respectively. Thus, it can be seen that the model proposed by the present invention can more accurately identify and locate the ground object targets in remote sensing images compared with the prior art, effectively improving the detection effect of weakly supervised targets in remote sensing images.
[0147] Weakly supervised remote sensing image target detection shows great potential in the fields of military, urban planning, disaster detection, traffic construction, and ecological research. The present invention can accurately detect specific targets in remote sensing images containing rich information and unique perspectives, solving the problem of the large amount of manpower and material resources required for supervised learning that depends on complete annotation information, making remote sensing image target detection more efficient and economical, which is crucial for decision-making support and practical operations in various fields. The proposed method of the present invention provides more accurate and efficient support for the application field of weakly supervised remote sensing image target detection.
[0148] The parts not described in detail in the present invention belong to the common general knowledge of those skilled in the art.
[0149] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Obviously, for professionals in the field, after understanding the content and principle of the present invention, various modifications and changes in form and details may be made without departing from the principle and structure of the present invention. However, these modifications and changes based on the idea of the present invention are still within the protection scope of the claims of the present invention.
Claims
1. A weakly supervised target detection method for remote sensing images based on weight redistribution and task alignment, characterized in that: The implementation steps include the following: (1) Given a remote sensing image Where H, W and B represent the width, height and number of channels of the image respectively. The selective search algorithm is used to generate M target candidate frames for the remote sensing image I, forming a target candidate frame set PR = {pr1, pr2, ...pr m ,…pr M }, where pr m Represents the mth target candidate box; (2) The remote sensing image I and the target candidate box set PR are input into the instance classifier online optimization model OICR, which includes a backbone network, a multi-instance detection network, an instance classification optimization branch, and a task alignment bounding box regression branch; the feature vectors of all target candidate boxes are obtained through the backbone network, the region of interest pooling, and two fully connected layers. And feed it into the multi-instance detection network, where f m Represents the feature vector of the mth target candidate box; (3) Target candidate box feature vector Input two parallel branches in the multi-instance detection network, and obtain two score matrices in category c after passing through the two branches respectively. Obtain the category confidence score matrix of the target candidate box in category c according to the following formula: Where ⊙ represents the Hadamard product, c is any one of all categories C; The image-level category prediction score φ is obtained according to the following formula c : in, Represents the category confidence score of the mth target candidate box in category c; (4) Add K parallel instance classification optimization branches to the baseline weakly supervised object detection model, and Input the kth instance classification optimization branch to obtain the category confidence score of the target candidate box Among them, C+1 represents the background category; (5) The category confidence score of the k-1th branch As the input for high-quality positive instance mining, we get the initial seed instance set Then perform non-maximum suppression operation on it to obtain a set of seed instances Implement seed instance mining; (6) Mining the initial neighborhood instances through spatial similarity, setting the spatial similarity judgment threshold ζ, when the jth seed instance When the spatial similarity IoU between the instance and other instances is greater than the threshold ζ, the instance is determined to be the initial neighborhood instance, otherwise it is determined to be the background instance; the seed instance is obtained. The initial neighborhood instance set (7) By using the feature similarity between instances, we can eliminate the misclassified initial neighborhood instances and set the feature similarity threshold δ. When the jth seed instance When the feature similarity between the initial neighborhood instance and the background instance is less than δ, the initial neighborhood instance is judged as a background instance and the jth seed instance is obtained. The set of neighborhood instances of Get the neighborhood instance set of all seed instances and obtain the instance-level supervision information y of the kth instance classification optimization branch k ; (8) The positive instance weight redistribution strategy is used to assign loss weights to positive instances that cover the entire target. First, a positive instance subset SN is constructed for the seed instance and its neighboring instances. Then, the spatial relationship between instances in SN is calculated to construct a spatial distance matrix Then sum up G along the row direction to get the spatial distance vector Find the maximum value in V and mine the complete instance r covering the target z , and find the instance r with the highest prediction category confidence score in the positive instance subset SN p , in the positive instance subset SN, instance r z and r p The loss weights are exchanged to achieve positive instance weight redistribution; (9) Construct the loss function of the kth instance classification optimization branch in, represents the loss weight of the target candidate box m in the kth branch, Represents y k The element in the cth row and mth column of When it is equal to 1, it means that the target candidate box m contains the target of category c; When it is equal to 0, it means that the target candidate box m does not contain the target of category c; Represents the predicted category confidence score of the target candidate box m on category c in the kth branch; (10) The task-aligned bounding box regression branch performs bounding box regression on the positive instance, taking into account the difficulty of the positive instance and the number of model iterations, and increases the coupling between the model classification task and the regression task, as follows: (10.1) Design a modulation factor μ that changes dynamically with model iteration to ensure the stability of training: Where t represents the current number of iterations, T represents the total number of model iterations; σ represents the hyperparameter used to control the growth rate; (10.2) Align the task of the k-th instance classification optimization branch with the bounding box regression loss The definition is as follows: Among them, P k represents the set of positive instances of the kth instance classification optimization branch, |P k | represents the number of elements in the positive instance set; represents the predicted category confidence score of the corresponding category of the lth positive instance in the kth instance classification optimization branch; and g l denote the target offset and predicted offset of the lth positive instance respectively; (11) Construct the overall loss L of the weakly supervised remote sensing image target detection model ALL : Among them, L WSDDN represents the loss function of the baseline weakly supervised object detection network, which is based on the image-level category prediction score φ obtained in step (3). c Calculated; K represents the number of instance classification optimization branches; (12) Using L ALL The entire weakly supervised remote sensing image target detection model is trained, and the remote sensing image to be tested and its target candidate box are input into the trained model to obtain the category and position of the target of interest in the image, thereby realizing weakly supervised target detection in remote sensing images.
2. The method according to claim 1, characterized in that: The two parallel branches in the multi-instance detection network described in step (3) are both composed of a fully connected layer and a Softmax classifier.
3. The method according to claim 1, characterized in that: The initial seed instance set described in step (5) Specifically, through Sort in descending order, and then select the first p percent of instances as the initial seed instance set of the cth category in the kth branch 4. The method according to claim 1, characterized in that: The jth seed instance in step (7) and the initial neighborhood instance ini i The feature similarity between them is obtained according to the following formula: in, for and its initial neighborhood instance The feature similarity between them, sim(·,·) represents the dot product operation, express The characteristic vector of Represents the initial neighborhood instance ini i The characteristic vector of When the initial neighborhood instance ini i are identified as neighborhood instances, whereas the others are identified as background instances.
5. The method according to claim 1, characterized in that: In step (8), the loss weights are assigned to the positive instances that completely cover the target through the positive instance weight redistribution strategy. The specific implementation steps are as follows: (8.1) For the jth seed instance and its neighborhood instances Construct a positive instance subset SN = {r1,...r n ,...,r N }, where r n represents the nth positive instance in the positive instance subset; (8.2) By calculating the spatial relationship between instances in the positive instance subset SN, a spatial distance matrix is constructed. (8.3) For the spatial distance matrix Sum along the row direction to get the spatial distance vector (8.4) Find the maximum value in the spatial distance vector V and mine the complete instance r covering the target z , whose index is obtained by: z=argmaxV, Find the instance r with the highest prediction category confidence score in the positive instance subset SN p , whose index is obtained by: (8.5) In the positive instance subset SN, instance r z and r p The loss weights between are exchanged to achieve positive instance weight redistribution, as follows: in, Respectively represent the instance r p and r z The predicted class confidence score for class c in the k-1th instance classification optimization branch.
6. The method according to claim 5, characterized in that: The spatial relationship between the instances described in step (8.2) is established according to the following formula: Among them, g n,n′ ∈G is a positive instance r n and r n' The spatial distance between them, ∩ and ∪ represent the intersection and union of two target candidate boxes, respectively, r n′ represents the n'th positive instance in the positive instance subset, and n'≠n.
7. The method according to claim 1, characterized in that: The loss function L of the baseline weakly supervised target detection network described in step (11) is WSDDN , as follows: Among them, C represents the number of target categories, y c =1 or 0 means that the input image contains or does not contain the target of the cth category.
Citation Information
Patent Citations
Weak supervision target detection method and system
CN116452877A
Remote sensing image weak supervision target detection method based on complete multi-instance learning and soft label definition
CN118608839A