Method and system for realizing remote sensing image target detection based on deep learning
By distinguishing the determination box from the uncertain box in the remote sensing image object detection and rotating, scaling and probability adjustment, the problems of background changes and multi-scale targets are solved, and the accuracy of object detection is improved.
Patent Information
- Application Number
- CN202510094228.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-16
AI Technical Summary
Remote sensing image object detection faces problems of background changes and multi-scale targets, resulting in increased difficulty in target recognition, and there is an imbalance in the detection effect of deep neural network models.
By obtaining the target probability of the recommended box selection area in the remote sensing image, distinguish it into a certain box and an uncertain box, forming a box pair and filtering out the retained box pair. Rotate and scale the uncertain box, obtain the final uncertain box, and calculate the target probability adjustment coefficient based on its difference from the determination box, and filter out the determination box.
Through the comparative analysis of the determination box and the uncertain box, the accuracy of object detection in the remote sensing image is improved, ensuring further judgment and processing of uncertain box that may be missed.
Smart Images

Figure CN120014494A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image recognition technology, specifically to a method and system for remote sensing image target detection based on deep learning. Background Technology
[0002] Target detection in remote sensing images based on deep learning refers to the process of locating and identifying targets in remote sensing images using deep neural networks. This method automatically extracts image features to identify and locate various targets in images, such as buildings, roads, and vehicles, and is widely used in fields such as urban planning, environmental monitoring, and agricultural management.
[0003] Existing problems: In practical applications, target detection in remote sensing images faces several challenges, one of the main issues being the impact of background changes on target recognition. Due to the complex and varied backgrounds of targets in remote sensing images, such as lighting conditions, seasonal changes, and topography, factors that could make normally easily identifiable targets difficult to identify. Furthermore, the multi-scale nature of targets is also a limiting factor for model performance. For example, targets such as vehicles and ships may be very small in the image, directly affecting the detection performance of deep neural network models for these types of targets. Moreover, the scale difference between large and small targets makes it difficult for the relatively fixed perceptual field of view of deep neural networks to achieve a balance between these two types of targets. Summary of the Invention
[0004] This invention provides a method and system for target detection in remote sensing images based on deep learning, in order to solve existing problems.
[0005] The method and system for target detection in remote sensing images based on deep learning of the present invention adopts the following technical solution:
[0006] One embodiment of the present invention provides a method for target detection in remote sensing images based on deep learning, the method comprising the following steps:
[0007] Acquire remote sensing imagery, use a target detection network to obtain several suggested bounding boxes in the remote sensing imagery and the target probability of each suggested bounding box; based on the target probability of each suggested bounding box, classify the suggested bounding boxes into confirmed boxes and uncertain boxes;
[0008] A box pair is formed by combining any defined box and any uncertain box, and several unique box pairs are obtained. Based on the difference between the edge pixels of the defined box and the uncertain box in each box pair, several box pairs are selected from all box pairs to be retained.
[0009] Rotate and scale the uncertain boxes in each pair of retained boxes to obtain the final uncertain boxes in each pair of retained boxes;
[0010] Based on the difference between the final uncertain bounding box and the edge pixels of the determined bounding box in each retained bounding box pair, the target probability adjustment coefficient of the uncertain bounding box in each retained bounding box pair is determined; based on the target probability adjustment coefficient and target probability of the uncertain bounding box in each retained bounding box pair, several determined bounding boxes are selected from the uncertain bounding boxes in all retained bounding box pairs; all determined bounding boxes in the remote sensing image are used as the target detection area in the remote sensing image.
[0011] Furthermore, the specific steps for distinguishing the suggested selection area into a confirmed box and an uncertain box based on the target probability of each suggested selection area are as follows:
[0012] The suggested selection area with a target probability greater than the preset probability threshold is denoted as the confirmation box;
[0013] The suggested selection area with a target probability less than or equal to a preset probability threshold is denoted as the uncertain box.
[0014] Furthermore, the specific steps for selecting several retained box pairs from all box pairs based on the difference between the edge pixels of the determined box and the uncertain box in each box pair are as follows:
[0015] The Canny edge detection operator is used to perform edge detection on the i-th suggested selection area to obtain several edge pixels within the i-th suggested selection area.
[0016] Based on the grayscale values of the pixels within the i-th suggested selection area, the LBP algorithm is used to obtain the LBP feature value of each edge pixel within the i-th suggested selection area;
[0017] Within the i-th suggested selection area, based on the LBP feature values of the edge pixels, a hierarchical clustering algorithm is used to cluster all edge pixels to obtain the hierarchical clustering tree corresponding to the i-th suggested selection area.
[0018] The APTED+ algorithm is used to obtain the edit distance between the hierarchical clustering tree corresponding to the determined box and the hierarchical clustering tree corresponding to the uncertain box in each box pair, which is used as the non-retention factor for each box pair.
[0019] Determine whether each box pair is a retained box pair based on the non-retention factor for each box pair.
[0020] Furthermore, the specific steps for determining whether each box pair is a retained box pair based on the non-retention factor of each box pair are as follows:
[0021] Pairs of boxes whose non-retention factor is less than the preset retention threshold are denoted as retained pairs.
[0022] Furthermore, the specific steps involved in rotating and scaling the uncertain boxes in each pair of retained boxes to obtain the final uncertain boxes in each pair are as follows:
[0023] Based on the edge pixels within each suggested selection area, the PCA algorithm is used to obtain the principal direction of the edge pixels, which is then used as the principal direction of each suggested selection area.
[0024] Rotate the uncertain frame in any pair of retained frames so that the main direction of the uncertain frame in the pair of retained frames is consistent with the main direction of the determined frame, and obtain the rotated uncertain frame, which is used as the rotated uncertain frame of the uncertain frame in the pair of retained frames.
[0025] The hierarchical clustering tree contains several layers, each layer contains several categories, and each category contains several edge pixels;
[0026] In each category of the hierarchical clustering tree corresponding to each suggested selection area, the number of edge pixels of each category is sorted in ascending order to obtain the ascending order sequence of the number;
[0027] In the ascending sequence of quantities, calculate the ratio of each quantity to the first quantity to obtain the category quantity ratio sequence;
[0028] In each reserved box pair, the left node is the sequence of the number of categories in each layer of the hierarchical clustering tree corresponding to the determined box, and the right node is the sequence of the number of categories in each layer of the hierarchical clustering tree corresponding to the uncertain box. The KM matching algorithm is used to obtain several matching pairs and the edge value of each matching pair; each matching pair contains a left node and a right node.
[0029] Among all matching pairs of the j-th reserved box pair, the matching pair with the largest edge value is denoted as the scale matching pair;
[0030] Based on the ascending sequence of the number of left and right nodes in the scale matching pair, obtain the scaling factor A of the j-th preserved box pair. j ;
[0031] Using bilinear interpolation, the rotation uncertainty of the uncertain box in the j-th retained box pair is scaled by A. j The uncertainty box is obtained by rotating and scaling the size of the uncertainty box, which serves as the final uncertainty box.
[0032] Furthermore, the scaling factor A of the j-th retained box pair is obtained based on the ascending sequence of the number of left and right side nodes in the scale matching pair. j The specific steps include the following:
[0033] The average of all quantities in the ascending sequence corresponding to the left node in the scale matching pair is denoted as the target scale N of the bounding box in the j-th reserved box pair. j ;
[0034] The mean of all quantities in the ascending sequence corresponding to the right node in the scale matching pair is denoted as the target scale M of the uncertain box in the j-th reserved box pair. j ;
[0035] The N j With the M j The ratio, denoted as the scaling factor A of the j-th reserved frame pair. j .
[0036] Furthermore, the specific steps for determining the target probability adjustment coefficient of the uncertain box in each pair of retained boxes based on the difference between the final uncertain box and the edge pixels within the determined box are as follows:
[0037] In the j-th reserved box pair, based on the length C1 and width K1 of the determined box and the length C2 and width K2 of the final uncertain box of the uncertain box, the segmentation determined box and the segmentation uncertain box are obtained;
[0038] Set the grayscale value of edge pixels within the segmentation box to 1 and the grayscale value of non-edge pixels to 0 to obtain the target binary image;
[0039] Set the grayscale value of edge pixels within the segmentation uncertainty box to 1 and the grayscale value of non-edge pixels to 0 to obtain a reference binary image;
[0040] Calculate the product of the target binary image and the reference binary image to obtain a binary image with the same edge;
[0041] Calculate the product of the binary image with the same edge and the segmentation bounding box to obtain the target segmentation bounding box;
[0042] Calculate the product of the binary image with the same edge and the segmentation uncertainty box to obtain the target segmentation uncertainty box;
[0043] The inversely proportional normalized value of the mean square error between the target segmentation determined box and the target segmentation uncertain box is denoted as the target probability adjustment coefficient for the j-th retained box to center the uncertain box.
[0044] Further, in the j-th retained box pair, the specific steps for obtaining the segmentation determined box and the segmentation uncertain box based on the length C1 and width K1 of the determined box and the length C2 and width K2 of the final uncertain box of the uncertain box are as follows:
[0045] In the j-th reserved box pair, obtain the length C1 and width K1 of the determined box and the length C2 and width K2 of the final uncertain box of the uncertain box;
[0046] Within the defined bounding box, a segmented bounding box with length min{C1,C2} and width min{K1,K2} is formed with the center of the defined bounding box as the center; the segmented bounding box is parallel to the width side and the length side of the defined bounding box; where min{} is a function for taking the minimum value;
[0047] Within the final uncertain frame of the uncertain frame, a segmented uncertain frame with length min{C1,C2} and width min{K1,K2} is formed with the center of the final uncertain frame as the center; the segmented uncertain frame is parallel to the wide side and the long side of the final uncertain frame.
[0048] Furthermore, the specific steps for selecting several certain boxes from the uncertain boxes in all retained box pairs based on the target probability adjustment coefficient and target probability of the uncertain boxes in each retained box pair are as follows:
[0049] In the j-th reserved box pair, calculate the sum of the target probability adjustment coefficients of 1 and the uncertain box, and multiply the sum by the target probability of the uncertain box as the updated target probability;
[0050] When the probability of the updated target is greater than a preset probability threshold, the uncertain box is recorded as a determined box.
[0051] This invention also proposes a remote sensing image target detection system based on deep learning, including a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor executes the computer program stored in the memory to implement the steps of the aforementioned remote sensing image target detection method based on deep learning.
[0052] The beneficial effects of the technical solution of the present invention are:
[0053] In this embodiment of the invention, defined bounding boxes and uncertain bounding boxes are obtained from remote sensing images. A box pair is formed by any defined bounding box and any uncertain bounding box. Several retained box pairs are selected from all box pairs. Through similarity analysis of the edge textures within the defined and uncertain bounding boxes, uncertain bounding boxes that may have been missed are identified, thus ensuring the accuracy of subsequent target detection. The uncertain bounding boxes in each retained box pair are rotated and scaled to obtain the final uncertain bounding box. This ensures that the edge texture direction and scale within the defined and uncertain bounding boxes are consistent, guaranteeing the reliability of subsequent comparison analysis between uncertain and defined bounding boxes. Based on the difference between the final uncertain bounding box and the defined bounding box, a target probability adjustment coefficient for the uncertain bounding box is obtained. Combined with the target probability, several defined bounding boxes are selected from the uncertain bounding boxes in all retained box pairs. This updates the target probability of the uncertain bounding boxes, re-evaluating uncertain bounding boxes that may have been missed, thereby improving the accuracy of target detection. All defined bounding boxes in the remote sensing image are used as the target detection area in the remote sensing image. Thus, this invention, through comparative analysis of defined and uncertain bounding boxes, further identifies uncertain bounding boxes that may have been missed, improving the accuracy of target detection in remote sensing images. Attached Figure Description
[0054] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0055] Figure 1 This is a flowchart illustrating the steps of the remote sensing image target detection method based on deep learning in this invention.
[0056] Figure 2 This is a remote sensing image showing the detection results of vehicles as targets.
[0057] Figure 3 This is a schematic diagram of a hierarchical clustering tree;
[0058] Figure 4 A schematic diagram of the concentric division within the defined box. Detailed Implementation
[0059] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of the deep learning-based remote sensing image target detection method and system proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.
[0060] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0061] The following description, in conjunction with the accompanying drawings, details the specific scheme of the remote sensing image target detection method and system based on deep learning provided by this invention.
[0062] Please see Figure 1 The diagram illustrates a flowchart of a method for remote sensing image target detection based on deep learning, according to an embodiment of the present invention. The method includes the following steps:
[0063] Step S001: Acquire remote sensing imagery, use a target detection network to acquire several suggested bounding boxes in the remote sensing imagery and the target probability of each suggested bounding box; based on the target probability of each suggested bounding box, classify the suggested bounding boxes into confirmed boxes and uncertain boxes.
[0064] Remote sensing images of any factory area are collected by a camera mounted on a drone. A target detection network is used to obtain several suggested bounding boxes in the remote sensing images, and each suggested bounding box corresponds to a target probability.
[0065] The suggested selection area with a target probability greater than the preset probability threshold is denoted as the "confirmed" area. The suggested selection area with a target probability less than or equal to the preset probability threshold is denoted as the "uncertain" area.
[0066] It should be noted that in this embodiment, the preset probability threshold is 0.7, which is used as an example for description. The suggested bounding boxes in the remote sensing image are rectangular image blocks in the remote sensing image. The remote sensing image is converted to grayscale, which is a well-known technique. The target detection network is a well-known technique. The specific operation of the target detection network in target recognition is as follows: multiple remote sensing images are collected as training datasets, and then the network extracts features from the input remote sensing images. These features can capture information such as texture, color, and shape in the remote sensing images, providing a basis for subsequent target recognition. For target detection networks (such as Faster R-CNN, whose Chinese name is Fast Region Convolutional Neural Network), multiple region proposals are generated, which are the suggested bounding boxes in the remote sensing images mentioned above. These proposals define the regions in the remote sensing image that may contain targets, and the network classifies each region proposal (suggested bounding box), calculates the probability that each box contains a specific target category, and sets a probability threshold. Only when the probability of a target exceeds this threshold is the target considered detected, i.e., a bounding box is determined. Ultimately, the network outputs a series of bounding boxes, each containing the target's category and probability. These boxes are used to "frame" targets in remote sensing imagery as detection results. An example of remote sensing imagery detection results with vehicles as targets is shown below. Figure 2 As shown, Figure 2 The vehicles detected are shown within the rectangular frame.
[0067] Step S002: Form a box pair with any defined box and any uncertain box, and obtain several non-repeating box pairs; based on the difference between the edge pixels of the defined box and the uncertain box in each box pair, select several box pairs to keep from all box pairs.
[0068] It should be noted that when using neural networks for target detection, especially in remote sensing image analysis, missed detections often occur due to complex backgrounds, large variations in target size, and low contrast between the target and background. These problems typically prevent the model from accurately identifying all targets, especially when the target is small, has a complex shape, or is similar in color to the background. Therefore, this embodiment further identifies low-probability targets by using high-probability targets to avoid missed detections; that is, it performs a comparative analysis of defined and undefined bounding boxes in the remote sensing image.
[0069] Preferably, in one embodiment of the present invention, the method for obtaining the reserved frame pairs includes:
[0070] Taking the i-th suggested bounding box region in the remote sensing image as an example, the Canny edge detection operator is used to perform edge detection on the i-th suggested bounding box region to obtain several edge pixels within the i-th suggested bounding box region.
[0071] Based on the grayscale values of the pixels within the i-th suggested selection area, the LBP algorithm is used to obtain the LBP feature value of each edge pixel within the i-th suggested selection area.
[0072] Within the i-th suggested bounding box region, based on the LBP feature values of the edge pixels, a hierarchical clustering algorithm is used to cluster all edge pixels, obtaining the hierarchical clustering tree corresponding to the i-th suggested bounding box region. The hierarchical clustering tree contains several layers, each layer contains several categories, and each category contains several edge pixels. Therefore, the number of edge pixels in each category within each layer of the hierarchical clustering tree can be determined.
[0073] It should be noted that the Canny edge detection operator, LBP (Local Binary Pattern) algorithm, and hierarchical clustering algorithm are all well-known techniques, and their specific methods will not be described here. The LBP feature value of an edge pixel represents the local texture information of that edge pixel. A schematic diagram of a hierarchical clustering tree is shown below. Figure 3 As shown, Figure 3 The horizontal axis represents the edge pixels within the suggested selection area, and the vertical axis represents the difference in LBP feature values between the edge pixels within the suggested selection area. In this embodiment, agglomerative hierarchical clustering (bottom-up) is used, with each level of the hierarchical clustering tree representing the result of one clustering operation.
[0074] Following the above method, obtain the hierarchical clustering tree corresponding to each suggested selection area, as well as the number of edge pixels of each category in each layer of the hierarchical clustering tree.
[0075] In remote sensing imagery, a pair of boxes is formed by combining any defined box and any indeterminate box, resulting in several unique box pairs.
[0076] The APTED+ algorithm is used to obtain the edit distance between the hierarchical clustering tree corresponding to the determined box and the hierarchical clustering tree corresponding to the uncertain box in each box pair, which is used as the non-retention factor for each box pair.
[0077] It should be noted that APTED+ (Approximate Tree Edit Distance Plus) is an algorithm for calculating tree edit distance. This is a well-known technology, and the specific method will not be introduced here. The smaller the edit distance, the more similar the two trees are. This means that the edge information within the determined box and the uncertain box in the bounding box is more similar. The uncertain box is more likely to be missed due to background influence.
[0078] The preset retention threshold is 0.5, and we will use this as an example for explanation.
[0079] Pairs of boxes whose non-retention factor is less than the preset retention threshold are denoted as retained pairs.
[0080] Step S003: Perform rotation and scaling operations on the uncertain boxes in each pair of retained boxes to obtain the final uncertain boxes of the uncertain boxes in each pair of retained boxes.
[0081] Preferably, in one embodiment of the present invention, the method for obtaining the final uncertain box of the uncertain box in each pair of retained boxes includes:
[0082] In remote sensing imagery, the PCA algorithm is used to obtain the principal direction of the edge pixels within each proposed bounding box area, which is then used as the principal direction of each proposed bounding box area.
[0083] It should be noted that the PCA (Principal Component Analysis) algorithm is a well-known technique, and its specific method will not be described here. The principal direction of the edge pixels obtained by the PCA algorithm is the principal direction of the distribution of edge pixels within the proposed selection area.
[0084] In each pair of retained frames, the uncertain frame is rotated so that the principal direction of the uncertain frame is consistent with the principal direction of the determined frame, resulting in a rotated uncertain frame, which is used as the rotated uncertain frame of the uncertain frame.
[0085] It should be noted that the uncertain bounding box in remote sensing imagery is a rectangular image patch. Rotating the uncertain bounding box is equivalent to rotating the rectangular image patch. For example, if the main direction of the uncertain bounding box centered by the retained bounding box is horizontal to the right, and the main direction of the determined bounding box is vertically upward, then the center of the uncertain bounding box is used as the reference point for rotation, and the uncertain bounding box is rotated 90 degrees counterclockwise to obtain the rotated uncertain bounding box. This ensures that the main directions of the pixel distribution within the determined bounding box and the uncertain bounding box are the same.
[0086] Since it is uncertain whether the target size is the same in the defined box and the uncertain box, it is necessary to determine the target size ratio, then scale the target in the uncertain box, and then compare it with the edge texture in the defined box.
[0087] In each level of the hierarchical clustering tree corresponding to each suggested selection area, the number of edge pixels of each category is sorted in ascending order to obtain an ascending sequence of numbers. In the ascending sequence of numbers, the ratio of each number to the first number is calculated to obtain a sequence of category number ratios.
[0088] For example, if the ascending sequence of quantities is {2, 3, 6, 8}, then the sequence of category quantity ratios is {1, 1.5, 3, 4}.
[0089] In each reserved box pair, the left node is the sequence of the number of categories in each layer of the hierarchical clustering tree corresponding to the determined box, and the right node is the sequence of the number of categories in each layer of the hierarchical clustering tree corresponding to the uncertain box. The KM matching algorithm is used to obtain several matching pairs and the edge values of each matching pair, wherein each matching pair contains a left node and a right node.
[0090] It should be noted that the KM matching algorithm refers to the Kuhn-Munkres algorithm, also known as the Hungarian algorithm. It is a well-known and efficient algorithm for solving assignment problems. Specifically, it involves assuming that every node on the left is connected to all nodes on the right by an edge, with the edge value being the cosine similarity between the two sequences (cosine similarity is a known calculation method). Then, through KM matching, following the principle of maximum matching, one-to-one matching pairs are obtained between the left and right nodes. These one-to-one matching pairs indicate at which scales the textures in the determined bounding box and the textures in the uncertain bounding box are similar, thus obtaining the texture scale ratio.
[0091] Taking the j-th reserved box pair as an example, among all matching pairs of the j-th reserved box pair, the matching pair with the largest edge value is denoted as the scale matching pair. The average of all quantities in the ascending sequence of the quantities corresponding to the left nodes in the scale matching pair is denoted as the target scale N of the bounding box in the j-th reserved box pair. j Let the mean of all quantities in the ascending sequence corresponding to the right node in the scale matching pair be denoted as the target scale M of the uncertain box in the j-th preserved box pair. j .
[0092] The target scale N of the bounding box is determined by centering the j-th bounding box. j The target scale M of the uncertain box aligned with the j-th reserved box j The ratio, denoted as the scaling factor A of the j-th reserved frame pair. j .
[0093] In the j-th retained box pair, the rotational uncertainty of the uncertainty box is scaled A using bilinear interpolation. j The uncertainty box is obtained by rotating and scaling the size of the uncertainty box, which serves as the final uncertainty box.
[0094] It should be noted that the bilinear interpolation algorithm is a well-known technique, and its specific method will not be described here. When N j Greater than M j At that time, A j The value is greater than 1, therefore the rotated uncertain frame is enlarged by A. j Times, when N j Less than M j At that time, A j The value is less than 1, therefore the rotated uncertainty box is reduced by A. j Times, when N jEqual to M j When the size of the uncertain box remains unchanged after rotation, the main direction of the distribution of edge pixels in the final uncertain box and the determined box in the retained box pair is the same, and the scale of the edge texture is the same. Then, the Canny edge detection operator is used to re-obtain the edge pixels in the final uncertain box of the uncertain box.
[0095] Step S004: Based on the difference between the final uncertain box and the edge pixels of the determined box in each retained box pair, determine the target probability adjustment coefficient of the uncertain box in each retained box pair; based on the target probability adjustment coefficient and target probability of the uncertain box in each retained box pair, select several determined boxes from the uncertain boxes in all retained box pairs; use all determined boxes in the remote sensing image as the target detection area in the remote sensing image.
[0096] The above only adjusts the scale and main direction of the edge texture of the centering bounding box and the uncertain bounding box to be the same. Due to the influence of the background inside the box, the final size of the uncertain bounding box and the determined bounding box may be different. However, it is recommended that the selected area be a box with the target as the center and the background as the surrounding area. Therefore, the centering bounding box and the final uncertain bounding box can be divided. While minimizing the impact on the edge texture inside the box, the two boxes can be made to be the same size.
[0097] Preferably, in one embodiment of the present invention, the method for obtaining the target detection region in a remote sensing image includes:
[0098] In the j-th reserved box pair, obtain the length C1 and width K1 of the defined box, and the length C2 and width K2 of the final uncertain box of the uncertain box. Within the defined box, using the center of the defined box as the center, construct a segmented defined box with length min{C1,C2} and width min{K1,K2}, where the segmented defined box is parallel to both the width and length sides of the defined box. Within the final uncertain box of the uncertain box, using the center of the final uncertain box as the center, construct a segmented uncertain box with length min{C1,C2} and width min{K1,K2}, where the segmented uncertain box is parallel to both the width and length sides of the final uncertain box of the uncertain box. Here, min{} is the minimum value function.
[0099] It should be noted that, taking the defined box as an example, the schematic diagram of the concentric division defined box within the defined box is as follows: Figure 4 As shown. At this point, in the j-th retained box pair, the defined segmentation box and the uncertain segmentation box are the same size, and the middle region containing the target is retained, while the boundary region with a large amount of background is removed.
[0100] In the j-th retained box pair, set the grayscale value of edge pixels within the segmentation determined box to 1 and the grayscale value of non-edge pixels to 0 to obtain the target binary image. Set the grayscale value of edge pixels within the segmentation uncertain box to 1 and the grayscale value of non-edge pixels to 0 to obtain the reference binary image. Calculate the product of the target binary image and the reference binary image to obtain the binary image with the same edge. Calculate the product of the binary image with the same edge and the segmentation determined box to obtain the target segmentation determined box. Calculate the product of the binary image with the same edge and the segmentation uncertain box to obtain the target segmentation uncertain box. The inversely proportional normalized value of the mean square error (MSE) of the target segmentation determined box and the target segmentation uncertain box is denoted as the target probability adjustment coefficient of the uncertain box in the j-th retained box pair.
[0101] It should be noted that: in a binary image with the same edge, a pixel with a grayscale value of 1 represents a pixel that is an edge pixel at the same position within both the defined and uncertain segments. Since both the defined and uncertain segments are grayscale images, the grayscale values of pixels at the same position within both the defined and uncertain segments are retained, while the grayscale values of other pixels are 0, thus removing the influence of the background within the frame. The calculation of the mean squared error (MSE) is a well-known method. The smaller the MSE between images, the smaller the difference between the two images, meaning they are more similar. Therefore, a smaller MSE means that the retained frames are more similar to the uncertain and defined frames, and the more likely the uncertain frame is to be a missed detection. Since the uncertain frame contains the target, it should be considered a defined frame. Regarding the inversely proportional normalized value of the MSE, this embodiment uses 1-norm(MSE) to present the inverse proportional relationship of MSE and the normalization process. norm() is a linear normalization function, and this will be used as an example for description.
[0102] In the j-th reserved box pair, calculate the sum of the target probability adjustment coefficients of 1 and the uncertain box, and multiply the sum by the target probability of the uncertain box as the updated target probability. When the updated target probability is greater than the preset probability threshold, the uncertain box is recorded as the determined box.
[0103] Using the method described above, determine whether the uncertain box in each pair of retained boxes is a definite box.
[0104] Therefore, several definite bounding boxes are selected from all uncertain bounding boxes, and combined with the original definite bounding boxes, all definite bounding boxes in the remote sensing image are obtained.
[0105] All bounding boxes in the remote sensing image are used as the target detection area in the remote sensing image.
[0106] The present invention also provides a remote sensing image target detection system based on deep learning, including a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor executes the computer program stored in the memory to implement the steps of the aforementioned remote sensing image target detection method based on deep learning.
[0107] This invention is now complete.
[0108] In summary, in this embodiment of the invention, determined bounding boxes and uncertain bounding boxes are obtained from remote sensing images. A bounding box pair is formed by any one determined bounding box and any one uncertain bounding box. Several retained bounding box pairs are selected from all bounding box pairs. Rotation and scaling operations are performed on the uncertain bounding boxes in each retained bounding box pair to obtain the final uncertain bounding box. Based on the difference between the final uncertain bounding box and the determined bounding box, a target probability adjustment coefficient for the uncertain bounding box is obtained. Combined with the target probability, several determined bounding boxes are selected from the uncertain bounding boxes in all retained bounding box pairs. All determined bounding boxes in the remote sensing image are used as the target detection region in the remote sensing image. This invention improves the accuracy of target detection in remote sensing images by comparing and analyzing determined bounding boxes and uncertain bounding boxes to further determine uncertain bounding boxes that may be missed.
[0109] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A remote sensing image target detection method based on deep learning, characterized in that: The method comprises the following steps: Acquire a remote sensing image, and use a target detection network to acquire a plurality of suggested frame selection areas in the remote sensing image and a target probability of each suggested frame selection area; and divide the suggested frame selection areas into confirmed frames and uncertain frames according to the target probability of each suggested frame selection area; A frame pair is formed by any determined frame and any uncertain frame, and several non-repeating frame pairs are obtained; according to the difference between the edge pixels in the determined frame and the uncertain frame in each frame pair, several reserved frame pairs are selected from all the frame pairs; Rotate and scale the uncertain box in each reserved box pair to obtain the final uncertain box of the uncertain box in each reserved box pair; According to the difference between the final uncertain box of the uncertain box in each reserved box pair and the edge pixel points in the determined box, the target probability adjustment coefficient of the uncertain box in each reserved box pair is determined; according to the target probability adjustment coefficient and the target probability of the uncertain box in each reserved box pair, several determined boxes are screened out from the uncertain boxes in all reserved box pairs; all the determined boxes in the remote sensing image are used as the target detection area in the remote sensing image.
2. According to claim 1, the method for remote sensing image target detection based on deep learning is characterized in that: The step of dividing the suggested box selection area into a determined box and an uncertain box according to the target probability of each suggested box selection area includes the following specific steps: The suggested box selection area with a target probability greater than the preset probability threshold is recorded as a confirmed box; The suggested box selection area whose target probability is less than or equal to the preset probability threshold is recorded as an uncertain box.
3. According to claim 1, the method for remote sensing image target detection based on deep learning is characterized in that: The specific steps of selecting a plurality of reserved frame pairs from all frame pairs according to the difference between the edge pixels in the determined frame and the uncertain frame in each frame pair are as follows: Use the Canny edge detection operator to perform edge detection on the i-th suggested box selection area and obtain several edge pixel points in the i-th suggested box selection area; According to the gray value of the pixel point in the i-th suggested box selection area, the LBP algorithm is used to obtain the LBP feature value of each edge pixel point in the i-th suggested box selection area; In the i-th suggested selection area, all edge pixels are clustered using a hierarchical clustering algorithm according to the LBP feature values of edge pixels to obtain a hierarchical clustering tree corresponding to the i-th suggested selection area; Use the APTED+ algorithm to obtain the edit distance between the hierarchical clustering tree corresponding to the determined box in each box pair and the hierarchical clustering tree corresponding to the uncertain box as the non-retention factor for each box pair; According to the non-retaining factor of each box pair, it is determined whether each box pair is a retained box pair.
4. According to claim 3, the method for remote sensing image target detection based on deep learning is characterized in that: The specific steps of judging whether each frame pair is a retained frame pair according to the non-retaining factor of each frame pair are as follows: The box pairs whose non-retention factor is less than the preset retention threshold are recorded as retained box pairs.
5. According to claim 3, the method for remote sensing image target detection based on deep learning is characterized in that: The step of rotating and scaling the uncertain box in each reserved box pair to obtain the final uncertain box of the uncertain box in each reserved box pair includes the following specific steps: According to the edge pixels in each suggested selection area, the PCA algorithm is used to obtain the main direction of the edge pixels as the main direction of each suggested selection area; Rotate the uncertain frame in any one of the reserved frame pairs so that the main direction of the uncertain frame in the any one of the reserved frame pairs is consistent with the main direction of the determined frame, and obtain a rotated uncertain frame as the rotated uncertain frame of the uncertain frame in the any one of the reserved frame pairs; The hierarchical clustering tree includes several layers, each layer includes several categories, and each category has several edge pixels; In all categories of each layer of the hierarchical clustering tree corresponding to each suggested frame selection area, the number of edge pixels of each category is arranged in ascending order to obtain an ascending sequence of numbers; In the ascending sequence of quantities, calculate the ratio of each quantity to the first quantity to obtain a category quantity ratio sequence; In each reserved frame pair, the category quantity ratio sequence of each layer of the hierarchical clustering tree corresponding to the determined frame is used as the left node, and the category quantity ratio sequence of each layer of the hierarchical clustering tree corresponding to the uncertain frame is used as the right node, and the KM matching algorithm is used to obtain a plurality of matching pairs and the edge value of each matching pair; each matching pair includes a left node and a right node; Among all the matching pairs of the jth retained box pair, the matching pair with the largest margin value is recorded as the scale matching pair; According to the ascending sequence of the number of left and right nodes in the scale matching pair, the scaling factor A of the jth reserved box pair is obtained. j ; Use the bilinear interpolation algorithm to scale the rotation uncertainty box of the uncertainty box in the jth retained box pair by A. j times, and obtain the uncertainty box after rotation and scaling as the final uncertainty box of the uncertainty box.
6. The method for remote sensing image target detection based on deep learning according to claim 5, characterized in that: According to the ascending sequence of the number of left and right nodes corresponding to the scale matching pair, the scaling factor A of the jth reserved frame pair is obtained. j The specific steps are as follows: The average of all the numbers in the ascending sequence of the numbers corresponding to the left nodes in the scale matching pair is recorded as the target scale N of the determined box in the jth retained box pair. j ; The mean of all the quantities in the ascending sequence of the quantities corresponding to the right nodes in the scale matching pair is recorded as the target scale M of the uncertain box in the jth reserved box pair. j ; The N j With the M j The ratio of is recorded as the scaling factor A of the jth reserved frame pair. j .
7. The method for remote sensing image target detection based on deep learning according to claim 1, characterized in that: The step of determining the target probability adjustment coefficient of the uncertain box in each reserved box pair according to the difference between the final uncertain box of the uncertain box in each reserved box pair and the edge pixel points in the determined box includes the following specific steps: In the jth reserved box pair, a segmentation determination box and a segmentation uncertainty box are obtained according to the length C1 and width K1 of the determination box and the length C2 and width K2 of the final uncertainty box of the uncertainty box; Set the grayscale value of the edge pixel points in the segmentation determination box to 1, and the grayscale value of the non-edge pixel points to 0, to obtain a target binary image; Set the grayscale value of the edge pixel points in the segmentation uncertainty box to 1, and the grayscale value of the non-edge pixel points to 0, to obtain a reference binary image; Calculating the product of the target binary image and the reference binary image to obtain a binary image with the same edge; Calculate the product of the binary image with the same edge and the segmentation determination box to obtain the target segmentation determination box; Calculate the product of the binary image with the same edge and the segmentation uncertainty box to obtain the target segmentation uncertainty box; The inversely proportional normalized value of the mean square error between the target segmentation certain frame and the target segmentation uncertain frame is recorded as the target probability adjustment coefficient of the uncertain frame in the jth retained frame pair.
8. The method for remote sensing image target detection based on deep learning according to claim 7, characterized in that: In the jth reserved frame pair, according to the length C1 and width K1 of the determined frame and the length C2 and width K2 of the final uncertain frame of the uncertain frame, the segmentation determined frame and the segmentation uncertain frame are obtained, and the specific steps include the following: In the jth reserved box pair, obtain the length C1 and width K1 of the determined box and the length C2 and width K2 of the final uncertain box of the uncertain box; In the determination box, with the center of the determination box as the center, a segmentation determination box with a length of min{C1, C2} and a width of min{K1, K2} is formed; the segmentation determination box is parallel to the wide side and the long side of the determination box; wherein min{} is a minimum value function; Within the final uncertain box of the uncertain box, a segmentation uncertain box with a length of min{C1, C2} and a width of min{K1, K2} is formed with the center of the final uncertain box as the center; the segmentation uncertain box is parallel to the wide side and the long side of the final uncertain box.
9. The method for remote sensing image target detection based on deep learning according to claim 1, characterized in that: The specific steps of selecting a plurality of determined frames from the uncertain frames in all the reserved frame pairs according to the target probability adjustment coefficient and the target probability of the uncertain frames in each reserved frame pair are as follows: In the jth reserved frame pair, the sum of 1 and the target probability adjustment coefficient of the uncertain frame is calculated, and the product of the sum and the target probability of the uncertain frame is recorded as the updated target probability; When the updated target probability is greater than a preset probability threshold, the uncertain box is recorded as a certain box.
10. A remote sensing image target detection system based on deep learning, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the computer program is executed by a processor, the steps of a method for realizing remote sensing image target detection based on deep learning as described in any one of claims 1 to 9 are implemented.