Method and device for detecting objects in any direction in remote sensing images by anchor box-independent corner point regression
By using the anchor frame-independent corner point regression method in the remote sensing image, the feature expression is enhanced by using semantic adjacency nodes, the problem of detecting objects in the remote sensing image is solved, and high-precision object detection is achieved.
Patent Information
- Application Number
- CN202210630115.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-04-08
- Filing Date
- 2022-06-06
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-06-06
AI Technical Summary
Objects in remote sensing images appear in any direction, and the prior art is difficult to effectively detect and locate, resulting in mis-detection and missed detection problems.
The corner point regression method is adopted for the anchor box independent corner point regression method, and multiple dense points are sampled on the candidate box boundary, and feature expression is enhanced by using semantic adjacency nodes, thereby regressing the corner points of the quadrilateral surrounding the box in any direction.
Accurate detection of objects in any direction in remote sensing images is realized, which reduces interference from background information and improves detection accuracy and efficiency.
Smart Images

Figure CN115240077B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of remote sensing information processing, and particularly relates to a method and device for detecting objects in any direction in remote sensing images with anchor box-free corner regression. Background Technique
[0002] With the development of satellite technology, observing the ground using artificial satellites can obtain data within a very large spatial range in a very short time. Therefore, remote sensing technology has been increasingly valued by countries. According to the types of sensors carried on artificial satellites, the types of remote sensing images obtained are also different. Currently, they are mainly divided into visible light remote sensing images, infrared remote sensing images, synthetic aperture radar images, and multispectral remote sensing images. Since visible light remote sensing images are easy to obtain and the target forms are easy to identify by the human eye, remote sensing images based on visible light are widely used. Detecting target objects of interest (such as airplanes, ships, vehicles, bridges, seaports, storage warehouses, etc.) from such remote sensing images has very important application value.
[0003] Nowadays, optical remote sensing images can reach nanometer-level resolution, which can provide very fine texture and spatial information, thus facilitating the detection of independent target individuals. At the same time, high-quality remote sensing images also highlight irrelevant targets and background region textures, making the detection of target objects of interest face huge challenges. For the region of interest, the scales of its targets vary greatly, and there are many small-target objects; in addition, due to the shooting angle factor, the target regions show layouts in any direction. These factors all increase the difficulty of detecting target objects in optical remote sensing images and are prone to false detection and missed detection of the region of interest.
[0004] In recent years, for natural scene images, many deep learning-based object detection methods have achieved exciting performance. However, the outputs of these methods are usually horizontal bounding boxes. However, in remote sensing images, due to the continuous rotation of the earth, the objects in the images taken by satellites will show layouts in different directions. Using horizontal bounding boxes to locate objects in any direction in remote sensing images will result in many background information being included in the located boxes, which is not conducive to downstream decision-making tasks. Therefore, the detection of objects in any direction in remote sensing images is particularly important and is also a very challenging task. Currently, the methods for detecting objects in any direction in remote sensing images can be roughly divided into three categories: segmentation-based methods, angle-based methods, and key-point-based methods.
[0005] Segmentation-based methods, such as CenterMap-Net, predict the probability map of the center of the object bounding box in the candidate box. However, such methods rely on the localization accuracy of the candidate box. And dense pixel-level segmentation requires more storage space. Angle-based methods predict a rotation angle based on predicting the center point of the horizontal bounding box of the object and the scale of the horizontal bounding box. Due to the arbitrary orientation of the object, it is difficult to effectively capture the feature expression of the object using standard convolutional neural networks. Therefore, some methods (such as ROI-Trans, DRN, R3Det, etc.) are dedicated to solving the problem of misalignment of rotated object features. In addition, some methods (such as SCRNet, PIoU, RSDet, etc.) are dedicated to designing loss functions to promote the learning of angles. Additionally, CSL and DCL adopt a cyclic smoothing label technique to handle the periodicity problem of angles and convert angle regression into a classification problem. Key point regression-based methods can be further divided into corner-based methods and midpoint-based methods. Corner-based methods directly predict the four vertices of the directional bounding box; while midpoint-based methods reconstruct the bounding box of an object in an arbitrary direction by predicting the midpoint of the directional bounding box. For example, IENet directly predicts the distances from the foreground pixels of the object to the four sides. Gliding-Vertex and TOSO predict the sliding distances from the four corners of the horizontal candidate box to the four corners of the rotated rectangle box. RIL-Q designs a loss function with expression invariance to optimize the regression of corners. In addition, O2-DNet and BBAVectors not only predict the distances from the center point of the object to the midpoints of the four sides, but also estimate the width and height of the rotated rectangle box.
[0006] Therefore, the present invention generates a small number of candidate boxes through a network without prior box design, then samples multiple dense points on the boundaries of the candidate boxes, and enhances the key feature expression through the sampled points of semantic neighbors of each point and the geometric topology of the boundaries, so as to regress the corners of the accurate quadrilateral bounding box in an arbitrary direction. Summary of the Invention
[0007] The present invention proposes a method and device for detecting objects in an arbitrary direction in remote sensing images with corner regression independent of anchor boxes. This method directly regresses the key points on the boundary of the candidate box to the corners of the quadrilateral bounding box of the object on the basis of the candidate box. In the process of generating candidate boxes, the present invention estimates the center point of the horizontal bounding box of the object and the height and width of the horizontal bounding box to generate a horizontal bounding box in a manner independent of anchor boxes. In the process of corner regression, the present invention uses semantic adjacent nodes to enhance the feature expression of the boundary sampling points, so that the key points on the boundary regress to more accurate corner positions.
[0008] To achieve the above object, the technical solution of the present invention includes:
[0009] A method for detecting objects in any direction in an image based on anchor-free corner regression, the steps of which include:
[0010] Extract the global feature representation of the input image;
[0011] Based on the global feature representation, predict the center point, height and width of the horizontal bounding box of the object to reconstruct the horizontal candidate box of the object;
[0012] Based on the global feature representation, extract the original feature representation of the boundary sampling points of the horizontal candidate box of the object;
[0013] Utilize the semantic adjacent nodes of the boundary sampling points to enhance the original feature representation;
[0014] Obtain the boundary key points of the horizontal candidate box of the object, and extract the feature representation of the boundary key points according to the enhanced feature representation to estimate the corner offset between the boundary key points and the corners of the bounding box of the object in any direction;
[0015] Based on the corner offset and the boundary key points, calculate the corner coordinates of the bounding box of the object in any direction;
[0016] Based on the constructed bounding box of the object in any direction, detect the objects in the input image.
[0017] Furthermore, the extraction of the global feature representation of the input image
[0018] Use the backbone network to extract the global visual feature representation F 0 of the input image, the global visual feature representation F 1 of the input image, the global visual feature representation F 2 of the input image and the global visual feature representation F 3 of the input image, where the backbone network includes: DLA34 network;
[0019] Fuse the global visual feature representation F 0 of the input image, the global visual feature representation F 1 of the input image, the global visual feature representation F 2 of the input image and the global visual feature representation F 3 of the input image to obtain the global feature representation of the input image.
[0020] Furthermore, the prediction of the center point, height and width of the horizontal bounding box of the object based on the global feature representation to reconstruct the horizontal candidate box of the object includes:
[0021] Based on the global feature representation, generate the center point response map of the horizontal box and the scale prediction map of the horizontal box
[0022] Response map of the center point of the object's horizontal frame Perform a max pooling operation and filter according to a confidence threshold to obtain the filtered center point;
[0023] According to the filtered center point and the scale prediction map of the object's circumscribed rectangle Generate horizontal candidate boxes for the object.
[0024] Furthermore, based on the global feature representation, extract the original feature representation of the boundary sampling points of the object's horizontal candidate box
[0025] Uniformly sample on the boundary of the horizontal text candidate box to obtain N c boundary sampling points and their corresponding sampling point coordinates P;
[0026] Extract the original feature representation of the boundary sampling points according to the global feature representation wherein, represents the feature mapping operation, represents the global feature representation, P j represents the position coordinate of the j-th boundary sampling point, Δx j and Δy j respectively represent the horizontal offset and vertical offset of the j-th sampling point from a corner point of the horizontal bounding box.
[0027] Furthermore, enhancing the original feature representation by using the semantic adjacent nodes of the boundary sampling points includes:
[0028] Use a 1D convolutional network to map the original feature representation to feature X, and use 1D convolutions with different dilation rates to map feature X into different feature representations U;
[0029] For each feature representation U, based on the feature similarity of the boundary sampling points, select K semantic adjacent points for each boundary sampling point, and perform feature expression aggregation according to the weights between the boundary sampling points and the K semantic adjacent points to obtain the boundary sampling point feature expression under this feature representation U;
[0030] For each boundary sampling point, concatenate the boundary sampling point feature expressions generated under different dilation rates to obtain the enhanced feature expression.
[0031] Furthermore, obtaining the boundary key points of the object's horizontal candidate box includes:
[0032] Set the number N of boundary key points k ;
[0033] Select from the boundary sampling points at a rate of N c / N kSample at intervals to obtain the boundary key points, where N c represents the number of boundary sampling points.
[0034] Furthermore, the feature expression of the boundary key points where i represents the serial number of the object in the input image, m ∈ [1, N k , F refined represents the enhanced feature.
[0035] A convolutional network for implementing any of the above methods, and the loss function for training the convolutional network where N k represents the number of boundary key points, represents the predicted value of the j-th coordinate of the m-th sampling key point of the quadrilateral bounding box of the object in any direction, T ij represents the true value of the j-th coordinate of the m-th sampling key point of the quadrilateral bounding box of the object in any direction, represents the smooth-L1 loss function
[0036] A storage medium stores a computer program, where the computer program is configured to execute any of the above methods when running.
[0037] An electronic device includes a memory and a processor. The memory stores a computer program, and the processor is configured to run the computer program to execute any of the above methods.
[0038] Advantages of the present invention:
[0039] 1. The present invention can regress from the key points in the candidate box to the corner points of the bounding box of the object in any direction, and a more compact quadrilateral bounding box can be formed.
[0040] 2. The present invention adopts a dynamic boundary information aggregation technology, which can enhance the feature expression of the key points and make the positioning of the corner points more accurate.
[0041] 3. The present invention does not need to design prior anchor boxes, so that the model has better generalization for objects of different scales.
[0042] 4. The number of horizontal candidate boxes generated by the present invention is significantly less than the number of candidate boxes generated by regression based on prior boxes; in addition, the structure of the entire network is relatively simple, so that the execution speed of the model can be effectively improved.
[0043] 5. The present invention has strong detection ability and excellent detection performance for objects in different directions, different scales and different types in remote sensing images. Description of the Drawings
[0044] Figure 1 Flowchart of object detection in any direction for remote sensing image anchor box-independent corner regression. Specific implementation manner
[0045] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present invention.
[0046] The method for object detection in any direction of remote sensing images of the present invention, as Figure 1 shown, includes:
[0047] Step 1: Extract the feature representation of the input image.
[0048] In an embodiment of the present invention, the input is an RGB image with a size of H*W, where H is the height of the RGB image and W is the width of the RGB image.
[0049] The feature extraction of the image input into the network is specifically as follows:
[0050] 1.1) Use the backbone network (i.e., the network pre-trained on ImageNet, such as DLA34, etc.) to extract the visual features of the input image, and its output is expressed as For DLA34, the feature dimensions D0, D1, D2, and D3 are 64, 128, 256, and 512 respectively, and F 0 、F 1 、F 2 、F 3 are multi-level global visual feature representations.
[0051] 1.2) Fuse the features F 0 , F 1 , F 2 and F 3 to obtain the fused feature The specific fusion method is expressed as:
[0052]
[0053] where represents the fused feature of the l-th layer, represents the upsampling operation with a factor of 2 ι , and represent the feature dimensionality reduction operation, and Θ1 and Θ2 respectively represent and the learnable parameters in. α represents the image downsampling factor, which is set to 4 in one example.
[0054] Step 2: Predict the center point of the horizontal bounding box of the object and the height and width of the horizontal bounding box to reconstruct the horizontal candidate box of the object.
[0055] In one embodiment of the present invention, based on the fused feature Generate the response map of the center point of the horizontal box of the object The network Consists of a 3×3 convolutional layer (256 convolutional kernels), a rectified linear unit (ReLU), and a 1×1 convolutional layer (1 convolutional kernel).
[0056] In one embodiment of the present invention, based on the fused feature Generate the scale prediction map of the horizontal box of the object The network Consists of a 3×3 convolutional layer (256 convolutional kernels), a rectified linear unit (ReLU), and a 1×1 convolutional layer (2 convolutional kernels).
[0057] During training, the loss function of the network contains the loss of the center point of the horizontal box of the object And the scale loss These two parts are respectively expressed as:
[0058]
[0059]
[0060] Where N represents the number of objects in the image, i represents the position index on the prediction map And Are the corresponding ground truths, And Respectively represent the corresponding ground truths, Is the Smooth-L1 loss function, r and β respectively represent the first penalty factor and the second penalty factor.
[0061] During the testing process, according to the obtained response map of the center point of the horizontal box of the object The present invention first uses a 3×3 max pooling operation to highlight the center point; then filters the center points with low scores through a threshold τ c The reconstructed horizontal candidate box of the object can be expressed as:
[0062]
[0063] Where And Are the horizontal and vertical coordinates corresponding to the point with the largest i-th response. and respectively represent the width and height of the horizontal box prediction corresponding to the i-th center point.
[0064] Step 3: Map the boundary sampling points of the horizontal candidate box to the image feature map to obtain the feature expression of the boundary sampling points.
[0065] In an embodiment of the present invention, in order to obtain the feature expression of the sampling points on the boundary of the object's horizontal bounding box, the present invention first uniformly samples N points on the boundary of each horizontal candidate box c , and then extracts its semantic information and spatial layout information according to the sampling point positions, and extracts the original feature expression of the boundary sampling points from the fused feature . The specific implementation method is as follows: Specifically,
[0066]
[0067] where j represents the sampling point index, P j represents the position coordinates of the j-th sampling point, represents the feature mapping operation, Δx j and Δy j represent the horizontal and vertical offsets of the j-th sampling point from the upper left corner point of the horizontal bounding box. [;] represents the feature concatenation operation.
[0068] Step 4: Use the dynamic information aggregation technology to enhance the feature expression of the boundary sampling points by using the semantic adjacent nodes of the boundary points.
[0069] In an example of the present invention, in order to enhance the feature expression of the horizontal candidate box boundary, the present invention uses the dynamic information aggregation technology to enhance the feature expression of the boundary sampling points by using the semantic adjacent nodes of the boundary points (D r is the feature dimension), and the specific implementation process is as follows:
[0070] 4.1) Use 1D convolution to map F init into the feature Then use two 1D convolutions with a dilation rate of r to map X into the feature expressions and where D g represents the feature dimension of the sampling points.
[0071] 4.2) Use the nearest neighbor algorithm to select K semantic adjacent points for each sampling point according to the similarity of the feature U of the sampling points, and the corresponding feature expressions of these adjacent points are represented as
[0072] 4.3) Calculate the weight W between each node and the K adjacent points, and the calculation method is:
[0073]
[0074] where j is the index of the boundary sampling point, and k is the index of the semantic nearest neighbor point of the j-th sampling point. U j represents the feature expression of the j-th sampling point. Z j represents the normalization factor.
[0075] 4.4) Aggregate the feature expressions of the K adjacent points of each boundary sampling point, and the aggregation method is:
[0076]
[0077] where d is the feature dimension index.
[0078] 4.5) Concatenate the boundary feature expressions F r generated at different dilation rates to form F refined .
[0079] Step 5: Extract the feature expressions of the boundary key points based on the enhanced feature expressions of the boundary sampling points of the horizontal candidate boxes, so as to estimate the offsets of the key points from the corner points of the quadrilateral bounding box of the object in any direction.
[0080] In an example of the present invention, according to the feature expression F refined of the boundary sampling points of the horizontal candidate box, extract the feature expressions of the boundary key points which is expressed as:
[0081]
[0082] where N k represents the number of corner points of the quadrilateral bounding box. In the present invention, N k = 4. In formula (8), i represents the index of the bounding box number. m * N c / N k represents the position of the boundary key point, that is, sampling from N c boundary sampling points at an interval of N c / N k . The coordinates of the N k key points are represented as a matrix
[0083] Step 6: Obtain the coordinates of the corner points of the quadrilateral bounding box of the object in any direction according to the position coordinates of the boundary key points of the horizontal box and the offsets predicted by the network.
[0084] In an example of the present invention, use F cp to generate the corner point position offset O, and then combine it with the position of the boundary key point to obtain the new predicted position of the corner point
[0085] Network training:
[0086] During the training process, the loss function is
[0087]
[0088] where represents the predicted value of the j-th coordinate of the m-th sampled key point of the quadrilateral bounding box of an object in any direction. T ij is the corresponding ground truth.
[0089] Experimental data
[0090] For the method for detecting objects in any direction in remote sensing images with anchor box-free corner regression of the present invention, its test environment and experimental results are as follows:
[0091] (1) Test environment:
[0092] System environment: ubuntu16.04.
[0093] Hardware environment: Memory: 15GB, GPU: NVIDIA RTX 2080Ti, CPU: 4.00GHz Intel(R)Xeon(R)W-2125, Hard disk: 2TB.
[0094] (2) Experimental data:
[0095] The present invention conducts experiments on two remote sensing image datasets, namely HRSC2016 (617 training images, 444 test images) and DOTA (1869 training images, 937 test images). Due to the extremely large resolution of the images in the DOTA dataset, during the training and testing processes, for each image, we crop it with a sliding window of size 800*800 and a step size of 200. For the dataset HRSC2016, the images are scaled to 640*640 during the training and testing processes.
[0096] (3) Optimization method:
[0097] The Adam optimizer is used for optimization. For both the HRSC2016 and DOTA models, 300 epochs are trained. The initial learning rate of the model is 0.0001. Its learning rate is multiplied by 0.1 after the 120th and 240th epochs. During the training process, the images are randomly rotated {0°, 90°, 180°, 270°}, and randomly flipped, scaled, and color jittered. For HRSC2016 and DOTA, due to the limitation of video memory, the training batch sizes are set to 10 and 8 respectively.
[0098] (4) Experimental results:
[0099] 1) Ablation experiment:
[0100] This experiment was completed on the HRSC2016 dataset, and the experimental results are shown in Table 1. In the experiment, three baseline models were set up. The first one is to directly predict the center and scale of the horizontal circumscribed rectangle of the object to obtain the horizontal bounding box (HBB). The second baseline model is to obtain the rotated rectangle bounding box by additionally predicting the angle. The third baseline model is to predict the offsets of the center point from the four corner points of the object's quadrangle bounding box to reconstruct the quadrangle bounding box. When the dynamic information aggregation technology (DIG) is not adopted in the present invention to enhance the feature expression of the boundary sampling points, the F-measure and mAP of our model under the two evaluation criteria have significantly exceeded the baseline models. Especially when DIG is adopted, the detection performance of the present invention reaches the best.
[0101] Table 1: Verification of the effectiveness of the proposed module
[0102]
[0103]
[0104] 2) Performance comparison:
[0105] As can be seen from Table 2 and Table 3, the method of the present invention achieves the state-of-the-art performance.
[0106] Table 2: Performance comparison on HRSC2016
[0107]
[0108] Table 3: Performance comparison on DOTA
[0109]
[0110]
[0111] The above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit them. Those of ordinary skill in the art can modify or equivalently replace the technical solutions of the present invention without departing from the scope of the present invention. The protection scope of the present invention shall be subject to what is described in the claims.
Claims
1. An object detection method for arbitrary - direction objects in images based on anchor - free corner regression, the steps of which include: Extract the global feature representation of the input image; Based on the global feature representation, predict the center point, height, and width of the object's horizontal bounding box to reconstruct the object's horizontal candidate bounding box; Based on the global feature representation, extract the original feature representation of the boundary sampling points of the object's horizontal candidate bounding box; wherein, the extracting the original feature representation of the boundary sampling points of the object's horizontal candidate bounding box based on the global feature representation includes: Uniformly sample on the boundary of the horizontal text candidate box to obtain N c boundary sampling points and the corresponding sampling point coordinates P; Extract the original feature representation of the boundary sampling points according to the global feature representation Among them, represents the feature mapping operation, represents the global feature representation, P j represents the position coordinates of the j-th boundary sampling point, Δx j and Δy j respectively represent the horizontal offset and vertical offset between the j-th sampling point and a corner point of the horizontal bounding box; Enhance the original feature representation by using the semantic adjacent nodes of the boundary sampling points; Obtain the boundary key points of the object's horizontal candidate bounding box, and extract the feature representation of the boundary key points according to the enhanced feature representation to estimate the corner offset between the boundary key points and the corners of the arbitrary - direction object bounding box; wherein, the obtaining the boundary key points of the object's horizontal candidate bounding box includes: Set the number N of boundary key points k ; Sample at an interval of N c / N k from the boundary sampling points to obtain the boundary key points, where N c represents the number of boundary sampling points, and the feature expression of the boundary key points where i represents the serial number of the object in the input image, m ∈ [1, N k , and F refined represents the enhanced feature; Calculate the corner coordinates of the arbitrary - direction object bounding box based on the corner offset and the boundary key points; Based on the constructed arbitrary - direction object bounding box, detect the objects in the input image.
2. The method according to claim 1, characterized in that, The extracting the global feature representation of the input image Extract the global visual feature representation F of the input image using the backbone network 0 、the global visual feature representation F 1 、the global visual feature representation F 2 and the global visual feature representation F 3 , where the backbone network includes: DLA34 network; Fuse the global visual feature representation F 0 and the global visual feature representation F 1 and the global visual feature representation F 2 with the global visual feature representation F 3 to obtain the global feature representation of the input image.
3. The method according to claim 1, characterized in that, The predicting the center point, height, and width of the object's horizontal bounding box based on the global feature representation to reconstruct the object's horizontal candidate bounding box includes: Based on the global feature representation, generate an object horizontal box center point response map respectively and a horizontal box scale prediction map For the object horizontal box center point response map perform a max pooling operation and filter according to a confidence threshold to obtain the filtered center points; Generate horizontal candidate bounding boxes for objects according to the filtered center points and the scales of the minimum bounding rectangles of the objects Generate horizontal candidate bounding boxes for objects.
4. The method according to claim 1, characterized in that The enhancing the original feature representation by using the semantic adjacent nodes of the boundary sampling points includes: Use a 1D convolutional network to map the original feature representation to feature X, and use 1D convolutions with different dilation rates to map feature X into different feature representations U; For each feature representation U, based on the feature similarity of the boundary sampling points, select K semantic adjacent points for each boundary sampling point, and perform feature - expression aggregation according to the weights between the boundary sampling points and the K semantic adjacent points to obtain the boundary sampling point feature representation under this feature representation U; For each boundary sampling point, concatenate the boundary sampling point feature representations generated under different dilation rates to obtain the enhanced feature representation.
5. A storage medium storing a computer program therein, wherein, The computer program is set to execute any one of the methods recited in claims 1 - 4 when running.
6. An electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is set to run the computer program to execute any one of the methods recited in claims 1 - 4.
Citation Information
Patent Citations
Random-shape scene character detection method and device of asymptotic regression boundary
CN113139539A
Rotation equivariant space local attention remote sensing image target detection method
CN113850129A