Remote sensing image target detection method based on feature pyramid and boundary perception vector

Through the combination of feature pyramids and boundary perception vectors, the problems of low detection accuracy and high calculation cost in the rotation object detection of remote sensing images are solved, and the detection effect of higher accuracy, lower cost and faster speed is achieved. It is suitable for remote sensing image object detection of multi-scale and complex backgrounds.

CN120451516AActive Publication Date: 2025-08-08NAT INNOVATION INST OF DEFENSE TECH PLA ACAD OF MILITARY SCI
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510942006.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-09
Publication Date
2025-08-08
Estimated Expiration
2045-07-09

AI Technical Summary

Technical Problem

When using multi-scale and complex backgrounds, existing remote sensing image rotation object detection methods are difficult to take into account both detection accuracy and calculation efficiency. Traditional horizontal bounding boxes are difficult to accurately describe the true shape of the rotation object, resulting in low detection accuracy and high calculation cost.

Method used

Using a combination of feature pyramids and boundary-aware vectors, multi-scale features are extracted through feature pyramid networks, a dynamic selection mechanism is used to select appropriate kernels for different targets, and a rotation boundary box prediction is combined with boundary-aware vectors to avoid complex joint learning in traditional methods.

Benefits of technology

It significantly improves detection accuracy, reduces computing costs, enhances multi-scale adaptability and real-timeness, and can efficiently perform remote sensing image rotation object detection in resource-constrained environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120451516A_ABST
    Figure CN120451516A_ABST
Patent Text Reader

Abstract

The invention discloses a remote sensing image target detection method based on a feature pyramid and a boundary perception vector. The method comprises the following steps: S1, feeding an input image into a feature pyramid backbone to extract multi-scale features; s2, adopting a dynamic selection mechanism, and adaptively selecting proper kernels for different targets according to the extracted multi-scale features; and S3, inputting the comprehensive features into a decoder to obtain a prediction coordinate and a target category. According to the method, the feature pyramid and the learning frame boundary perception vector are combined to capture the rotation boundary frame of the target, the calculation amount is remarkably reduced, higher precision is obtained, and the application requirement for remote sensing image rotation target detection under the condition that resources are limited in the aerospace field is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of remote sensing image processing, and in particular to a method for detecting rotating targets in optical remote sensing images based on feature pyramids and boundary perception vectors. Background Art

[0002] Rotating object detection in optical remote sensing images is a technique for detecting rotating objects in remote sensing images. Remote sensing images, acquired through platforms such as satellites and drones, capture the Earth's surface. They are characterized by wide coverage, high resolution, and rich geographic information. Since objects in remote sensing images (such as aircraft, ships, and vehicles) can appear at arbitrary angles, they have important applications in a variety of fields, including military reconnaissance, intelligent transportation, environmental monitoring, and land use planning.

[0003] With the increasing application of convolutional neural networks (CNNs) in computer vision, object detection in aerial imagery has rapidly developed. Related models typically use horizontal bounding boxes (HBBs) to locate objects. Most objects in remote sensing images are densely distributed, subject to occlusion, and have high aspect ratios. Consequently, HBB-based models can result in significant misalignment and overlap between the detected bounding boxes and the objects, with the size and aspect ratio failing to reflect the true shape of the objects. To address this issue, rotated bounding boxes are used to process these objects. These boxes contain fewer background pixels, making it easier to classify the object from the background, improving object capture accuracy, and significantly reducing the overlap between adjacent objects compared to horizontal bounding boxes. Therefore, rotated bounding boxes are an effective solution for capturing objects in aerial imagery.

[0004] Rotated object detection in aerial imagery presents unique challenges due to the bird's-eye view, complex backgrounds, and varying object appearance. This often results in models predicting high confidence in object classification but inaccurate localization. Remote sensing object detection methods often employ detectors based on two-stage anchors. For example, on the DOTA dataset, the two-stage anchor-based Faster R-CNN algorithm achieved a mean average precision (mAP) of 0.85 when detecting aircraft targets, while the single-stage anchor-based YOLOv3 algorithm achieved a mAP of only 0.72. This is primarily because the two-stage anchor approach generates high-quality region proposals through the RPN in the first stage, which more accurately localizes the approximate location of the target and provides a solid foundation for subsequent fine-grained detection. In the second stage, ROI feature merging and refinement further improve detection accuracy. Using center points instead of anchor boxes, the spatial information and semantic information of key regions are used to adaptively localize the target, resulting in improved detection accuracy. In the first stage, these detectors typically densely spread anchor boxes on the feature map and then regress the offset between the target box and the anchor box parameters to produce region proposals. In the second stage, the region of interest (ROI) features are combined to refine the box parameters and classify the object category. They use the center, width, height and angle as the description of the rotated bounding box.

[0005] In recent years, keypoint-based object detectors have been developed. Their computational efficiency and fast processing capabilities address the slow inference speed of anchor-based methods in horizontal object detection. These methods detect bounding box corners and group them by comparing their embedding distance or center distance. While these strategies have shown improved performance, they are computationally expensive due to the grouping process. To address this issue, CenterNet detects object centers and directly regresses the width (w) and height (h) of the bounding box, achieving faster detection with comparable accuracy. Intuitively, by learning an additional angle θ along with w and h, CenterNet can be extended to rotated object detection. However, since w and h are measured in a different rotated coordinate system for each arbitrarily rotated object, jointly learning these parameters poses a significant challenge for the model.

[0006] Traditional horizontal object detection methods often struggle to accurately identify and localize these rotated objects, necessitating specialized rotated object detection techniques to address this issue. Detection of rotated objects in remote sensing images involves densely packed objects, partial occlusion, and high aspect ratios. Traditional horizontal bounding box detection methods struggle to accurately describe the true shape of these objects, resulting in large alignment errors between the detection box and the object and excessive background pixels, which reduces object detection accuracy. Compared to horizontal bounding boxes, oriented bounding boxes (OBBs) can better adapt to the object's rotation angle, reduce background pixel interference, and improve object localization accuracy. However, rotated object detection faces the following technical challenges: the parameters of the rotated bounding box (such as width, height, and angle) are typically defined in different rotational coordinate systems, making it challenging for the model to jointly learn these parameters. Existing detection methods often struggle to balance detection accuracy and computational efficiency when dealing with objects of multiple scales and complex backgrounds, failing to meet application requirements. Summary of the Invention

[0007] In response to the problems of low experimental accuracy, slow inference speed, large model parameters and computational complexity in the existing technology, the purpose of the present invention is to provide a method for detecting rotated targets in optical remote sensing images based on feature pyramid and boundary perception vector. The feature pyramid is combined with the learned box boundary perception vector to capture the rotated bounding box of the target, which significantly reduces the computational complexity and achieves higher accuracy.

[0008] To achieve the above object, the present invention provides a method for remote sensing image target detection based on feature pyramid and boundary perception vector, the method comprising the following steps: S1. Feed the input image into the feature pyramid backbone to extract multi-scale features; S2. A dynamic selection mechanism is used to adaptively select appropriate kernels for different targets based on the extracted multi-scale features; S3. Input the comprehensive features into the decoder to obtain the predicted coordinates and target category, thereby obtaining the rotated bounding box of the target.

[0009] Furthermore, the backbone network of the method includes a backbone network, a neck module, an LSK module and a detection head (the detection head here is the decoder or prediction head of the four vectors below); the backbone network is implemented based on the ResNet architecture, and the ResNet101 backbone processes the RGB input image; the backbone network extracts hierarchical features by performing a series of convolution, batch normalization and ReLU activation layers on the input image.

[0010] Further, step S1 includes: S1-1. First, through the backbone network, four sets of feature maps of different scales are obtained through layer-by-layer convolution; S1-2. To address the issue of different feature maps, we use a feature pyramid structure to construct bottom-up and top-down feature fusion paths, fusing feature maps of different scales to generate a feature pyramid with multi-scale information. S1-3. Splice the feature pyramids after FPN to obtain a feature map of 152×152×256 shape, where 152×152 is the image resolution and 256 is the number of feature maps.

[0011] Further, step S2 includes: S2-1. The input feature map passes through the large kernel selection module, and its receptive field is expanded through 5×5 and 7×7 convolutions to obtain information around the target. S2-2. Using a spatial selection mechanism, these features are weighted and fused according to the input features to generate an expanded receptive field of multiple features.

[0012] Furthermore, on top of the backbone network, a feature pyramid network (FPN) is used as the neck module. The convolutional layers refine the upsampled feature maps.

[0013] Furthermore, the core of the feature pyramid network is to effectively integrate multi-scale features. Through feature pyramids of different scales, the network can capture fine-grained and coarse-grained information; during the upsampling process, jump connections are used to combine deep and shallow features.

[0014] Furthermore, the feature pyramid network first upsamples the deep feature map to the size of the shallow feature map by bilinear interpolation, and then uses The convolutional layer refines the upsampled map; after connecting with the shallow feature map, The convolutional layers refine the channel features, and batch normalization and ReLU activation are applied in the latent layers.

[0015] Furthermore, the LSK module processes the input feature map and enhances the representation capability through a series of convolution and attention mechanisms.

[0016] Furthermore, the LSK module converts the feature map into four branches: heat map, offset, box parameter and direction map.

[0017] Furthermore, when the LSK module transforms the feature map, it uses two Convolutional layer implementation with kernel and 256 channels.

[0018] Furthermore, the LSK module consists of two sub-blocks: a large kernel selection sub-block and a feed-forward network sub-block; in the large kernel selection sub-block of the LSK module, a large kernel convolution is first constructed by decomposing the large kernel convolution into a series of depthwise convolutions with increasing kernel size and expansion rate; this process produces multiple features with different receptive fields; then a spatial selection mechanism is adopted to weight and fuse these features according to the input features.

[0019] The beneficial effects of the present invention are as follows: This paper uses BBAVectors as a baseline method, focusing on regressing four vectors in a Cartesian coordinate system. Comparing this method with BBAVectors demonstrates the advantages of leveraging feature pyramids and boundary vectors. Experimental evaluation on the DOTA and HRSC2016 datasets demonstrates the superiority of this method over existing techniques. This method has the following four advantages: a) Higher detection accuracy. This invention achieves more accurate detection and identification of rotated objects by integrating a novel architecture called a feature pyramid network and boundary-aware vectors. By predicting the vector information of the object's bounding box, the boundary-aware vector avoids the complex joint learning of width, height, and angle required in traditional methods, significantly improving detection accuracy.

[0020] b) Lower computational cost. This paper adopts a single-stage, anchor-free detection method, avoiding the traditional two-stage detection method, such as the complex region proposal generation and refined detection steps in Faster R-CNN, significantly reducing computational complexity.

[0021] c) Enhanced multi-scale adaptability: The Feature Pyramid Network efficiently integrates multi-scale features, enabling the model to excel in processing small objects, large objects, and objects with large aspect ratios. This multi-scale feature extraction capability makes the present invention more adaptable in complex scenarios. The present invention also uses a dynamic selection mechanism to adaptively select appropriate feature kernels based on the object's scale and orientation, further enhancing its ability to detect multi-scale objects.

[0022] d) Improved real-time performance. This invention prioritizes computational efficiency and real-time performance in its design. By optimizing the network architecture and reducing the number of model parameters, it significantly improves detection speed. Experiments show that this invention significantly outperforms existing technologies in detection speed on standard datasets (DOTA and HRSC2016), meeting the requirements of real-time application scenarios.

[0023] In general, the method of the present invention can strike a good balance between detection accuracy and speed, and meet the application requirements of the aerospace field for rotating target detection in remote sensing images under resource-constrained conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 1 is a flow chart of a method for detecting target in remote sensing images based on feature pyramid and boundary perception vector according to the present invention; Figure 2 This is a schematic diagram of the detection results of FLBA-R on the DOTA V1.0 dataset; Figure 3 This is a schematic diagram of the detection results of FLBA-R on the HRSC 2016 dataset; Figure 4 Schematic diagram of comparison of detection results, including (a) missed detection comparison diagram; (b) error detection comparison diagram; (c) redundant detection comparison diagram; (d) angle correction comparison diagram; (e) confidence increase comparison diagram; Figure 5 This is a detailed diagram of the detection results under the DOTA dataset; Figure 6 Figure 1 is a schematic diagram comparing the visualization of feature maps before and after feature enhancement, where (a) shows the detection results on the DOTA V1.0 test dataset; (b) and (c) show the visualization of feature maps before and after LSK module processing in the proposed method, further verifying its effectiveness. Figure 7 It is a schematic diagram of the prediction results. DETAILED DESCRIPTION

[0025] The following will clearly and completely describe the technical solution of the present invention in conjunction with the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0026] In the description of the present invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings and are intended solely to facilitate and simplify the description of the present invention. They are not intended to indicate or imply that the devices or components referred to must have, be constructed, or operate in a specific orientation, and therefore should not be construed as limitations on the present invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0027] In the description of the present invention, it should be noted that, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood in a broad sense. For example, they may refer to fixed, detachable, or integral connections; mechanical or electrical connections; direct or indirect connections through an intermediate medium; and internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on specific circumstances.

[0028] The following combination Figure 1-Figure 7 The specific embodiments of the present invention are described in detail. It should be understood that the specific embodiments described herein are only used to illustrate and explain the present invention and are not used to limit the present invention.

[0029] This method first implements rotated object detection in optical remote sensing images using a feature pyramid structure. A feature integration module then integrates the features extracted from the feature pyramid to better separate the target from background information. Finally, boundary-aware vectors are used for post-processing to obtain the final target. Compared to the original boundary-aware vector (BBAVector) method, this method offers higher detection accuracy, lower computational cost, and faster detection speed. By optimizing the network architecture and reducing the number of model parameters, this method improves scalability and adaptability, enabling efficient operation in resource-constrained environments.

[0030] The difference between the existing model and the present invention lies in the different model architectures. The existing model, i.e., the baseline model, uses a U-Net network for feature extraction. The U-Net structure first downsamples the image and then upsamples it. The upsampling part has certain parameter redundancy and poor effect. Therefore, the feature pyramid is used to directly fuse the downsampled features of different stages, which can improve the feature extraction effect while reducing parameters. The LSK module is then used to improve the overall effect of the network while increasing the limited number of parameters and calculations. Finally, the boundary perception vector is used to process the extracted features to obtain the final rotation box result. The present invention removes the upsampling of the baseline model, uses the feature pyramid to obtain multi-scale features, and uses the feature selection module LSK to further optimize the features. Finally, the boundary perception module of the baseline model is used to obtain better results.

[0031] The use of spatial pyramids in the task of rotating target detection in remote sensing images can better capture multi-scale targets, and predicting the vectors of the four quadrants of the Cartesian coordinate system is better than directly predicting the spatial parameters of the bounding box. Figure 1 As shown, the remote sensing image target detection method based on feature pyramid and boundary perception vector of the present invention includes the following steps: S1. Initially, the input image is fed into a feature pyramid backbone to extract multi-scale features. Multi-scale features refer to feature maps that are processed by the backbone network, ResNet101, and then upsampled and stacked using the feature pyramid. These feature maps are sized 152×152×256 and contain feature information of objects at different scales. Compared to the existing U-Net architecture, one improvement of this method is that it replaces the U-Net's upsampling operation with a feature pyramid. This improvement saves a significant number of parameters and effectively improves computational speed without compromising detection accuracy. Specific parameters are shown in the #P column of Tables 1 and 2. Both this method and the U-Net employ upsampling, but there are certain differences. This method uses nearest neighbor interpolation to upsample the feature maps of various sizes extracted by the backbone. This is a simple operation and requires no parameter learning. In contrast, the U-Net architecture uses transposed convolution or interpolation (such as bilinear interpolation) plus convolution for upsampling.

[0032] These include: S1-1. First, through the backbone network, i.e., ResNet-101, four sets of feature maps of different scales are obtained through layer-by-layer convolution.

[0033] S1-2. To solve the problem of different feature maps, a feature pyramid structure is used to construct bottom-up and top-down feature fusion paths, fusing feature maps of different scales to generate a feature pyramid with rich multi-scale information.

[0034] S1-3. Splice the feature pyramids after FPN to obtain a feature map of 152×152×256 shape, where 152×152 is the image resolution and 256 is the number of feature maps.

[0035] S2. Subsequently, a dynamic selection mechanism is used to adaptively select appropriate kernels for different targets based on the extracted multi-scale features, including: S2-1. The input feature map passes through the large kernel selection module, and its receptive field is expanded through 5×5 and 7×7 convolutions to obtain information around the target.

[0036] S2-2. A spatial selection mechanism is then used to weight and fuse the results obtained from different large-core selection modules. This process expands the receptive fields of multiple features.

[0037] The features here are the features after 5×5 and 7×7 processing, and the features after these two large kernel convolutions are weighted and fused.

[0038] S2-3 fuses the features obtained through the spatial selection mechanism with the original features obtained by FPN to obtain the final comprehensive features.

[0039] S3. Finally, the comprehensive features are input into the decoder to obtain the predicted coordinates and target categories. The decoder, i.e., the prediction head, decodes the feature map output by LSK to obtain four branches: heat map, offset, box parameter, and direction map. Among them, the tensor shape obtained by the heat map is: 15×152×152, the tensor shape obtained by the offset is: 2×152×152, the tensor shape obtained by the box parameter is: 10×152×152, and the tensor shape obtained by the direction map is: 1×152×152. The present invention combines the feature pyramid with the boundary perception vector to capture the rotated bounding box of the target, significantly reducing the amount of calculation and obtaining higher accuracy, and uses large kernel convolution to strengthen the intermediate features to further improve the final accuracy.

[0040] like Figure 1 As shown, the FLBANet proposed in the present invention is built based on the feature pyramid network (FPN) architecture and is specifically designed for target detection and related tasks. It includes a backbone network, a neck module, an LSK module, and a head branch (the backbone network (also the main network) is used to extract features. In the present invention, the backbone network adopts ResNer101. The feature pyramid is FPN, which helps the model better handle targets of different sizes by constructing multi-scale feature representations. Target detection is divided into the following four parts: the backbone network (backbone), the neck network (neck), and the detection head (Head). The backbone network uses ResNet101, the neck network uses FPN, and the detection head uses the detection head of BBAVector. The model of the present invention is based on the feature extraction and target perception parts of the baseline model. The FPN module is used to improve the feature extraction capability of targets of different sizes, and the LSK module is used to perceive the surrounding information of the target to improve the overall accuracy.

[0041] In step S1, the backbone network is implemented based on the ResNet architecture. Specifically, the ResNet101 backbone processes the RGB input image. ,in and W denote the height and width of the input image, respectively. The backbone network extracts hierarchical features by subjecting the input image to a series of convolution, batch normalization, and ReLU activation layers. Each residual block in the backbone network includes convolutional layers with skip connections, which enables the network to learn residual mappings and facilitates the flow of information between layers.

[0042] At the top of the backbone network, the Feature Pyramid Network (FPN) is used as the neck module. The convolutional layer refines the upsampled feature map. The role of the neck module is to fuse the hierarchical features extracted by the backbone.

[0043] Feature Pyramid Networks (FPNs) are a key component in many computer vision tasks, particularly object detection and semantic segmentation. In this paper, the core of FPNs is the efficient integration of multi-scale features. By using feature pyramids at different scales, the network can capture both fine-grained and coarse-grained information. The original input image has a resolution of 608×608. After backbone downsampling, the resolution of the feature maps gradually decreases. Features at different scales are obtained, with resolutions of 152×152, 76×76, 38×38, and 19×19, respectively. Among these features at different scales, high-resolution feature maps typically contain more specific, low-level semantic information, such as texture and edges, while low-resolution feature maps contain more abstract, high-level semantic information, such as object and scene categories. The feature pyramid combines features at different scales and upsamples them to 152×152, thus integrating features at different scales while maintaining consistent resolution. Coarse-grained features, such as the smallest feature map in the figure above, have a resolution of 19×19. Each point contains features of a large area and is therefore often interpreted as high-level semantic information. Fine-grained information, on the other hand, corresponds to the largest feature map in the figure above, with a resolution of 152×152. It contains texture information of objects, such as edges, and is well-suited for processing.

[0044] In the upsampling process of FPN, a stacking operation is used to combine deep and shallow features. This allows the network to utilize both high-level semantic and low-level texture information. The process first upsamples the deep feature map to the size of the shallow feature map through bilinear interpolation, and then uses The convolutional layer refines the upsampled map. After connecting with the shallow feature map, The convolutional layer refines the channel features, and batch normalization and ReLU activation are applied in the latent layer to obtain better training stability and representation ability. Deep features and shallow features are relative. Shallow features usually refer to features extracted in the first few layers of the neural network (close to the input layer), and deep features refer to features extracted in the last few layers of the neural network (close to the output layer). The resolution of the output from shallow features to deep features in the present invention is 152×152, 76×76, 38×38, and 19×19, respectively. Shallow features are fine-grained information because they focus on local details and specific visual elements in the image. Deep features are coarse-grained information because they focus on the overall structure and semantics of the image rather than specific details.

[0045] FPNs typically have a top-down pathway and lateral connections. The top-down pathway upsamples coarse-scale features, and the lateral connections merge these features with bottom-up features of the same scale. Mathematically speaking, 、 and Represent bottom-up, top-down upsampling and output feature maps respectively ( is the pyramid level), the process is as follows: ; in is an upsampling operation based on bilinear interpolation, and: ; in, Is a learnable weight matrix (can be an identity or a simple convolutional layer) that adjusts the bottom-up feature contributions. Output feature map ,in and (s is the scale factor, Make the output 4 times smaller than the input image). Indicates the pyramid level is The top-down upsampled feature map of represents the output feature map of pyramid level l, represents the spatial dimension of the output feature map, where C is the number of channels, and are the height and width of the feature map, respectively, after scaling by the scale factor s.

[0046] The FPN's function is to fuse feature maps of different sizes output by the backbone. To achieve horizontal connections and bottom-up feature merging, images with smaller feature maps are amplified through upsampling and convolution, and then stitched together to create feature maps of a uniform size of 152×152.

[0047] In step S2, the LSK module processes the input feature map and enhances the representation capability through a series of convolution and attention mechanisms. Output feature map (In this article ) is then transformed into four branches: heatmap ( ), offset ( ), frame parameters ( ) and directional patterns ( Here, K represents the number of dataset categories and s=4 represents the scaling factor. This transformation is done using two The kernel and the convolutional layer with 256 channels are implemented.

[0048] LSK inputs and outputs feature maps, and the size of the feature maps remains unchanged. Only the numerical values of the feature maps are altered, aiming to increase the response to the target's feature map and suppress the relevant values in the background's feature map. After processing, decoding is required. The decoding head consists of a heat map, offset, box parameters, and a directional map.

[0049] The LSK module primarily adjusts the multi-scale features extracted by the FPN, selecting appropriate kernels for different objects. The integrated features are then fed into the decoder to obtain predicted coordinates and object categories. The LSK component integrates the features, which are then used by the detection head to separate them into four components: the heatmap and offset. Finally, the decoder decodes these four components to produce the final rotated bounding box and predicted category. The LSK module consists of two sub-blocks: the large kernel selection (LK selection) sub-block and the feed-forward network (FFN) sub-block. In the LK selection sub-block of the LSK module, a large kernel convolution is first constructed by factorizing the large kernel convolution into a series of depthwise convolutions with increasing kernel size and dilation rate. This process generates multiple features with different receptive fields (the multiple features here refer to two features obtained with different receptive fields, 5×5 and 7×7). Then, a spatial selection mechanism is used to weight and fuse these features according to the input features (the various features obtained previously are obtained based on the feature pyramid and are fused by the features of different backbone stages. The various feature inputs obtained by the LSK part are the feature maps of the feature pyramid of the previous stage. The various features of different receptive fields mentioned here refer to targets of different sizes, such as cars and boats, which have large scale differences. LSK increases the surrounding information of these targets and takes into account the features of the surrounding water and land, so it is described as multiple features of different receptive fields). In large kernel convolution, the kernel size is , expansion rate and receptive field The expansion formula is as follows: ; ; Let's use a specific example to illustrate. A large kernel convolution can be divided into several small kernel convolutions to achieve the same receptive field effect. For example, using a large kernel convolution k=23, d=1, the receptive field RF=23. Using two small kernel convolutions but Get the same receptive field.

[0050] in, is the size of the i-th convolution kernel, which means the width and height of the convolution kernel. is the dilation rate of the i-th convolution. The dilation rate is used to control the size of the receptive field of the convolution kernel, and the receptive field is expanded by inserting holes between the convolution kernel elements. is the receptive field of the i-th convolution. The receptive field refers to the range of the input feature map that the convolution kernel can cover, which indicates the area size of the input image that the convolution kernel can perceive.

[0051] It means that the size of the i-th convolution kernel is not less than the size of the previous convolution kernel. This shows that the size of the convolution kernel is increasing, thereby gradually expanding the receptive field.

[0052] It means that the expansion rate of the first convolution is 1, which means there is no hole and the receptive field is determined only by the size of the convolution kernel. The expansion rate of subsequent convolutions is increasing, but it will not exceed the receptive field size of the previous convolution. This ensures that the receptive field is gradually expanded while avoiding receptive field overlap caused by excessive expansion.

[0053] It means that the receptive field of the i-th convolution is determined by the current convolution kernel size, the dilation rate and the receptive field of the previous convolution. Indicates the additional receptive field expansion brought about by the expansion of the current convolution kernel. It means that the receptive field of the previous convolution is further expanded.

[0054] This design enables the LSK module to better simulate the detection of objects of different scales in remote sensing scenes and achieve state-of-the-art performance on multiple remote sensing tasks. In addition, this design enables the method of the present invention to achieve an improvement in mAP with a limited increase in the number of parameters.

[0055] Tables 1 and 2 compare the effect of adding LSK on the number of parameters. Using the DOTA dataset as an example, FBA-R and FLBA-R represent the comparison between using LSK and not using LSK, respectively. With ResNet-101 as the backbone, the computational overhead increases by 131.39-128.68=2.71GB, and the number of parameters increases by 459,600-458,400=1,200. This results in a performance improvement of 75.84-75.72=0.12%.

[0056] For the four branches of heatmap, offset, box parameter and direction map, the details are as follows: Heatmap ( ): The heat map branch is used to predict the target at different locations in the image. The output of this branch is ,in is the number of dataset categories. Heatmaps are usually used to locate specific key points in the input image. This paper applies it to the center point detection of targets in any direction in aerial images. For ground truth, the center point of the given rotating target box is The present invention places a 2D Gaussian around the center point Generate ground truth heatmap ,in The frame size is adaptive.

[0057] The loss function formula is ; in, and are the true and predicted heatmap values, Index pixel location, is the target number, and is an empirically selected hyperparameter. In a specific embodiment, the preferred hyperparameter is =2, =4, one of the pixels is actually the target, and the predicted value is close to the background and is taken as 0.01, that is: =0.01, =1, then the loss of this pixel is , finally all pixel losses need to be accumulated and multiplied by (-1 / 15), where category N is 15. Figure 7 is the predicted result. Figure 7 The left side is the input image, and the right side is the output heat map result. It can be seen that for the center of the target, the heat map tends to be 1, and for the background part, the heat map prediction is 0.

[0058] K is the number of categories in the dataset. middle, The coordinates of the center point of the rotation target frame. is the pixel position coordinate on the heat map. is the standard deviation of the Gaussian distribution. In the loss function, Represents the loss value for each target. Used to adjust the weight of positive samples in the loss function. Used to adjust the weight of negative samples in the loss function.

[0059] Offset ( ): The Offset branch predicts the offset value of each target position. Its output is In the inference phase of heatmap-based aerial image object detection, the predicted offset map Compensates for differences between quantized floating-point and integer center points.

[0060] For the true center point On the input image, the offset o between the scaled floating point center point and the quantized center point is calculated as follows: ; Use smooth Loss function optimization offset: ; Smooth L1 loss is the most commonly used loss function in machine learning. For example, if the two values of one category are 0.3 and 0.5, then , then the offsets of all pixels of this class are accumulated, and then the loss function of this branch is obtained by averaging all classes.

[0061] in, is the coordinate of the ground truth center point on the input image. The offset o is the offset between the scaled floating point center point and the quantized center point. This is a round-down operation, which means taking the largest integer that is less than or equal to the value. is the number of targets. is the true offset of the k-th target. is the predicted offset of the k-th target. is the loss function of the offset branch.

[0062] Box parameters ( ): This branch predicts the box parameters of the target bounding box. The output is The present invention uses box parameter mapping Captures bounding box perception vectors (BBAVectors) representing the object's rotation bounding box (OBB). Each object's bounding box is represented by five two-dimensional vectors, with a total of ten channels. Given an object's center point and direction angle , in In the Cartesian coordinate system centered at is the distance from the center point to the bounding box, the vector The calculation is as follows ;

[0063] The ground truth value of the box parameter map is generated by the bounding box ground truth value, and its training loss is the smoothing loss

[0064] ; The purpose of setting a training loss function is to ensure that the final predicted value is closer to the actual value. For example, if a box is predicted to be in the upper left corner, but the actual value is in the lower right corner, a simple L1 loss is used to gradually move the predicted value closer to the right corner. The closer the predicted value is, the smaller the loss is, until the predicted value and the actual value overlap. At this point, the loss reaches a minimum of 0.

[0065] is the center coordinate of the target. is the i-th boundary perception vector. is the ground-truth box parameter of the k-th target. is the predicted box parameter of the kth target. is the loss function of the box parameter branch. is the direction angle of the target. is the distance from the center point c to the bounding box.

[0066] Directional diagram ( ): The directional map branch predicts the direction of the target. Its output is , the present invention utilizes the directional pattern The bounding box is divided into horizontal bounding box (HBB) and rotation bounding box (RBB). HBB is defined as a bounding box with a certain threshold. Orientation angle inside , while RBB has a value outside this threshold Apply the sigmoid function to map the predicted direction value to the range , the ground truth of the directional map is set to HBB is 0, RBB is 1. The training loss of the directional map is the binary cross entropy loss ; If this part determines that a rotation box is a horizontal box, it is assigned a value of 0, and if it is determined to be a rotation box, it is assigned a value of 1.

[0067] The total loss of the network is the weighted sum of all output graph losses and is given by ; in 、 and They are the weights of offset loss, box parameter loss and direction map loss, respectively, and these weights can be adjusted based on experience.

[0068] is the predicted direction value of the i-th pixel position. is the true direction value of the i-th pixel position. If it is a horizontal bounding box (HBB), then = 0. If it is a rotated bounding box (RBB), then =1. is the loss function of the heatmap branch. is the loss function of the offset branch, is the loss function of the box parameter branch. is the loss function of the directional graph branch. 、 、 is the weight coefficient used to balance the losses of different branches.

[0069] The loss function is calculated and back-propagated by the GPU. The ultimate goal of this invention is to get the overall loss function target close to 0, which means that the network fits the data.

[0070] In this paper, we developed a method for rotating object detection in optical remote sensing images by integrating an FPN backbone with a BBAVectors head. The FPN backbone effectively addresses the challenge of multi-scale object detection and enables robust feature extraction at different levels. Furthermore, the LSK feature integration module effectively separates the background from the target. Finally, the BBAVectors head accurately determines the target's rotated bounding box.

[0071] The decoder is equivalent to post-processing the output of the above four parts.

[0072] The tensor shape for the heatmap is 15×152×152, the offset is 2×152×152, the box parameters is 10×152×152, and the orientation map is 1×152×152. The post-processing pipeline is as follows: First, the heatmap tensor represents 15 categories, each of which is a 152×152 resolution center point heatmap. The offset is due to the fourfold reduction in heatmap resolution (608 to 152). Therefore, a floating-point number, the offset, is used to represent the center point offset, correcting the target center point. The box parameters have a shape of 10×152×152. The first eight parameters represent the xy distances from the center point to the four bounding boxes, and the last two parameters represent the width and height of the rotated box. The orientation map has only one parameter. If the horizontal angle is less than 5°, it is treated as a horizontal box; otherwise, it is processed as a rotated box.

[0073] The present invention is verified on the public datasets HRSC 2016 and DOTA V1.0.

[0074] HRSC 2016: Released in 2016, the HRSC 2016 dataset contains 1,680 images, 1,061 of which are fully annotated. The dataset features 2,976 ship objects, divided into 436 training images, 181 validation images, and 444 test images. The spatial resolution of the images ranges from 0.4 to 2 meters, clearly capturing ship details. Image sizes range from 300 × 300 pixels to 1,500 × 900 pixels. The dataset uses the rotated bounding box (OBB) annotation format, which is used to accurately capture ship orientation and pose.

[0075] DOTA V1.0: The DOTA-V1.0 dataset consists of 2,806 aerial images from various sensors and platforms, yielding rich and diverse image features. Objects in these images exhibit a wide range of scales, orientations, and shapes, making detection algorithms extremely challenging. Image resolutions range from 800×800 pixels to 4,000×4,000 pixels. There are 188,282 fully annotated instances belonging to 15 different classes: airplane (PL), baseball field (BD), bridge (BR), ground track and field (GTF), small vehicle (SV), large vehicle (LV), ship (SH), tennis court (TC), basketball court (BC), storage tank (ST), soccer field (SBF), roundabout (RA), harbor (HA), swimming pool (SP), and helicopter (HC). The dataset is divided into 1,411 training images, 458 validation images, and 937 test images.

[0076] Throughout the training process, the present invention trained the network in batches of 12 on two NVIDIA RTX 3090 TI GPUs. For the DOTA dataset, the present invention trained the network for approximately 200 iterations to ensure full convergence and learning of complex patterns in the data. Given the relatively small size and unique features of the HRSC2016 dataset, the present invention trained the network for 100 iterations, achieving excellent performance. In addition, the present invention filtered out images without objects to promote better network convergence. The present invention measured the speed of the proposed network on a single NVIDIA 3090 TI GPU using the HRSC 2016 and DOTA V1.0 datasets, illustrating the computational efficiency of the present invention's method in practical scenarios.

[0077] Quantitative Results: To verify the effectiveness of each component in our proposed framework, Tables 1 and 2 summarize the results of ablation experiments conducted on the DOTA V1.0 and HRSC2016 datasets. Specifically, FBA-R represents the results without the LSK module, while FLBA-R represents the results with the LSK module. FLBA-D represents the results with a lightweight backbone network.

[0078] Table 1: DOTA V1.0 experimental results ; Table 1 summarizes the results, where we present an ablation study of different components of the DOTA dataset. The evaluated methods include the baseline method (BBAVectors), our FBA-R method (using ResNet + FPN), the FLBA-R method (using ResNet + FPN + LSK), and the FLBA-D method (using a decoupled network with FPN and LSK).

[0079] A baseline method using ResNet-101 achieved a mAP of 75.36% with 176.61 G FLOPs, 53.43 million parameters, and an inference speed of 39.41 FPS. By integrating the FPN architecture with the ResNet backbone, the proposed FBA-R method demonstrated competitive performance while reducing computational complexity and improving inference speed. For example, using ResNet-101, FBA-R achieved a mAP of 75.72% with a reduced FLOP of 128.68 G, 45.84 million parameters, and an inference speed of 43.14 FPS. This demonstrates that FPN effectively improves feature extraction efficiency without significantly degrading performance. Further integrating the LSK module into the FBA-R method yielded the FLBA-R method, which achieved higher accuracy at a slightly increased computational cost. For example, using ResNet-101, FLBA-R achieved a mAP of 75.84% with 131.39 G FLOPs, 45.96 million parameters, and an inference speed of 43.43 FPS. This demonstrates that the LSK module effectively increases the model's sensitivity to local features, thereby improving detection performance. The FLBA-D method, based on a decoupled network, achieves a balance between accuracy and computational efficiency. Using DecoupleNet D2, FLBA-D achieves a mAP of 73.59%, 81.11 G FLOPs, 9.49 million parameters, and an inference speed of 43.46 FPS. This highlights the effectiveness of decoupled networks in reducing model complexity while maintaining competitive performance.

[0080] The proposed methods, FBA-R and FLBA-R, consistently outperform a baseline method (BBAVectors) on different ResNet backbones. The integration of FPN and LSK modules significantly improves detection accuracy while reducing computational complexity and increasing inference speed. Furthermore, the FLBA-D method, utilizing a disaggregated network, achieves a favorable trade-off between accuracy and computational efficiency, making it suitable for resource-constrained environments. Overall, these results validate the effectiveness of the proposed methods in improving oriented object detection on the DOTA dataset and highlight their potential for real-world applications.

[0081] Table 2: HRSC2016 experimental results ; As shown in Table 2, ablation experiments on the HRSC2016 dataset demonstrate the effectiveness of our proposed method. Compared to the baseline method (ResNet-101 backbone, mAP 88.22%, 176.53 G FLOPs, 53.43M parameters, 49.87 FPS), our FBA-R method (ResNet+FPN) achieves comparable or higher mAP values while reducing computational complexity and improving inference speed. For example, FBA-R using ResNet-101 achieves 89.68% mAP, 128.59 G FLOPs, 45.83M parameters, and 54.14 FPS. The FLBA-R method (ResNet+FPN+LSK) further improves performance, achieving 89.73% mAP using ResNet-101 while maintaining reasonable computational efficiency (131.31 G FLOPs, 45.95M parameters, 49.48 FPS). The FLBA-D method using DecoupleNet D2 achieves 86.05% mAP using only 9.49M parameters and 53.85 FPS, highlighting its efficiency in resource-constrained applications. Overall, our proposed FLBA-R and FLBA-D show significant improvements in accuracy, computational efficiency, and inference speed compared to the baseline.

[0082] Qualitative results: In order to demonstrate the advantages of the feature pyramid and feature integration architecture, the present invention is also compared with the baseline method. The method of the present invention uses the same decoder and post-processing steps as the baseline method, and the training process is the same as the baseline method, and no data enhancement is used. As can be seen from Tables 1 and 2, the method proposed in the present invention outperforms the baseline method by 1.91% and 0.48% on the HRSC 2016 and DOTA datasets, respectively. Compared with the U-Net structure used in the baseline method, the feature pyramid structure is more suitable for rotated target detection. It not only has faster inference speed and fewer parameters, but also has higher accuracy.

[0083] Figure 4 The qualitative results of the proposed method are compared in detail with the baseline methods in various detection scenarios. Figure 4 (a) shows that the baseline method cannot handle objects with low contrast or partially blurred. In contrast, our method successfully recognizes these objects, demonstrating its robustness in capturing a wider range of object appearances. Figure 4 (b) shows an incorrect detection by the baseline method, where a small car is mistakenly detected as a large car. This problem may be due to the loss of detail during feature extraction and the limited ability to distinguish between different objects. However, the proposed method can accurately distinguish between objects and non-objects, achieving more accurate and reliable detection. Figure 4Panel (c) shows that the baseline method produces redundant detections, where multiple bounding boxes are applied to the same object and background is mistakenly detected as a different object. This redundancy leads to confusion and inefficiency in the subsequent analysis stage. The proposed method effectively alleviates this problem, thereby improving the clarity and practicality of the detection results. Figure 4 (d) shows that in the case of rotating target detection in aerial images, the baseline method will randomly predict the angle of the non-directional target. The method of the present invention tends to unify the targets with random angles. Figure 4 (e) shows that the baseline method often assigns low confidence scores to correct detections, which may lead to the rejection of valid detections during post-processing. The method of the present invention assigns higher confidence scores to correct targets, thereby reducing the possibility of false positives and improving overall detection performance. In summary, the method of the present invention shows significant advantages over the baseline method in these scenarios. These improvements are crucial for improving the reliability and effectiveness of remote sensing applications.

[0084] The present invention further conducts a qualitative comparison of the visualization results on the DOTA V1.0 dataset. Figure 5 Detailed image information is presented. The input image contains objects of varying categories and sizes (including common objects, small objects, and objects with large aspect ratios). For small objects like a car, medium-sized objects like a swimming pool, and large-aspect-ratio harbors, the method accurately predicts rotated bounding boxes aligned with the objects. This demonstrates the effectiveness of the proposed method for detecting rotated objects in optical remote sensing images.

[0085] exist Figure 6 The figure shows the visualization of feature maps. The present invention shows the image detection results and a comparison of the feature maps of the same channel before and after processing through the LSK module. After processing through the LSK module, the feature maps generated at each stage are more closely aligned with the semantic meaning of the corresponding prediction parameters. Before processing through the LSK module, the features are more dispersed when obtaining the center position information x and y. After processing through the LSK module, the features become more distinct, more concentrated at the center of the target, and less background information. This phenomenon is likely the result of the gradual guidance of different feature kernels in the LSK module, further confirming the effectiveness of the proposed architecture approach.

[0086] Any process or method described in the flowchart of the present invention or in other ways herein can be understood as representing a module, segment or portion of code including one or more executable instructions for implementing specific logical functions or process steps, which can be implemented in any computer-readable medium for use by an instruction execution system, device or apparatus. The computer-readable medium can be any medium that stores, communicates, propagates or transmits a program for use by an execution system, device or apparatus, including read-only memory, magnetic disk or optical disk, etc.

[0087] Throughout this specification, reference to terms such as "embodiment" and "example" indicates that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the schematic representations of these terms do not necessarily refer to the same embodiment or example. Furthermore, those skilled in the art may combine or integrate different embodiments or examples described in this specification, as well as features therein, without creating any inconsistency.

[0088] Although the above content has shown and described the embodiments of the present invention, it can be understood that the above embodiments are exemplary and cannot be understood as limitations of the present invention. Ordinary technicians in this field can perform update operations such as changes, modifications, replacements and variations on the above embodiments within the scope of the present invention.

Claims

1. A remote sensing image target detection method based on feature pyramid and boundary perception vector, characterized in that: The method comprises the following steps: S1. Feed the input image into the feature pyramid backbone to extract multi-scale features; S2. A dynamic selection mechanism is used to adaptively select appropriate kernels for different targets based on the extracted multi-scale features; S3. Input the comprehensive features into the decoder to obtain the predicted coordinates and target category, thereby obtaining the rotated bounding box of the target.

2. The method for remote sensing image target detection based on feature pyramid and boundary perception vector according to claim 1, characterized in that: The backbone network of the method includes a backbone network, a neck module, an LSK module and a detection head; the backbone network is implemented based on the ResNet architecture, and the ResNet101 backbone processes RGB input images; the backbone network extracts hierarchical features by performing a series of convolution, batch normalization and ReLU activation layers on the input image.

3. The method for remote sensing image target detection based on feature pyramid and boundary perception vector according to claim 2, characterized in that: Step S1 includes: S1-1. First, through the backbone network, four sets of feature maps of different scales are obtained through layer-by-layer convolution; S1-2. To address the issue of different feature maps, we use a feature pyramid structure to construct bottom-up and top-down feature fusion paths, fusing feature maps of different scales to generate a feature pyramid with multi-scale information. S1-3. Splice the feature pyramids after FPN to obtain a feature map of 152×152×256 shape, where 152×152 is the image resolution and 256 is the number of feature maps.

4. The method for remote sensing image target detection based on feature pyramid and boundary perception vector according to claim 2, characterized in that: Step S2 includes: S2-1. The input feature map passes through the large kernel selection module, and its receptive field is expanded through 5×5 and 7×7 convolutions to obtain information around the target. S2-2. Using a spatial selection mechanism, these features are weighted and fused according to the input features to generate an expanded receptive field of multiple features.

5. The method for remote sensing image target detection based on feature pyramid and boundary perception vector according to claim 4, characterized in that: The core of the feature pyramid network is to effectively integrate multi-scale features. Through feature pyramids of different scales, the network can capture fine-grained and coarse-grained information. During the upsampling process, jump connections are used to combine deep and shallow features.

6. The method for remote sensing image target detection based on feature pyramid and boundary perception vector according to claim 5, characterized in that: The feature pyramid network first upsamples the deep feature map to the size of the shallow feature map by bilinear interpolation, and then uses The convolutional layer refines the upsampled map; After connecting with the shallow feature map, The convolutional layers refine the channel features, and batch normalization and ReLU activation are applied in the latent layers.

7. The method for remote sensing image target detection based on feature pyramid and boundary perception vector according to claim 5, characterized in that: The LSK module processes the input feature map and enhances the representation capability through a series of convolution and attention mechanisms.

8. The method for remote sensing image target detection based on feature pyramid and boundary perception vector according to claim 7, characterized in that: The LSK module converts the feature map into four branches: heat map, offset, box parameter and direction map.

9. The method for remote sensing image target detection based on feature pyramid and boundary perception vector according to claim 8, characterized in that: When the LSK module transforms the feature map, it uses two Convolutional layer implementation with kernel and 256 channels.

10. The method for remote sensing image target detection based on feature pyramid and boundary perception vector according to claim 7, characterized in that: The LSK module consists of two sub-blocks: a large kernel selection sub-block and a feed-forward network sub-block. In the large kernel selection sub-block of the LSK module, a large kernel convolution is first constructed by decomposing the large kernel convolution into a series of depthwise convolutions with increasing kernel size and dilation rate. This process produces multiple features with different receptive fields. A spatial selection mechanism is then employed to weight and fuse these features according to the input features.

Citation Information

Patent Citations

  • Remote sensing image target detection method based on deep learning

    CN113468993A

  • Rotary SAR ship target detection method based on bidirectional feature pyramid network

    CN115294452A

  • Mixed anchor point remote sensing image target detection method based on multi-scale large kernel convolution

    CN120279310A

  • Commutator inner side image defect detection method based on fusible feature pyramid

    WO2024208100A1