A target detection method and device for remote sensing images

By using intensive feature pyramid network and attention module in remote sensing image object detection, the problem of low accuracy of detection of targets with large scale differences in the prior art is solved, and efficient detection of targets at different scales is achieved.

CN114359709BActive Publication Date: 2025-06-27BEIJING NORTH INTELLIGENT MAP INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111484783.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-07
Publication Date
2025-06-27
Estimated Expiration
2041-12-07

AI Technical Summary

Technical Problem

The existing convolutional neural networks have low accuracy when detecting targets with large scale differences in remote sensing images, especially in complex scenarios, which are prone to missed detection of small targets.

Method used

Using an object detection model based on intensive feature pyramid networks, multi-scale features are extracted through the fusion of upsampling and downsampling feature pyramid networks, and feature extraction capabilities are enhanced through attention modules.

Benefits of technology

The detection accuracy of targets at different scales is improved, especially when detecting small targets, the accuracy is significantly improved, and the robustness and adaptability of the model are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114359709B_ABST
    Figure CN114359709B_ABST
Patent Text Reader

Abstract

The present invention provides a target detection method and device for remote sensing images, including: determining a target remote sensing image; inputting the target remote sensing image into a target detection model to obtain a target detection result output by the target detection model; the target detection result includes the target type and target location in the target remote sensing image; the target detection model is trained based on sample remote sensing images, as well as target type samples and target location samples in the sample remote sensing images, and is used to detect the target type and target location in the target remote sensing image; the target detection model is constructed based on a dense feature pyramid network; the dense feature pyramid network includes an upsampling feature pyramid network and a downsampling feature pyramid network. The target detection method and device for remote sensing images provided by the present invention can fuse the features of targets at different scales by constructing a target detection model based on a dense feature pyramid network, thereby improving the detection accuracy of targets at different scales.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method and device for target detection of remote sensing images. Background Art

[0002] With the development of satellite imaging and deep learning, remote sensing target detection has become a hot issue in computer vision research and can be widely applied to fields such as navigation, disaster warning, building detection, etc.

[0003] In the problem of deep learning for remote sensing target detection, since the convolutional neural network has a strong ability to mine spatial context information, it is widely applied to the target detection of remote sensing images.

[0004] Existing convolutional neural networks cannot accurately detect targets with large scale differences in remote sensing images. Summary of the Invention

[0005] In view of the problems existing in the prior art, embodiments of the present invention provide a method and device for target detection of remote sensing images.

[0006] The present invention provides a method for target detection of remote sensing images, including: determining a target remote sensing image;

[0007] Inputting the target remote sensing image into a target detection model to obtain a target detection result output by the target detection model; the target detection result includes the target type and target position in the target remote sensing image;

[0008] The target detection model is trained based on sample remote sensing images, as well as target type samples and target position samples in the sample remote sensing images, and the target detection model is used to detect the target type and target position in the target remote sensing image;

[0009] The target detection model is constructed based on a dense feature pyramid network; the dense feature pyramid network includes an upsampling feature pyramid network and a downsampling feature pyramid network.

[0010] According to the method for target detection of remote sensing images provided by the present invention, the target detection model includes a feature extraction network, the dense feature pyramid network, and a detection network;

[0011] The step of inputting the target remote sensing image into the target detection model to obtain the target detection result output by the target detection model includes:

[0012] Inputting the target remote sensing image into the feature extraction network to obtain feature images of multiple scales of the target remote sensing image output by the feature extraction network;

[0013] Input each feature image into the upsampling scale layer corresponding to the scale in the upsampling feature pyramid network, and obtain the upsampling output features output by each upsampling scale layer;

[0014] Input each upsampling output feature into the downsampling scale layer corresponding to the scale in the downsampling feature pyramid network, and obtain the fused feature map output by each downsampling scale layer;

[0015] Input each fused feature map into the detection network to obtain the target detection result.

[0016] According to the target detection method for remote sensing images provided by the present invention, the step of inputting each feature image into the upsampling scale layer corresponding to the scale in the upsampling feature pyramid network and obtaining the upsampling output features output by each upsampling scale layer includes:

[0017] Input the feature image of any scale into the corresponding scale layer of the upsampling feature pyramid network. The corresponding scale layer of the upsampling feature pyramid network fuses the feature image of any scale with the upsampled features of the feature image of the previous scale of any scale, and the upsampling output features output by the scale layer of the previous scale, to obtain the upsampling output features output by the scale layer of any scale.

[0018] According to the target detection method for remote sensing images provided by the present invention, the step of inputting each upsampling output feature into the downsampling scale layer corresponding to the scale in the downsampling feature pyramid network and obtaining the fused feature map output by each downsampling scale layer includes:

[0019] Input the upsampling output feature of any scale into the corresponding scale layer of the downsampling feature pyramid network. The corresponding scale layer of the downsampling feature pyramid network fuses the feature image of any scale with the downsampled features of the upsampling output features of the next scale of any scale, and the fused feature map output by the scale layer of the next scale, to obtain the fused feature map output by the scale layer of any scale.

[0020] According to the target detection method for remote sensing images provided by the present invention, the feature extraction network includes a plurality of sequentially connected residual modules; the step of inputting the target remote sensing image into the feature extraction network and obtaining the feature images of multiple scales of the target remote sensing image output by the feature extraction network includes:

[0021] Input the target remote sensing image into the feature extraction network, and obtain the feature images of multiple scales output by a plurality of residual modules in the feature extraction network;

[0022] Each residual module in the feature extraction network includes an attention module.

[0023] According to the object detection method for remote sensing images provided by the present invention, the determination of the target remote sensing image includes:

[0024] Obtain an initial remote sensing image;

[0025] Perform size normalization processing on the initial remote sensing image to determine the target remote sensing image.

[0026] The present invention also provides an object detection device for remote sensing images, including:

[0027] A determination unit for determining a target remote sensing image;

[0028] An acquisition unit for inputting the target remote sensing image into an object detection model and obtaining an object detection result output by the object detection model; the object detection result includes the object type and object position in the target remote sensing image;

[0029] The object detection model is trained based on sample remote sensing images, as well as object type samples and object position samples in the sample remote sensing images, and the object detection model is used to detect the object type and object position in the target remote sensing image;

[0030] The object detection model is constructed based on a dense feature pyramid network; the dense feature pyramid network includes an upsampling feature pyramid network and a downsampling feature pyramid network.

[0031] According to the object detection device for remote sensing images provided by the present invention, it further includes:

[0032] A normalization module for performing size normalization processing on the initial remote sensing image to determine the target remote sensing image.

[0033] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the steps of the object detection method for remote sensing images as described in any one of the above.

[0034] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the object detection method for remote sensing images as described in any one of the above.

[0035] The object detection method and device for remote sensing images provided by the present invention can fuse the features of objects at different scales by constructing an object detection model based on a dense feature pyramid network, thereby improving the detection accuracy of objects at different scales. Brief Description of the Drawings

[0036] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0037] Figure 1 is a schematic flowchart of the object detection method for remote sensing images provided by the present invention;

[0038] Figure 2 is one of the schematic structural diagrams of the object detection model provided by the present invention;

[0039] Figure 3 is another schematic structural diagram of the object detection model provided by the present invention;

[0040] Figure 4 is a schematic structural diagram of the dense pyramid network provided by the present invention;

[0041] Figure 5 is a schematic structural diagram of the feature pyramid network provided by the present invention;

[0042] Figure 6 is a schematic structural diagram of the residual unit based on SGE attention provided by the present invention;

[0043] Figure 7 is a schematic structural diagram of the SGE attention module provided by the present invention;

[0044] Figure 8 is a schematic structural diagram of the object detection device for remote sensing images provided by the present invention;

[0045] Figure 9 is a schematic structural diagram of the electronic device provided by the present invention. Detailed Embodiments

[0046] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention fall within the scope of protection of the present invention.

[0047] It should be noted that in the description of the embodiments of the present invention, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element. The orientation or positional relationship indicated by terms such as "upper", "lower", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be construed as a limitation on the present invention. Unless otherwise expressly specified and limited, the terms "mounted", "connected" and "coupled" shall be construed broadly. For example, it may be a fixed connection, a detachable connection or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the internal communication of two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention may be understood according to specific circumstances.

[0048] The mainstream neural networks for object detection include region proposal-based neural networks and bounding box regression-based neural networks.

[0049] Most of the region proposal-based neural networks are two-stage networks. First, the approximate object location is obtained according to the region proposal network, and then the category of the object is accurately predicted and the accurate prediction box is regressed. This step-by-step learning strategy makes the detection accuracy of the network relatively high, but it also results in a long detection time, making it difficult to achieve efficient processing, and the training time for remote sensing images with a large input image size is too long. Typical representatives of such networks include the Recursive Convolutional Neural Network (Region with Convolutional Neural Network feature) series, including Fast RCNN, Faster RCNN, and Mask RCNN, etc.

[0050] Most of the bounding box regression-based neural networks are single-stage networks, regarding the entire prediction process as a regression process. The simplification of the process not only does not lose too much accuracy, but also brings an improvement in speed. Representatives of such networks include: Single Shot Detection (SSD), single-stage object detection algorithms (YOLO series), EfficientDet, etc.

[0051] The YOLO series of networks are typical regression-based neural networks, and multiple versions such as YOLO, YOLOv2, YOLOv3, and YOLOv4 have been developed. Among these versions, YOLOv3 and YOLOv4 achieve a good compromise between speed and accuracy when facing the application requirements of traditional object detection, meeting the real-time application requirements while achieving good accuracy. However, when directly applying them to remote sensing image detection, there are problems such as low detection accuracy for targets with large scale differences, and often missed detection of small targets in complex scenes.

[0052] How to effectively design an efficient object detection algorithm with strong robustness and adaptability for the characteristics of remote sensing images is a key problem that urgently needs to be solved.

[0053] Next, in combination with Figures 1 to 9 Describe the object detection method and device for remote sensing images provided by the embodiments of the present invention.

[0054] Figure 1 It is a schematic flowchart of the object detection method for remote sensing images provided by the present invention. As Figure 1 shown, it includes but is not limited to the following steps:

[0055] First, in step S1, determine the target remote sensing image.

[0056] Specifically, select the initial remote sensing image to be recognized, perform denoising and image enhancement processing on the initial remote sensing image, and determine the processed target remote sensing image.

[0057] Furthermore, in step S2, input the target remote sensing image into the target detection model, and obtain the target detection result output by the target detection model; the target detection result includes the target type and target position in the target remote sensing image;

[0058] The target detection model is trained based on the sample remote sensing images, as well as the target type samples and target position samples in the sample remote sensing images. The target detection model is used to detect the target type and target position in the target remote sensing image;

[0059] The target detection model is constructed based on a dense feature pyramid network; the dense feature pyramid network includes an upsampling feature pyramid network and a downsampling feature pyramid network.

[0060] In order to make each detection result output contain the feature information of targets of different scales, compared with the traditional feature pyramid which only has the top-down feature fusion method, the dense feature pyramid network also adds bottom-up feature fusion and a way similar to the skip connection of the Densely connected convolutional Networks (DenseNet), which is more conducive to the propagation of gradients. Finally, each detection result input can deeply contain the target information of different scales, improving the detection performance of the target detection model for targets of different scales.

[0061] Based on the Dense-Feature Pyramid Networks (Dense-FPN), a target detection model is constructed to achieve the detection of multi-scale targets in images.

[0062] The target detection model can be obtained after being trained using remote sensing images with target category labels and target location labels.

[0063] In the target detection model, first, the feature extraction network extracts feature images of 4 sizes.

[0064] Furthermore, the 4 feature images are fused through the dense feature pyramid network. Finally, the 4 scale layers of the dense feature pyramid network output 4 fused feature maps after feature fusion, and each scale layer outputs one fused feature map.

[0065] Furthermore, through channel convolution on the 4 fused feature maps, 4 detection results are output.

[0066] Furthermore, the parameters in the 4 detection results all include prior box correction parameters and category parameters for target localization and classification. After prior box correction on the feature maps of the 4 detection results, the relative positions of the targets in the feature maps are obtained.

[0067] Furthermore, after mapping the relative positions back to the original image coordinates, non-maximum suppression is used to merge the overlapping results between different layers to obtain the final target detection result.

[0068] The detection results output by the target detection model include target categories and target positions. Among them, the target categories can include: targets such as airplanes and ships; the target positions can be the coordinate positions of each identified target on the target remote sensing image.

[0069] The target detection method for remote sensing images provided by the present invention can fuse the features of targets of different scales by constructing a target detection model based on the dense feature pyramid network, thereby improving the detection accuracy of targets of different scales.

[0070] Optionally, the determining of the target remote sensing image includes:

[0071] Obtain an initial remote sensing image;

[0072] Perform size normalization processing on the initial remote sensing image to determine the target remote sensing image.

[0073] Specifically, select the initial remote sensing image that needs to be detected. The size normalization processing method for the initial remote sensing image can use the nearest neighbor interpolation method or the bilinear interpolation method to obtain a target remote sensing image with a pixel size of 416×416.

[0074] The dense feature pyramid network based on YOLO proposed by the present invention, namely Dense-FPN-YOLO, can specifically solve the problems of complex scenes, large target scale differences, large field of view and small targets existing in large field-of-view optical remote sensing images.

[0075] Figure 2 is one of the structural schematic diagrams of the target detection model provided by the present invention, as Figure 2 shown, the target detection model includes a feature extraction network, a dense feature pyramid network and a detection network;

[0076] The inputting of the target remote sensing image into the target detection model to obtain the target detection result output by the target detection model includes:

[0077] Input the target remote sensing image into the feature extraction network to obtain feature images of multiple scales of the target remote sensing image output by the feature extraction network;

[0078] Input each feature image into the upsampling scale layer corresponding to the scale in the upsampling feature pyramid network to obtain the upsampling output feature output by each upsampling scale layer;

[0079] Input each upsampling output feature into the downsampling scale layer corresponding to the scale in the downsampling feature pyramid network to obtain a fused feature map output by each downsampling scale layer;

[0080] Input each fused feature map into the detection network to obtain the target detection result.

[0081] Figure 3 is the second structural schematic diagram of the target detection model provided by the present invention, Figure 4 is the structural schematic diagram of the dense pyramid network provided by the present invention, as Figures 2 to 4As shown, the dense feature pyramid network includes: a first scale layer, a second scale layer, and a third scale layer, and also includes a fourth scale layer (Scale 4 for small object detection).

[0082] The first scale layer includes a first convolutional block, a first connection block, a second connection block, and a third connection block connected in sequence. In the first scale layer, the first convolutional block includes: 5 convolutional blocks (CBL); the first connection block, the second connection block, and the third connection block each include: 5 CBL and 1 tensor concatenation layer (concat).

[0083] The second scale layer includes a fourth connection block, a fifth connection block, and a sixth connection block connected in sequence; in the second scale layer, the fourth connection block, the fifth connection block, and the sixth connection block each include: 5 CBL and 1 concat.

[0084] The third scale layer includes a seventh connection block, an eighth connection block, and a ninth connection block connected in sequence; in the third scale layer, the seventh connection block, the eighth connection block, and the ninth connection block each include: 5 CBL and 1 concat.

[0085] The fourth scale layer includes a tenth connection block, an eleventh connection block, a first connection layer, and a second convolutional block connected in sequence; in the fourth scale layer, the tenth connection block and the eleventh connection block each include: 5 CBL and 1 concat; the first connection layer includes 1 concat; the second convolutional block includes 5 CBL.

[0086] The input end of the first convolutional block serves as the input end of the first scale layer;

[0087] The output end of the third connection block serves as the output end of the first scale layer;

[0088] The output end of the first convolutional block is connected to the input end of the third connection block;

[0089] The output end of the first convolutional block is connected to the input end of the fourth connection block through a first upsampling module;

[0090] The input end of the fourth connection block serves as the input end of the second scale layer;

[0091] The output end of the sixth connection block serves as the output end of the second scale layer;

[0092] The input end of the second scale layer is connected to the input end of the seventh connection block through a second upsampling module;

[0093] The output end of the fourth connection block is connected to the input end of the eighth connection block through a third upsampling module;

[0094] The output end of the fourth connection block is connected to the input end of the first connection block through the first downsampling module;

[0095] The output end of the fourth connection block is also connected to the input end of the sixth connection block;

[0096] The output end of the fifth connection block is connected to the input end of the second connection block through the second downsampling module;

[0097] The output end of the sixth connection block is connected to the input end of the third connection block through the third downsampling module;

[0098] The input end of the seventh connection block serves as the input end of the third scale layer;

[0099] The output end of the ninth connection block serves as the output end of the third scale layer;

[0100] The output end of the eighth connection block is connected to the input end of the fifth connection block through the fourth downsampling module;

[0101] The output end of the ninth connection block is connected to the input end of the sixth connection block through the fifth downsampling module.

[0102] The input end of the tenth connection block serves as the input end of the fourth scale layer;

[0103] The output end of the second convolutional block serves as the output end of the fourth scale layer;

[0104] The input end of the third scale layer is connected to the input end of the tenth connection block through the fourth upsampling module;

[0105] The output end of the seventh connection block is connected to the input end of the eleventh connection block through the fifth upsampling module;

[0106] The output end of the eighth connection block is connected to the input end of the first connection layer through the sixth upsampling module;

[0107] The output end of the second convolutional block is connected to the input end of the ninth connection block through the sixth downsampling module.

[0108] The first upsampling module, the second upsampling module, the third upsampling module, the fourth upsampling module, the fifth upsampling module, and the sixth upsampling module all include: 1 CBL and 1 up sample.

[0109] The first downsampling module, the second downsampling module, the third downsampling module, the fourth downsampling module, the fifth downsampling module, and the sixth downsampling module all include: 1 down sample.

[0110] Optionally, the step of respectively inputting each feature image into the upsampling scale layer corresponding to the corresponding scale in the upsampling feature pyramid network to obtain the upsampling output features output by each upsampling scale layer includes:

[0111] Input a feature image of any scale into the corresponding scale layer of the upsampling feature pyramid network, and the corresponding scale layer of the upsampling feature pyramid network fuses the upsampling features of the feature image of any scale and the feature image of the previous scale of any scale, and the upsampling output features output by the scale layer of the previous scale, to obtain the upsampling output features output by the scale layer of any scale.

[0112] Optionally, the step of respectively inputting each upsampling output feature into the downsampling scale layer corresponding to the corresponding scale in the downsampling feature pyramid network to obtain the fused feature maps output by each downsampling scale layer includes:

[0113] Input the upsampling output feature of any scale into the corresponding scale layer of the downsampling feature pyramid network, and the corresponding scale layer of the downsampling feature pyramid network fuses the downsampling features of the feature image of any scale and the upsampling output feature of the next scale of any scale, and the fused feature maps output by the scale layer of the next scale, to obtain the fused feature maps output by the scale layer of any scale.

[0114] The object detection model includes a feature extraction network and a dense feature pyramid network. After the feature extraction network extracts features from the target remote sensing image, four feature images of different sizes are obtained.

[0115] Figure 5 is a schematic structural diagram of the feature pyramid network provided by the present invention. As Figure 5 shown, after the input image (Input Image), YOLOv3 originally only fuses the last three feature images to obtain the corresponding three detection results (predict). However, the network designed in this way has an unsatisfactory detection effect on small targets. The dense feature pyramid network provided by the present invention adds a fourth detection result on the basis of the original three detection results to fuse the feature information of small targets, so as to improve the detection performance of small targets.

[0116] YOLOv3 detects targets using different detection results for targets of different sizes. For an input image with a size of 416×416, the sizes of the three detection results are 13×13, 26×26, and 52×52 respectively. That is, the feature maps of the three detection results are downsampled 8 times, 16 times, and 32 times respectively. The smaller the size of the feature map, the larger the area corresponding to each grid unit in the input image. On the contrary, the larger the size of the feature map, the smaller the area corresponding to each grid unit in the input image. This means that the 13×13 detection result is suitable for detecting large targets, while the 52×52 detection result is suitable for detecting small targets. However, the 52×52 detection result is downsampled 8 times compared to the original image. That is, when the size of the target is less than 8×8, after feature extraction processing by the feature extraction network, the space it occupies in the feature map may be less than 1 pixel, which makes it difficult to detect small targets. Generally speaking, remote sensing images contain a large number of small targets. To further improve the detection performance of small targets in remote sensing images, a fourth detection result with a size of 104×104 is added to detect small targets, and the improved network structure adds a small target detection result of 104×104×255.

[0117] The detection network includes: a third convolutional block, a fourth convolutional block, a fifth convolutional block, and a sixth convolutional block;

[0118] The output end of the first scale layer is connected to the input end of the third convolutional block, and the output end of the third convolutional block is used to output the first detection result P5;

[0119] The output end of the second scale layer is connected to the input end of the fourth convolutional block, and the output end of the fourth convolutional block is used to output the second detection result P4;

[0120] The output end of the third scale layer is connected to the input end of the fifth convolutional block, and the output end of the fifth convolutional block is used to output the third detection result P3;

[0121] The output end of the fourth scale layer is connected to the input end of the sixth convolutional block, and the output end of the sixth convolutional block is used to output the fourth detection result P2.

[0122] Among them, the third convolutional block, the fourth convolutional block, the fifth convolutional block, and the sixth convolutional block all include: 1 CBL and 1 convolutional layer (Conv layer). Among them, CBL includes Conv, BN, and Leaky relu functions connected in sequence.

[0123] The parameters in each detection result include prior box correction parameters and class parameters to perform target localization and classification. Different detection results are used to detect targets of different sizes. For smaller targets, in order to improve the detection accuracy of small targets, a larger feature image with a resolution of 104×104 is added to detect small targets.

[0124] After adding the fourth detection result, Dense-FPN uses different detection results for targets of different sizes. The feature maps of the four detection results are downsampled 4 times, 8 times, 16 times, and 32 times respectively. The feature map downsampled 4 times is the newly added detection result for small targets. The smaller the size of the feature map, the larger the area corresponding to each grid unit in the input image. On the contrary, the larger the size of the feature map, the smaller the area corresponding to each grid unit in the input image. Generally speaking, remote sensing images contain a large number of small targets. To further improve the detection performance of small targets in remote sensing images, the detection result downsampled 4 times is used to detect small targets.

[0125] In the original YOLOv3, the structure of the feature pyramid network was used to horizontally fuse the semantic information before and after sampling. However, simple horizontal connections cannot well fuse the semantic information before sampling. Therefore, the dense feature pyramid network (Dense-FPN) proposed in the present invention, in order to solve the problem that small targets are difficult to detect in a large field of view, adds a larger fourth detection result on the basis of the three detection results of the original YOLOv3, making the dense feature pyramid network reach a depth of four layers. Through continuous upsampling and downsampling and feature fusion, the four finally generated detection results can deeply fuse the feature information of targets of different scales, so as to retain the feature information of smaller targets and improve the detection accuracy and detection ability for targets of different scales, especially small targets.

[0126] Since the algorithm ability of deep learning is closely related to the feature expressions extracted during its training, different from traditional optical images, the background of remote sensing image datasets is relatively complex. Therefore, in the field of processing optical remote sensing images, directly using a convolutional neural network for feature extraction has relatively poor effects. This patent improves the ability of the feature extraction network to extract features by adding an attention mechanism to the feature extraction network to enhance the network detection performance.

[0127] Optionally, the feature extraction network includes a plurality of sequentially connected residual modules; the inputting the target remote sensing image into the feature extraction network to obtain multiple-scale feature images of the target remote sensing image output by the feature extraction network includes:

[0128] Inputting the target remote sensing image into the feature extraction network to obtain the multiple-scale feature images output by multiple residual modules in the feature extraction network;

[0129] Each residual module in the feature extraction network includes an attention module.

[0130] In Figure 2 and Figure 3Among them, the feature extraction network includes a first residual module, a second residual module, a third residual module, a fourth residual module, and a fifth residual module connected in sequence;

[0131] The output end of the first residual module is connected to the input end of the first connection layer;

[0132] The output end of the second residual module is simultaneously connected to the input end of the fourth scale layer and the input end of the eighth connection block;

[0133] The output end of the third residual module is connected to the input end of the third scale layer;

[0134] The output end of the fourth residual module is connected to the input end of the second scale layer;

[0135] The output end of the fifth residual module is connected to the input end of the first scale layer.

[0136] In the feature extraction network, the first residual module includes: 1 CBL and 1 residual unit (SGERes1); the second residual module includes: 1 residual unit (SGERes2); the third residual module includes: 1 residual unit (SGERes8); the fourth residual module includes: 1 residual unit (SGERes8); the fifth residual module includes: 1 residual unit (SGERes4). The attention mechanism (Spatial Group-wise Enhance, SGE) module is a lightweight attention module, which can greatly improve the classification and detection performance with almost no increase in the number of parameters and computational complexity.

[0137] As Figure 2 shown, the SGEResX residual module includes a CBL and X SGERes units connected in sequence.

[0138] The SGERes unit includes a CBL and a CBSL connected in sequence.

[0139] The CBSL includes a combination of Conv, normalization layer (BN), SGE, and Leaky relu functions connected in sequence.

[0140] Figure 6 is the structural schematic diagram of the residual unit based on SGE attention provided by the present invention. As Figure 6 shown, 1 SGEResN residual module includes a CBL, and N combinations of CBL, Conv, normalization layer (BN), SGE, and Leaky relu functions connected in sequence.

[0141] The feature extraction network is constructed based on the backbone network Darknet53 and consists of a large number of residual units. Thanks to these residual structures, Darknet can be effectively trained even when stacked to 53 layers, without the problems of gradient explosion or gradient disappearance. However, the hierarchical connections in a single residual module only capture detailed information in the receptive field and lose the global characteristics, resulting in insufficient and ineffective extraction of features in each layer in complex scenarios.

[0142] The background of optical remote sensing images is complex and the target features are not obvious, which will lead to a low detection accuracy of the convolutional neural network. To improve the ability of the network to extract effective features in complex scenarios, the present invention adds an SGE module to the residual unit. Due to the lightweight of the SGE module and its effectiveness for high-order semantic features, the SGE module can perfectly fit with Darknet53. By improving the feature extraction ability of the feature extraction network, the network detection performance is enhanced, enabling the feature extraction network to quickly extract the key feature information of the target in complex scenarios.

[0143] Figure 7 It is a schematic diagram of the structure of the SGE attention module provided by the present invention. As Figure 7 shown, since a complete feature is composed of many sub-features, and these sub-features are distributed in the features of each layer in groups, but these sub-features will be processed in the same way and are affected by background noise, which will lead to incorrect recognition and localization results. Therefore, the addition of the SGE module can generate attention factors in each group, so as to obtain the importance of each sub-feature, and each feature group can also specifically learn and suppress noise. The specific steps are as follows:

[0144] Divide the feature map into G groups according to the channel dimension; perform attention learning on each group separately; perform global average pooling on the group to obtain the vector g; perform element-wise dot product of the pooled g and the original feature group; after normalization, use the Sigmoid function (activation to assign weights) for activation; finally, perform element-wise dot product with the original feature group.

[0145] In Figure 2 , the feature map obtained after continuous convolution of the original target remote sensing image passes through the SGE module according to the channel dimension, obtains the attention factors of each group of features and maps them to the positions of the corresponding feature maps, and finally outputs the feature image with enhanced semantic features.

[0146] After extracting features through the backbone network SGEDarknet53 and designing the fourth-scale layer for small object detection, Dense-FPN is required to continuously sample and fuse the feature images of different objects at four scales of the C2, C3, C4, and C5 layers. Among them, the feature images of each layer of C3, C4, and C5 need to be upsampled and fused with the previous layer. The fused feature maps are then upsampled and fused with the previous layer until reaching the top layer C2, thus generating the intermediate hidden layers H2, H3, H4, and H5.

[0147] Furthermore, the feature maps of each layer of the intermediate hidden layers H2, H3, and H4 need to be downsampled and fused with the next layer. The fused feature maps are then downsampled and fused with the next layer until reaching the bottom layer H5, obtaining 4 fused feature maps. Then, channel convolution is performed on the 4 fused feature maps to generate 4 detection results, namely P2, P3, P4, and P5. And through skip connections, all layers are merged in channels. Finally, the detection results are generated based on the detection results of the last four layers, realizing feature reuse. This connection method is more conducive to the backpropagation of gradients, can better utilize feature information, and improve the transfer efficiency of information between layers.

[0148] The K-means algorithm can be used to generate corresponding anchor boxes for the 4 detection results. The anchor boxes generated by the K-means algorithm and the labeled boxes have a large intersection over union, which is more conducive to the convergence of the network. The steps are as follows.

[0149] Step 1, select k samples from the data as the initial clustering centers: (W i , H i ), i ∈ {1, 2, … k}, where (W i , H i ) represents the width and height of the anchor box;

[0150] Step 2, calculate the distance from each ground truth box to the clustering center;

[0151] Step 3, calculate each clustering center (W i ’, H i ’), i ∈ {1, 2, … k} specifically as follows:

[0152] Step 4, repeat Step 2 to Step 3 until the clustering converges to obtain the anchor boxes.

[0153] Using the K-means algorithm, a total of 12 anchor boxes of 4 different scales are finally generated for the four detection results: (21, 25), (25, 31), (33, 39), (44, 51), (59, 81), (84, 95), (104, 116), (119, 148), (161, 184), (221, 201), (246, 213), (259, 278).

[0154] Among them, the three anchor boxes (21, 25), (25, 31), and (33, 39) are designed for the fourth detection result of the added size of 104×104. They can be used to detect small targets in remote sensing images, such as helicopters, cars, etc. Usually, these small targets are only a few pixels in size.

[0155] Furthermore, through the correction of the preset prior boxes, the detection result outputs the prediction result. Specifically, the detection result is predicted using prior boxes of 4 different scales. According to the object inclusion situation of the prior boxes, the forward propagation and backward propagation of the neural network will correct the coordinate parameters of the prior boxes and output the corrected result. Its final data result is the corrected anchor box coordinates. This operation offsets the center point and scales the width and height of the generated prior anchor boxes so that the generated prior anchor boxes can locate the targets in the input image.

[0156] Furthermore, after mapping the corrected result coordinates of the detection results of different scales back to the original image coordinates, multiple prediction bounding boxes of the same target are generated. Therefore, non-maximum suppression is used to merge the overlapping results between different layers to obtain the final target detection result of the original large field-of-view optical remote sensing image.

[0157] Furthermore, the detection result is output on the four-layer detection results after fusion, and non-maximum suppression is used for merging to obtain the final target detection result.

[0158] According to the object detection method for remote sensing images provided by the present invention, the object detection model has strong robustness and adaptability.

[0159] Figure 8 It is a schematic structural diagram of the object detection device for remote sensing images provided by the present invention, as Figure 8 shown, including:

[0160] A determination unit 801, configured to determine a target remote sensing image;

[0161] An acquisition unit 802, configured to input the target remote sensing image into an object detection model and acquire the object detection result output by the object detection model; the object detection result includes the object type and object position in the target remote sensing image;

[0162] The object detection model is trained based on a sample remote sensing image, as well as an object type sample and an object position sample in the sample remote sensing image. The object detection model is used to detect the object type and object position in the target remote sensing image;

[0163] The target detection model is constructed based on a dense feature pyramid network; the dense feature pyramid network includes an upsampling feature pyramid network and a downsampling feature pyramid network.

[0164] First, the determination unit 801 determines the target remote sensing image.

[0165] Specifically, an initial remote sensing image to be recognized is selected, and the initial remote sensing image is denoised and image-enhanced to determine the processed target remote sensing image.

[0166] Furthermore, the acquisition unit 802 inputs the target remote sensing image into the target detection model to obtain the target detection result output by the target detection model; the target detection result includes the target category and target location in the target remote sensing image; the target detection model is trained based on sample remote sensing images, as well as target category samples and target location samples in the sample remote sensing images, and the target detection model is used to detect the target category and target location in the target remote sensing image; the target detection model is constructed based on a dense feature pyramid network; the dense feature pyramid network includes an upsampling feature pyramid network and a downsampling feature pyramid network.

[0167] In order to make each output detection result contain the feature information of targets at different scales, compared with the traditional feature pyramid network which only has the top-down feature fusion method, the dense feature pyramid network also adds bottom-up feature fusion and a way similar to the DenseNet skip connection, which is more conducive to the propagation of gradients. Finally, each input detection result can deeply contain the target information at different scales, improving the detection performance of the target detection model for targets at different scales.

[0168] Construct a target detection model based on Dense-FPN to achieve the detection of multi-scale targets in an image.

[0169] The target detection model can be obtained by training with remote sensing images with target category labels and target location labels.

[0170] In the target detection model, first, the feature extraction network extracts feature images of 4 sizes.

[0171] Furthermore, the 4 feature images are fused through the dense feature pyramid network. Finally, the 4 scale layers of the dense feature pyramid network output 4 fused feature maps after feature fusion, and each scale layer outputs one fused feature map.

[0172] Furthermore, by performing channel convolution on the 4 fused feature maps, 4 detection results are output.

[0173] Furthermore, the parameters in the four detection results all include prior box correction parameters and class parameters for target localization and classification. After correcting the prior boxes for the feature maps of the four detection results, the relative positions of the targets in the feature maps are obtained.

[0174] Furthermore, after mapping the relative positions back to the original image coordinates, non-maximum suppression is used to merge the overlapping results between different layers to obtain the final target detection result.

[0175] The detection results output by the target detection model include the target category and the target location. The target category can include: targets such as airplanes and ships; the target location can be the coordinate position of each identified target on the target remote sensing image.

[0176] The target detection device for remote sensing images provided by the present invention can fuse the features of targets at different scales by constructing a target detection model based on a dense feature pyramid network, thereby improving the detection accuracy of targets at different scales.

[0177] Optionally, the target detection device for remote sensing images further includes: a normalization module for performing size normalization processing on the initial remote sensing image to determine the target remote sensing image.

[0178] Specifically, the normalization module selects the initial remote sensing image for which target detection is to be performed. The size normalization processing method for the initial remote sensing image can use the nearest neighbor interpolation method or the bilinear interpolation method to obtain a target remote sensing image with a pixel size of 416×416.

[0179] It should be noted that the target detection device for remote sensing images provided in the embodiments of the present invention can be implemented based on the target detection method for remote sensing images described in any of the above embodiments when specifically executed, and this embodiment will not be elaborated herein.

[0180] Figure 9 is a schematic structural diagram of the electronic device provided by the present invention, as Figure 9As shown, the electronic device may include: a processor 910, a communications interface 920, a memory 930, and a communication bus 940. Among them, the processor 910, the communication interface 920, and the memory 930 complete communication with each other through the communication bus 940. The processor 910 may call the logical instructions in the memory 930 to execute a target detection method for remote sensing images. The method includes: determining a target remote sensing image; inputting the target remote sensing image into a target detection model to obtain a target detection result output by the target detection model; the target detection result includes the target type and target location in the target remote sensing image; the target detection model is trained based on sample remote sensing images, as well as target type samples and target location samples in the sample remote sensing images, and the target detection model is used to detect the target type and target location in the target remote sensing image; the target detection model is constructed based on a dense feature pyramid network; the dense feature pyramid network includes an upsampling feature pyramid network and a downsampling feature pyramid network.

[0181] In addition, when the logical instructions in the above-mentioned memory 930 can be implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to enable a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0182] On the other hand, the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the object detection method for remote sensing images provided by the above-mentioned various methods. The method includes: determining a target remote sensing image; inputting the target remote sensing image into a target detection model to obtain a target detection result output by the target detection model; the target detection result includes the target type and target location in the target remote sensing image; the target detection model is trained based on sample remote sensing images, as well as target type samples and target location samples in the sample remote sensing images. The target detection model is used to detect the target type and target location in the target remote sensing image; the target detection model is constructed based on a dense feature pyramid network; the dense feature pyramid network includes an upsampling feature pyramid network and a downsampling feature pyramid network.

[0183] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it realizes the object detection method for remote sensing images provided by the above-mentioned various embodiments. The method includes: determining a target remote sensing image; inputting the target remote sensing image into a target detection model to obtain a target detection result output by the target detection model; the target detection result includes the target type and target location in the target remote sensing image; the target detection model is trained based on sample remote sensing images, as well as target type samples and target location samples in the sample remote sensing images. The target detection model is used to detect the target type and target location in the target remote sensing image; the target detection model is constructed based on a dense feature pyramid network; the dense feature pyramid network includes an upsampling feature pyramid network and a downsampling feature pyramid network.

[0184] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.

[0185] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0186] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features. And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A target detection method for remote sensing images, characterized in that, Including: Determine the target remote sensing image; Input the target remote sensing image into the target detection model to obtain the target detection result output by the target detection model; The target detection result includes the target type and target location in the target remote sensing image; The target detection model is trained based on the sample remote sensing image, as well as the target type samples and target location samples in the sample remote sensing image, and is used to detect the target type and target location in the target remote sensing image; The target detection model is constructed based on the dense feature pyramid network; the dense feature pyramid network includes an upsampling feature pyramid network and a downsampling feature pyramid network; The dense feature pyramid network includes: a first-scale layer, a second-scale layer, a third-scale layer, and further includes a fourth-scale layer; the first-scale layer includes a first convolutional block, a first connection block, a second connection block, and a third connection block connected in sequence; in the first-scale layer, the first convolutional block includes: 5 convolutional blocks CBL; the first connection block, the second connection block, and the third connection block all include: 5 CBL and 1 tensor concatenation layer concat; the second-scale layer includes a fourth connection block, a fifth connection block, and a sixth connection block connected in sequence; in the second-scale layer, the fourth connection block, the fifth connection block, and the sixth connection block all include: 5 CBL and 1 concat; the third-scale layer includes a seventh connection block, an eighth connection block, and a ninth connection block connected in sequence; in the third-scale layer, the seventh connection block, the eighth connection block, and the ninth connection block all include: 5 CBL and 1 concat; the fourth-scale layer includes a tenth connection block, an eleventh connection block, a first connection layer, and a second convolutional block connected in sequence; in the fourth-scale layer, the tenth connection block and the eleventh connection block both include: 5 CBL and 1 concat; the first connection layer includes 1 concat; the second convolutional block includes 5 CBL; the input end of the first convolutional block serves as the input end of the first-scale layer; the output end of the third connection block serves as the output end of the first-scale layer; the output end of the first convolutional block is connected to the input end of the third connection block; the output end of the first convolutional block is connected to the input end of the fourth connection block through a first upsampling module; the input end of the fourth connection block serves as the input end of the second-scale layer; the output end of the sixth connection block serves as the output end of the second-scale layer; the input end of the second-scale layer is connected to the input end of the seventh connection block through a second upsampling module; the output end of the fourth connection block is connected to the input end of the eighth connection block through a third upsampling module; the output end of the fourth connection block is connected to the input end of the first connection block through a first downsampling module; the output end of the fourth connection block is further connected to the input end of the sixth connection block; the output end of the fifth connection block is connected to the input end of the second connection block through a second downsampling module; the output end of the sixth connection block is connected to the input end of the third connection block through a third downsampling module; the input end of the seventh connection block serves as the input end of the third-scale layer; the output end of the ninth connection block serves as the output end of the third-scale layer; the output end of the eighth connection block is connected to the input end of the fifth connection block through a fourth downsampling module; the output end of the ninth connection block is connected to the input end of the sixth connection block through a fifth downsampling module; the input end of the tenth connection block serves as the input end of the fourth-scale layer; the output end of the second convolutional block serves as the output end of the fourth-scale layer; the input end of the third-scale layer is connected to the input end of the tenth connection block through a fourth upsampling module; the output end of the seventh connection block is connected to the input end of the eleventh connection block through a fifth upsampling module;The output end of the eighth connection block is connected to the input end of the first connection layer through the sixth upsampling module; the output end of the second convolution block is connected to the input end of the ninth connection block through the sixth downsampling module; the first upsampling module, the second upsampling module, the third upsampling module, the fourth upsampling module, the fifth upsampling module, and the sixth upsampling module each include: 1 CBL and 1 upsampling up sample; the first downsampling module, the second downsampling module, the third downsampling module, the fourth downsampling module, the fifth downsampling module, and the sixth downsampling module each include: 1 downsampling downsample.; 2. The object detection method for remote sensing images according to claim 1, wherein The target detection model includes a feature extraction network, the dense feature pyramid network, and a detection network; The step of inputting the target remote sensing image into the target detection model to obtain the target detection result output by the target detection model includes: Input the target remote sensing image into the feature extraction network to obtain feature images of multiple scales of the target remote sensing image output by the feature extraction network; Input each feature image into the upsampling scale layer corresponding to the scale in the upsampling feature pyramid network to obtain the upsampling output features output by each upsampling scale layer; Input each upsampling output feature into the downsampling scale layer corresponding to the scale in the downsampling feature pyramid network to obtain the fused feature maps output by each downsampling scale layer; Input each fused feature map into the detection network to obtain the target detection result output by the detection network.

3. The object detection method for remote sensing images according to claim 2, characterized in that, The step of inputting each feature image into the upsampling scale layer corresponding to the scale in the upsampling feature pyramid network to obtain the upsampling output features output by each upsampling scale layer includes: Input the feature image of any scale into the corresponding scale layer of the upsampling feature pyramid network, and the corresponding scale layer of the upsampling feature pyramid network fuses the feature image of any scale with the upsampling feature of the feature image of the previous scale of any scale, as well as the upsampling output features output by the scale layer of the previous scale, to obtain the upsampling output features output by the scale layer of any scale.

4. The object detection method for remote sensing images according to claim 2, wherein, The step of inputting each upsampling output feature into the downsampling scale layer corresponding to the scale in the downsampling feature pyramid network to obtain the fused feature maps output by each downsampling scale layer includes: Input the upsampling output feature of any scale into the corresponding scale layer of the downsampling feature pyramid network, and the corresponding scale layer of the downsampling feature pyramid network fuses the feature image of any scale with the downsampling feature of the upsampling output feature of the next scale of any scale, as well as the fused feature maps output by the scale layer of the next scale, to obtain the fused feature maps output by the scale layer of any scale.

5. The object detection method for remote sensing images according to claim 2, wherein The feature extraction network includes a plurality of sequentially connected residual modules; the step of inputting the target remote sensing image into the feature extraction network to obtain the feature images of multiple scales of the target remote sensing image output by the feature extraction network includes: Input the target remote sensing image into the feature extraction network to obtain the feature images of multiple scales output by multiple residual modules in the feature extraction network; Each residual module in the feature extraction network includes an attention module.

6. The object detection method for remote sensing images according to any one of claims 1 to 5, characterized in that, The determination of the target remote sensing image includes: Obtain the initial remote sensing image; Perform size normalization processing on the initial remote sensing image to determine the target remote sensing image.

7. An object detection device for remote sensing images, characterized in that, It includes: A determination unit for determining the target remote sensing image; An acquisition unit for inputting the target remote sensing image into the target detection model to obtain the target detection result output by the target detection model; the target detection result includes the target type and target location in the target remote sensing image; The target detection model is trained based on the sample remote sensing image, as well as the target type samples and target location samples in the sample remote sensing image, and the target detection model is used to detect the target type and target location in the target remote sensing image; The target detection model is constructed based on the dense feature pyramid network; the dense feature pyramid network includes an upsampling feature pyramid network and a downsampling feature pyramid network; The dense feature pyramid network includes: a first-scale layer, a second-scale layer, and a third-scale layer, and further includes a fourth-scale layer; the first-scale layer includes a first convolutional block, a first connection block, a second connection block, and a third connection block connected in sequence; in the first-scale layer, the first convolutional block includes: 5 convolutional blocks CBL; the first connection block, the second connection block, and the third connection block each include: 5 CBL and 1 tensor concatenation layer concat; the second-scale layer includes a fourth connection block, a fifth connection block, and a sixth connection block connected in sequence; in the second-scale layer, the fourth connection block, the fifth connection block, and the sixth connection block each include: 5 CBL and 1 concat; the third-scale layer includes a seventh connection block, an eighth connection block, and a ninth connection block connected in sequence; in the third-scale layer, the seventh connection block, the eighth connection block, and the ninth connection block each include: 5 CBL and 1 concat; the fourth-scale layer includes a tenth connection block, an eleventh connection block, a first connection layer, and a second convolutional block connected in sequence; in the fourth-scale layer, the tenth connection block and the eleventh connection block each include: 5 CBL and 1 concat; the first connection layer includes 1 concat; the second convolutional block includes 5 CBL; the input end of the first convolutional block serves as the input end of the first-scale layer; the output end of the third connection block serves as the output end of the first-scale layer; the output end of the first convolutional block is connected to the input end of the third connection block; the output end of the first convolutional block is connected to the input end of the fourth connection block through a first upsampling module; the input end of the fourth connection block serves as the input end of the second-scale layer; the output end of the sixth connection block serves as the output end of the second-scale layer; the input end of the second-scale layer is connected to the input end of the seventh connection block through a second upsampling module; the output end of the fourth connection block is connected to the input end of the eighth connection block through a third upsampling module; the output end of the fourth connection block is connected to the input end of the first connection block through a first downsampling module; the output end of the fourth connection block is also connected to the input end of the sixth connection block; the output end of the fifth connection block is connected to the input end of the second connection block through a second downsampling module; the output end of the sixth connection block is connected to the input end of the third connection block through a third downsampling module; the input end of the seventh connection block serves as the input end of the third-scale layer; the output end of the ninth connection block serves as the output end of the third-scale layer; the output end of the eighth connection block is connected to the input end of the fifth connection block through a fourth downsampling module; the output end of the ninth connection block is connected to the input end of the sixth connection block through a fifth downsampling module; the input end of the tenth connection block serves as the input end of the fourth-scale layer; the output end of the second convolutional block serves as the output end of the fourth-scale layer; the input end of the third-scale layer is connected to the input end of the tenth connection block through a fourth upsampling module; the output end of the seventh connection block is connected to the input end of the eleventh connection block through a fifth upsampling module;The output end of the eighth connection block is connected to the input end of the first connection layer through the sixth upsampling module; the output end of the second convolutional block is connected to the input end of the ninth connection block through the sixth downsampling module; the first upsampling module, the second upsampling module, the third upsampling module, the fourth upsampling module, the fifth upsampling module and the sixth upsampling module all include: 1 CBL and 1 upsampling up sample; the first downsampling module, the second downsampling module, the third downsampling module, the fourth downsampling module, the fifth downsampling module and the sixth downsampling module all include: 1 downsampling downsample.; 8. The target detection device for remote sensing images according to claim 7, characterized in that, It further includes: A normalization module for performing size normalization processing on the initial remote sensing image to determine the target remote sensing image.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the target detection method for remote sensing images according to any one of claims 1 to 6.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the target detection method for remote sensing images according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Multi-angle remote sensing ship image target detection method based on feature pyramid structure

    CN111753677A

  • Feature weaving method for remote sensing image target detection

    CN113111740A