Remote sensing image target detection method based on improved pyramid

By adopting the dual-stage feature pyramid structure based on ResNet50 and the attention extraction method of the enhancement module in the remote sensing image object detection, the problem of low detection accuracy of remote sensing image object in the prior art is solved, and higher detection accuracy and accurate recognition of edge features are achieved.

CN119919737AActive Publication Date: 2025-05-02NORTHWESTERN POLYTECHNICAL UNIV

Patent Information

Application Number
CN202510150057.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-02
Estimated Expiration
2045-02-11

AI Technical Summary

Technical Problem

When detecting high-resolution remote sensing images, existing pyramid networks are difficult to accurately identify the edges of dense targets, resulting in lower detection accuracy.

Method used

Feature extraction is performed using a two-stage feature pyramid structure based on ResNet50, and channel attention extraction and spatial attention extraction are performed on high-level and low-level feature maps through two enhancement modules, and a double-scale feature map is generated and input to the detection head for object detection.

Benefits of technology

While retaining high-level semantic information, it captures the underlying detailed information more comprehensively, improves the detection accuracy of dense small targets and can accurately identify the edges of objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119919737A_ABST
    Figure CN119919737A_ABST
Patent Text Reader

Abstract

The invention provides a remote sensing image target detection method based on an improved pyramid, and the method comprises the steps: in a constructed detection model, a feature extraction module is a two-stage feature pyramid structure based on ResNet50, and a layer of pyramid feature map fusion structure from bottom to top is added; the two enhancement modules respectively take the two feature maps of the high-level size and the two feature maps of the low-level size as input, channel attention extraction and space attention extraction are carried out, and the extracted feature maps are multiplied to obtain a dual-scale feature map; and then the two detection head modules respectively process the dual-scale feature map, and respectively output the category and bounding box information of the target through a classification branch and a regression branch. According to the method, rich semantic information of a high layer is reserved, detail information such as edges, corners and textures of a bottom layer is captured more comprehensively, each edge of an object can be accurately identified for a remote sensing target under dense and overlapping conditions, and higher detection precision is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of image processing, and in particular relates to a remote sensing image target detection method based on an improved pyramid. Background Art

[0002] In recent years, with the increasing number of space imaging devices, the spatial resolution of remote sensing images has been continuously improved. Compared with traditional medium and low-resolution images, high-resolution remote sensing images can more clearly express the characteristic distribution and spatial correlation of ground objects. The large amount of texture, color and other details they contain provide new information channels and intelligence sources, which facilitates the refined integrated management and decision-making of the region.

[0003] Pyramid network is a method of image processing by decomposing image scales. Specifically, the input image is divided into multiple different scales and then processed separately. In tasks such as image classification and target detection, pyramid network can improve the performance and robustness of the model, so it is widely used. However, high-resolution remote sensing images have a long imaging distance, and the target size of remote sensing images is small compared to the targets in natural scene images, so higher requirements are also placed on the accuracy of the detection method. The existing pyramid network only retains the rich semantic information in the high layer. For remote sensing images, there are often dense targets in the image. In this case, it cannot accurately identify the edges of the object, and the detection accuracy is low. Summary of the invention

[0004] The purpose of the present invention is to solve the problem of low target detection accuracy caused by using pyramid network to detect remote sensing images in the prior art, and to provide a remote sensing image target detection method based on improved pyramid. This method not only retains the rich semantic information of the high-level, but also captures the underlying detail information more comprehensively. For dense remote sensing targets, the object features can be highlighted, and dense small targets can be better detected, thereby improving the detection accuracy.

[0005] To achieve the above purpose, the technical solution provided by the present invention is:

[0006] A remote sensing image target detection method based on an improved pyramid is provided, comprising the following steps:

[0007] Step 1, obtaining a data set of remote sensing images;

[0008] Step 2, constructing a remote sensing image target detection model based on an improved pyramid, the target detection model includes a feature extraction module, two enhancement modules and two detection head modules;

[0009] The feature extraction module is a two-level feature pyramid structure based on ResNet50. The first-level pyramid structure is used to extract four feature maps of different sizes from the input image and obtain a fused feature map through the FPN structure. The second-level pyramid structure is used to perform feature fusion on the obtained fused feature map from bottom to top.

[0010] The two enhancement modules take the two high-level feature maps and the two low-level feature maps of the four feature maps obtained by fusion as input, perform channel attention extraction and spatial attention extraction on the input feature maps, and multiply the channel attention feature map with the spatial attention feature map to obtain the final dual-scale feature map;

[0011] The two detection head modules are used to process the dual-scale feature maps and output the target category and bounding box information through the classification branch and regression branch respectively;

[0012] Step 3, iteratively train the constructed target detection model using the acquired data set, and obtain the optimal network model parameters after the detection model converges;

[0013] Step 4: Use the obtained optimal network model parameters to detect the target remote sensing image and obtain the classification and positioning information of the target in the image.

[0014] Furthermore, in step 1, the acquired dataset is a public dataset or a self-photographed dataset.

[0015] Furthermore, step 1 also includes: dividing the acquired data into a training set and a test set, with a division ratio of 8:2 or 7:3.

[0016] Furthermore, in step 2, the sizes of the images extracted by the feature extraction module are 1 / 4, 1 / 8, 1 / 16 and 1 / 32 of the input image size respectively.

[0017] Furthermore, in step 2, the two enhancement modules are respectively a first enhancement fusion module and a second enhancement fusion module. The input of the first enhancement fusion module is two of the four feature maps output by the feature extraction module, whose sizes are 1 / 16 and 1 / 32 of the input image size. The input of the second enhancement fusion module is two of the four feature maps output by the feature extraction module, whose sizes are 1 / 8 and 1 / 4 of the input image size.

[0018] Furthermore, in step 2, the processing process of each enhancement module on the input feature map is as follows: keep the image size consistent by upsampling or downsampling, and then perform a splicing operation; perform a convolution operation on the spliced ​​feature map; perform channel attention extraction and spatial attention extraction on the feature map obtained by convolution respectively; and multiply the channel attention feature map with the spatial attention feature map.

[0019] Furthermore, in step 2, the spatial attention extraction process of each enhancement module is as follows: the feature map is subjected to global maximum pooling and global average pooling operations along the channel direction to obtain two feature maps, which are then subjected to element-wise operations after convolution operations, and finally the final spatial attention feature map is generated through sigmoid activation operations.

[0020] Furthermore, in step 2, the channel attention extraction process of each enhancement module is as follows: the feature map is subjected to global maximum pooling and global average pooling operations along the height and width directions to obtain two feature maps respectively, and then subjected to element-wise operations after convolution operations, and finally a sigmoid activation operation is performed to generate the final channel attention feature map.

[0021] Furthermore, in step 2, the detection head module adopts an SSD detection head, and the loss function includes category loss and positioning loss.

[0022] The advantages of the present invention are:

[0023] 1. The remote sensing image target detection method based on the improved pyramid of the present invention adopts a two-level feature pyramid structure based on ResNet50 for feature extraction, and adopts two enhancement modules to process the feature maps of high-level size and low-level size respectively. Compared with the traditional pyramid network, while retaining the rich semantic information of the high-level, it more comprehensively captures the edge, corner, texture and other detail information of the bottom layer. For remote sensing targets in dense and overlapping situations, the edges of the objects can be accurately identified, and the detection accuracy is higher.

[0024] 2. The present invention adds the steps of extracting spatial attention and channel attention after splicing two adjacent feature maps after extraction, comprehensively captures the key information in the features, and provides a more comprehensive and effective feature extraction capability for the target detection model.

[0025] 3. The present invention addresses the difficulty of large scale differences in remote sensing targets. The enhancement module outputs a dual-scale feature map that enters the detection head for detection. The dual-scale feature map retains the top-level and bottom-level key information in the feature extraction module respectively, taking into account both large and small targets in the image, and to a certain extent solves the problem of missed detection of small targets in the multi-scale target feature characterization problem. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The above and / or other features and advantages of the present invention will become more easily understood through the following description with reference to the accompanying drawings, in which:

[0027] Figure 1 It is a flow chart of a remote sensing image target detection method based on an improved pyramid of the present invention;

[0028] Figure 2 It is the overall structure diagram of the remote sensing image target detection model in the present invention;

[0029] FIG3(a) is a schematic diagram of the structure of the first enhanced fusion module, and FIG3(b) is a schematic diagram of the structure of the first enhanced fusion module;

[0030] Figure 4 is a remote sensing image diagram in an example of the present invention;

[0031] Figure 5 It is a diagram of the detection results of the present invention in an example. DETAILED DESCRIPTION

[0032] The present invention will be described in detail below with reference to the accompanying drawings by means of exemplary embodiments of the present invention. It should be noted that the following detailed description of the present invention is only for the purpose of illustration, and is not intended to limit the present invention.

[0033] The present invention provides a remote sensing image target detection method based on an improved pyramid, which can effectively extract a feature map of a target and improve the accuracy of remote sensing image target detection.

[0034] First, the overall reference Figure 1 As an exemplary embodiment of the present invention, a remote sensing image target detection method based on an improved pyramid comprises the following steps:

[0035] Step S1, obtaining a data set of remote sensing images;

[0036] Step S2, constructing a remote sensing image target detection model based on an improved pyramid, the target detection model includes a feature extraction module, two enhancement modules and two detection head modules;

[0037] The feature extraction module is a two-level feature pyramid structure based on ResNet50. The first-level pyramid structure is used to extract four feature maps of different sizes from the input image and obtain a fused feature map through the FPN structure. The second-level pyramid structure is used to perform feature fusion on the obtained fused feature map from bottom to top.

[0038] The two enhancement modules take the two high-level feature maps and the two low-level feature maps of the four feature maps obtained by fusion as input, perform channel attention extraction and spatial attention extraction on the input feature maps, and multiply the channel attention feature map with the spatial attention feature map to obtain the final dual-scale feature map;

[0039] The two detection head modules are used to process the dual-scale feature maps and output the target category and bounding box information through the classification branch and regression branch respectively;

[0040] Step S3, iteratively training the constructed target detection model using the acquired data set, and obtaining the optimal network model parameters after the detection model converges;

[0041] Step S4, using the obtained optimal network model parameters to detect the target remote sensing image, and obtain the classification and positioning information of the target in the image.

[0042] In step S1, the remote sensing dataset obtained can be a public dataset, including optical remote sensing datasets such as DIOR and NWPU VHR-10; it can also be a self-photographed dataset. If a self-photographed dataset is used, the image needs to be size normalized first, and then the normalized image needs to be color normalized. In order to prevent overfitting, the number of self-photographed training set images should be no less than 800.

[0043] In addition, the acquired data can be divided into a training set and a test set, and the ratio of the training set to the test set can be 8:2 or 7:3. Accordingly, in step S3, the processed training set data can be input into the model for iterative training, and the optimal network model parameters are obtained after the detection model converges, and then the test set data is detected using the obtained optimal network model parameters.

[0044] For step S2, Figure 2 The overall structure of the remote sensing image target detection model is shown. In order to effectively solve the gradient disappearance and model saturation problems in the network model, the feature extraction module uses ResNet-50 as the backbone to extract features from the image. A total of 4 layers of feature maps are extracted, and four feature maps C1, C2, C3, and C4 are obtained from bottom to top. The image size can be 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the input image size, respectively. The number of channels is 64, 128, 256, and 512, respectively. After obtaining feature maps of different sizes, the structure of FPN (pyramid network) is used to perform upsampling and downsampling operations to make the parts that need to be fused have the same size. Through lateral connection and fusion operations, the fused feature maps P1, P2, and P3 are obtained. , P4 image sizes are 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the input image size, and the number of channels are 64, 128, 256, and 512, respectively; the traditional FPN network retains rich semantic information in the high layer, but for remote sensing images, there are often dense targets in the images. In this case, it is necessary to accurately identify the edges of the objects, and a large amount of edge details are retained in the low-level network. In order to retain more detail information and ensure the detection accuracy of dense small targets in remote sensing images, the present invention continues to add a bottom-to-top pyramid feature map fusion structure after the above process, and the output 4 feature maps are M1, M2, M3, and M4, which are consistent with the size and channel number of the feature maps P1, P2, P3, and P4 in the previous step.

[0045] There are two enhancement modules after the feature extraction module, which can be the first enhancement fusion module and the second enhancement fusion module. Referring to Figure 3(a), the input of the first enhancement fusion module, that is, the enhancement module 1, is two of the four feature maps output by the feature extraction module, whose sizes are 1 / 16 and 1 / 32 of the input image size; referring to Figure 3(b), the input of the second enhancement fusion module, that is, the enhancement module 2, is two of the four feature maps output by the feature extraction module, whose sizes are 1 / 8 and 1 / 4 of the input image size. The input feature map of each enhancement module is first upsampled or downsampled to keep the image size consistent, and then the splicing operation is performed. When the feature map input to enhancement module 1 is unified in size, the high-level size is mainly used, that is, the size of the two feature maps after sampling is 1 / 32 of the input image. When the feature map input to enhancement module 2 is unified in size, the low-level size is mainly used, that is, the size of the two feature maps after sampling is 1 / 4 of the input image. Assuming that the size of the spliced ​​feature map is H×W×2C, a 3×3 convolution operation is then performed, and the number of convolution kernels is C, a feature map X1 of size H×W×C is obtained, where the feature maps of enhancement module 1 and enhancement module 2 are X 11 and X 12 , and then perform channel attention extraction and spatial attention extraction on the obtained feature maps respectively; there is no full connection step in the extraction process of channel attention and spatial attention, and the full convolution method is adopted to reduce the model parameters, which is more friendly to small devices or lightweight devices.

[0046] The present invention focuses on the difficulty of large scale differences of remote sensing targets. The enhancement module outputs a dual-scale feature map which enters the detection head for detection. The dual-scale feature map respectively retains the key information of the top and bottom layers in the feature extraction module, taking into account both large and small targets in the image, and to a certain extent solves the problem of missed detection of small targets in the problem of multi-scale target feature characterization.

[0047] The spatial attention extraction process is as follows: the feature map X1 is subjected to global maximum pooling (MP) and global average pooling (AP) operations along the C (channel) direction to obtain two H×W×1 feature maps, respectively. Then, they are subjected to 3×3 convolution operations and element-wise operations, and finally, the final spatial attention feature map S1 is generated by a sigmoid activation operation. The spatial attention feature maps of enhancement module 1 and enhancement module 2 are S 11 and S 12 , dimensions are H×W×1.

[0048] The spatial attention extraction formula is:

[0049] S1=σ(Conv(AP C (X1))+(Conv(MP C(X1))

[0050] In the formula, σ represents the sigmoid operation, Conv represents the convolution calculation, and AP C Indicates global average pooling operation along the C (channel) direction, MP C Indicates a global maximum pooling operation along the C (channel) direction.

[0051] The channel attention extraction process is as follows: the feature map X1 is subjected to global maximum pooling (MP) and global average pooling (AP) operations along the H and W (height and width) directions to obtain two 1×1×C feature maps, respectively. Then, they are subjected to 3×3 convolution operations and element-wise operations, and finally, the final channel attention feature map S2 is generated by a sigmoid activation operation. The spatial attention feature maps of enhancement module 1 and enhancement module 2 are S 21 and S 22 , with a size of 1×1×C. By splicing the two adjacent feature maps after extraction, the spatial attention and channel attention links are added respectively, which fully captures the key information in the features and provides a more comprehensive and effective feature extraction capability for the target detection model.

[0052] The channel attention extraction formula is:

[0053] S2=σ(Conv(AP HW (X1))+(Conv(MP HW (X1))

[0054] In the formula, σ represents the sigmoid operation, Conv represents the convolution calculation, and AP HW Indicates global average pooling operation along the H and W (height and width) directions, MP HW Indicates that the global maximum pooling operation is performed along the H and W (height and width) directions.

[0055] The above global maximum pooling (MP) operation formula is:

[0056]

[0057] In the formula, y kij Represents the k-th feature map in R ij The maximum pooling output value, x kpq Represents the rectangular area R ij The value of the element at (p,q) in .

[0058] The above global average pooling (AP) operation formula is:

[0059]

[0060] In the formula, g kij Represents the k-th feature map in R ij The average pooling output value, t kpq Represents the rectangular area R ij The value of the element at (p,q) in |R ij | represents the number of all elements.

[0061] After performing spatial attention extraction and channel attention extraction operations, the feature map X 11 (or X 12 ) are respectively related to the spatial attention feature map S 11 (or S 12 ) and channel attention feature map S 21 (or S 22 ) to obtain the final feature map X 12 (or X 22 ); The feature maps finally output by the first enhanced fusion module and the second enhanced fusion module are X 21 and X 22 , that is, a dual-scale feature map.

[0062] The enhancement module is followed by the detection head module. The detection head module adopts the SSD detection head, and the loss function includes category loss and positioning loss. The feature map X2 is sent to the detection head and finally compared with the label value to calculate the loss.

[0063] The category loss formula is:

[0064] F L (p t )=-α(1-p t ) γ log(p t )

[0065] In the formula, p t Indicates confidence, p is the prediction probability, y=1 indicates foreground, y≠1 indicates background, α is a weight parameter that can suppress the imbalance in the number of positive and negative samples, and γ is a hyperparameter with a value range of [0, 5].

[0066] The positioning loss formula is:

[0067]

[0068] In the formula, y i is the true value, is the predicted value, and n is the number of sample points.

[0069] Step S3 may specifically include the following steps:

[0070] Store the training data set and the test data set in the train and test folders respectively, and move them into the data folder of the algorithm package. Run the train.py file in the terminal to start training the network. During the training process, first call the feature extraction module to perform feature extraction on the input image, then call the enhancement module to fuse high-level semantic information with underlying detail information, and finally call the detection head network to output the fused feature map X2 to the loss function to calculate the loss value between the true value, and then complete a forward propagation. Subsequently, the neural network will adjust the corresponding parameter values ​​in the model according to the loss value, and continuously update and iterate until the preset number of iterations is completed. At this time, the loss value and time relationship curve tends to be flat and converged. At this time, the entire model is the trained network model with optimal parameters.

[0071] Step S4 includes the following steps:

[0072] The target remote sensing image is used as the detection model trained in step S3 for forward propagation to obtain the location information and category confidence information of the predicted target. Finally, the generated bounding box is screened using the non-maximum suppression method (NMS) to obtain the bounding box and category confidence of each target.

[0073] The present invention Figure 4 The remote sensing image example in is the input, and the detection result example is as follows Figure 5 As shown, the effectiveness of the method proposed in the present invention is verified.

[0074] As described above, the present invention adopts a two-level feature pyramid structure based on ResNet50 for feature extraction, and uses two enhancement modules to process the feature maps of high-level dimensions and low-level dimensions respectively. Compared with the traditional pyramid network, while retaining the rich semantic information of the high-level, it more comprehensively captures the edge, corner, texture and other detail information of the bottom layer. For remote sensing targets in dense and overlapping situations, it can accurately identify the edges of the objects and has higher detection accuracy.

[0075] Finally, it should be noted that the features mentioned and / or shown in the above description of the exemplary embodiments of the present invention may be combined in the same or similar manner into one or more other embodiments, combined with the features in other embodiments or substituted for the corresponding features in other implementations. The technical solutions obtained by these combinations or substitutions shall also be deemed to be included in the protection scope of the present invention.

Claims

1. A remote sensing image target detection method based on improved pyramid, characterized in that: The following steps are involved: Step 1, obtaining a data set of remote sensing images; Step 2, constructing a remote sensing image target detection model based on an improved pyramid, wherein the target detection model includes a feature extraction module, two enhancement modules and two detection head modules; The feature extraction module is a two-level feature pyramid structure based on ResNet50. The first-level pyramid structure is used to extract four feature maps of different sizes of the input image and obtain a fused feature map through the FPN structure. The second-level pyramid structure is used to perform feature fusion on the obtained fused feature map from bottom to top. The two enhancement modules respectively take the two feature maps of the high-level size and the two feature maps of the low-level size in the four feature maps obtained by the fusion again as input, perform channel attention extraction and spatial attention extraction on the input feature maps, and multiply the channel attention feature map with the spatial attention feature map to obtain the final dual-scale feature map; The two detection head modules are used to process the dual-scale feature map, and output the category and bounding box information of the target through the classification branch and the regression branch respectively; Step 3, iteratively train the constructed target detection model using the acquired data set, and obtain the optimal network model parameters after the detection model converges; Step 4: Use the obtained optimal network model parameters to detect the target remote sensing image and obtain the classification and positioning information of the target in the image.

2. The remote sensing image target detection method based on improved pyramid according to claim 1, characterized in that: In step 1, the acquired dataset is a public dataset or a self-photographed dataset.

3. The remote sensing image target detection method based on improved pyramid according to claim 2, characterized in that: Step 1 also includes: dividing the acquired data into a training set and a test set, with a division ratio of 8:2 or 7:

3.

4. The remote sensing image target detection method based on improved pyramid according to claim 1, characterized in that: In step 2, the sizes of the images extracted by the feature extraction module are 1 / 4, 1 / 8, 1 / 16 and 1 / 32 of the input image size respectively.

5. The remote sensing image target detection method based on improved pyramid according to claim 4, characterized in that: In step 2, the two enhancement modules are respectively a first enhancement fusion module and a second enhancement fusion module. The input of the first enhancement fusion module is two of the four feature maps output by the feature extraction module, whose sizes are 1 / 16 and 1 / 32 of the input image size. The input of the second enhancement fusion module is two of the four feature maps output by the feature extraction module, whose sizes are 1 / 8 and 1 / 4 of the input image size.

6. The remote sensing image target detection method based on improved pyramid according to claim 4 or 5, characterized in that: In step 2, each enhancement module processes the input feature map as follows: keep the image size consistent by upsampling or downsampling, and then perform a splicing operation; perform a convolution operation on the spliced ​​feature map; perform channel attention extraction and spatial attention extraction on the feature map obtained by convolution respectively; and multiply the channel attention feature map with the spatial attention feature map.

7. The remote sensing image target detection method based on improved pyramid according to claim 1, characterized in that: In step 2, the spatial attention extraction process of each enhancement module is as follows: the feature map is subjected to global maximum pooling and global average pooling operations along the channel direction to obtain two feature maps, which are then subjected to element-wise operations after convolution operations, and finally the final spatial attention feature map is generated through sigmoid activation operations.

8. The remote sensing image target detection method based on improved pyramid according to claim 1, characterized in that: In step 2, the channel attention extraction process of each enhancement module is as follows: the feature map is subjected to global maximum pooling and global average pooling operations along the height and width directions to obtain two feature maps respectively, and then subjected to element-wise operations after convolution operations, and finally a sigmoid activation operation is performed to generate the final channel attention feature map.

9. The remote sensing image target detection method based on improved pyramid according to claim 1, characterized in that: In step 2, the detection head module uses the SSD detection head, and the loss function includes category loss and positioning loss.

Citation Information

Patent Citations

  • Pedestrian detection method based on semantic enhancement multi-scale feature pyramid fusion

    CN112446308A

  • Efficient real-time semantic segmentation method

    CN112529064A

  • Remote sensing image target detection method based on improved RetinaNet

    CN116310850A

  • Impact crater detection method

    CN118470509A

  • Object detection method based on attention-enhanced bidirectional feature pyramid network (a-bifpn)

    GB202217717D0

Cited By

  • Power transmission tower bolt detection method based on improved yov5s, electronic equipment and storage medium

    CN120563938A