Improved remote sensing image target detection method

By introducing PyConv4 module, PSConv module, CGAM module and decoupling head in remote sensing image object detection, the problem of multi-scale object detection in complex background is solved, and higher detection accuracy and robustness are achieved.

CN119992347AInactive Publication Date: 2025-05-13WUHAN TEXTILE UNIV
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202510471110.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art has poor detection of multi-scale targets and cover targets in complex contexts, especially in high target similarity and high resolution image scenarios, where detection of small targets and occluded targets still poses challenges.

Method used

An improved remote sensing image object detection method is designed, and the accuracy and robustness of the model in multi-scale object detection is improved by introducing PyConv4 module, PSConv module, CGAM module and decoupling head.

Benefits of technology

Through this method, the model can better adapt to multi-scale object detection tasks, improve the accuracy and robustness of the detection, and enhance the performance of the model in different scales and complex contexts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992347A_ABST
    Figure CN119992347A_ABST
Patent Text Reader

Abstract

The invention discloses an improved remote sensing image target detection method, and the method comprises the following steps: S1, obtaining a remote sensing image target detection data set, and dividing the data set into a training set, a verification set and a test set according to the proportion of 7: 2: 1; s2, training the data set by using an improved detection model; and S3, performing target detection on the remote sensing image based on the improved detection model. According to the invention, the PyConv4 module, the PSConv module and the CGAM module are designed, and the decoupling head is used to realize the classification and positioning of the target, so that the effectiveness and accuracy of remote sensing image target detection are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target detection, and in particular to an improved remote sensing image target detection method. Background Art

[0002] With the development of the big data era, various large-scale remote sensing image data sets are shrinking, and these data sets contain a lot of important information, such as buildings, roads, farmlands, natural resources, etc. This information plays a vital role in earth observation, environmental monitoring, and urban planning. Therefore, the process of accurately obtaining this important information from remote sensing images has become particularly important, and target detection technology plays a vital role in this process.

[0003] Object detection occupies a core position in computer vision. With the development of convolutional neural networks, object detection algorithms have made great progress. However, due to the different acquisition angles of remote sensing images and changeable weather, there are significant differences in the scale and image clarity of the target. Secondly, in scenes with complex backgrounds, high target similarity and high-resolution images, the detection of small targets and occluded targets is still a major challenge in remote sensing images, and further improvements are needed.

[0004] The Chinese patent publication number CN114998756B discloses "a remote sensing image detection method, device and storage medium based on yolov5". This method introduces the CBAM attention mechanism based on the yolov5 model to better extract features in the two dimensions of channel and space, improve detection accuracy, replace the original PANet of the yolov5 model with a dedicated feature fusion module, better perform feature fusion, and balance feature information of different scales. However, the above solution is only applicable to a small number of data sets, and the effectiveness and accuracy of the model still need to be further improved.

[0005] Therefore, target detection for multi-scale targets and covered targets in complex backgrounds still needs to be optimized, and more effective methods need to be proposed to overcome the limitations of existing technologies. Summary of the invention

[0006] In view of the above defects or improvement needs of the prior art, the present invention provides an improved remote sensing image target detection method, which improves the effectiveness and accuracy of remote sensing image target detection by designing a PyConv4 module, a PSConv module, a CGAM module and using a decoupling head to achieve target classification and positioning.

[0007] To achieve the above object, according to one aspect of the present invention, an improved remote sensing image target detection method is provided, the method comprising the following steps: S1: Obtain a remote sensing image target detection dataset and divide the dataset into a training set, a validation set, and a test set in a ratio of 7:2:1; S2: Using the data set to train an improved detection model, the improved detection model includes an input layer, a backbone network, a neck network and a prediction head, and the improved detection model is set as follows: Using the designed PyConv4 module at the last CBS in the backbone network, and integrating the SimAM module into the PyConv4 module through the PSConv module; In the neck network, the GAM module is introduced into the C3 module via the CGAM module; Use the decoupling head to classify and locate the final target; S3: Perform target detection on remote sensing images based on the trained improved detection model.

[0008] As an embodiment of the present application, step S3 includes: S31: The input layer inputs the image into the backbone network for feature extraction to obtain feature maps of different scales; S32: Inputting feature maps of different scales into the neck network for processing and fusion to obtain a fused feature map; S33: Input the fused feature map into the prediction head for detection, obtain the target detection result and output it.

[0009] As an embodiment of the present application, the network structure of the improved detection model is: The backbone network includes four CBS modules, four C3 modules, one PSConv module and one pyramid pooling layer, which extract features of different scales in turn; The neck network includes six CBS modules, three upsampling modules, four C3 modules, six fusion modules and one CGAM module. The neck network performs split-path fusion on semantic features of different scales and obtains feature maps of three different scales according to three fusion routes. The prediction head includes three convolution modules and three decoupling heads, which detect feature maps of three different scales respectively.

[0010] As an embodiment of the present application, the three fusion routes are respectively a first fusion route, a second fusion route and a third fusion route, and the structure of the feature fusion network according to the data flow direction is defined as: a first CBS module, a first upsampling module, a first fusion module, a first C3 module, a second CBS module, a second upsampling module, a second fusion module, a second C3 module, a third CBS module, a third fusion module, a third C3 module, a fourth CBS module, a fourth fusion module, a fourth C3 module, a fifth CBS module, a fourth fusion module, a fifth C3 module, a sixth CBS module, a sixth fusion module and a CGAM module; The first fusion route is: The semantic features of the third scale are sequentially input into the first CBS module and the first upsampling module, fused with the semantic features of the second scale in the first fusion module, and then input into the first C3 module and the second CBS module to obtain the first intermediate features. The first intermediate features are input into the second upsampling module, fused with the semantic features of the first scale in the second fusion module, and then input into the second C3 module and the third CBS module to obtain the second intermediate features. The second intermediate features are input into the third upsampling module, and then input into the third C3 module and the fourth CBS module after the third fusion module, and then fused with the second intermediate features in the fourth fusion module, and then input into the fourth C3 module to obtain the first feature map. The second fusion route is: The first fused feature is input into the fifth CBS module, fused with the first intermediate feature in the fifth fusion module, and then input into the third C3 module after fusion to obtain a second feature map; The third fusion route is: The second fusion feature is input into the sixth CBS module, and then into the CGAM module after passing through the sixth fusion module to obtain the third feature map.

[0011] As an embodiment of the present application, the PyConv4 module is a four-layer pyramid convolution, which includes a pyramid of four different types of convolution kernels. The convolution kernels of each layer are different in spatial size, and the size of the convolution kernels gradually increases from the bottom to the top of the pyramid. Increase to , while the depth gradually decreases from the first layer to the fourth layer.

[0012] As an embodiment of the present application, the PSConv module includes four layers of pyramid convolution and a SimAM module, and the PSConv module integrates the SimAM module into each layer of the PyConv4 module.

[0013] As an embodiment of the present application, the CGAM module includes three CBM modules, a residual unit, a GAM module, and a fusion module. The input feature map is divided into two branches through two parallel CBM modules, one branch passes through a residual unit containing multiple Bottleneck modules to extract deep features, and the other branch passes through the GAM module to enhance important features of channel and spatial dimensions. The outputs of the two branches are spliced ​​in the channel dimension through a fusion module, and the fused features are mapped to the target output channel number through a CBM module to generate an enhanced feature map.

[0014] As an embodiment of the present application, the decoupling head includes four 1×1 convolutional layers and two 3×3 convolutional layers. The input feature map is processed through multiple convolutional layers to extract features at different levels, and then calculated through category, regression and confidence branches respectively to output the category, bounding box regression information and confidence of each anchor box. According to the anchor box and grid information, the final target detection result is decoded, including coordinates, category and confidence.

[0015] As an embodiment of the present application, before step S3, the improved detection model is compared with the classical model and other improved models, and the improved detection model is compared with the classical model and other improved models. and The two evaluation indicators are used to compare and verify the improved detection model. The specific formula is as follows:

[0016]

[0017]

[0018]

[0019] in, It represents the average precision when the IoU threshold is 0.5; It represents the average precision in the range of IoU threshold from 0.5 to 0.95; represents the average precision; It represents the ratio of the number of samples correctly predicted by the algorithm to the total number of samples; Indicates the proportion of samples that are correctly predicted in positive samples; Indicates the number of positive samples that are correctly classified; Indicates the number of positive samples that are misclassified; Indicates the number of negative samples that are misclassified; Indicates the total number of categories.

[0020] As an embodiment of the present application, the improved detection model is compared with the classic model and other improved models. Specifically, in order to accurately verify the performance of the improved detection model in the remote sensing image target detection task, it is compared with a series of YOLOv5, YOLOv3, YOLOv6 and Gold-YOLO models respectively.

[0021] The beneficial effects of the present invention are: (1) The present invention designs a PyConv4 module, which is a four-layer pyramid convolution. The size of the convolution kernel can be expanded without adding additional cost, and convolution kernels of different spatial resolutions and depths can be applied in parallel. By using convolution kernels of different sizes to obtain target information of different sizes and scales, the model can better adapt to multi-scale target detection tasks, thereby improving the accuracy and robustness of detection and effectively solving the problem of multi-scale target detection.

[0022] (2) The present invention designs a PSConv module, which includes four layers of pyramid convolution and a SimAM module. The SimAM module is a simple, parameter-free attention mechanism that does not require learning additional parameters and is therefore easier to integrate into a neural network. While maintaining model accuracy, it also reduces the complexity of training and reasoning. Therefore, the parameter-free attention mechanism SimAM is integrated into the pyramid convolution through the PSConv module to suppress the interference of background information and emphasize the importance of targets at different scales.

[0023] (3) The present invention designs a CGAM module and introduces a global attention mechanism (GAM) after the CBM of the second branch of the traditional C3 module, thereby enhancing the dependency between layers within the C3 module, thereby improving the model's contextual understanding ability, improving model performance and accelerating the reasoning process.

[0024] (4) The present invention uses a decoupling head to classify and locate the final target. The decoupling detection head extracts the position and category information of the target respectively, learns through different network branches, and finally fuses the two. This design better meets the needs of classification and positioning tasks, improves the feature expression ability of the model in different tasks, and enhances its ability to adapt to remote sensing images of different types and environments. By learning and fusing the position and category information of the target respectively, the model can more effectively process the texture content and edge information of the target, thereby improving the performance and robustness of remote sensing image target detection, making the model more suitable for the detection needs of targets of different scales. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 A flowchart of an improved remote sensing image target detection method provided in an embodiment of the present invention; Figure 2 A network structure diagram of an improved detection model of an improved remote sensing image target detection method provided in an embodiment of the present invention; Figure 3 A PyConv4 module structure diagram of an improved remote sensing image target detection method provided in an embodiment of the present invention; Figure 4 A PSConv module structure diagram of an improved remote sensing image target detection method provided in an embodiment of the present invention; Figure 5 A CGAM module structure diagram of an improved remote sensing image target detection method provided in an embodiment of the present invention; Figure 6 A structural diagram of a decoupling head of an improved remote sensing image target detection method provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0026] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0027] It should be noted that all directional indications (such as up, down, left, right, front, back, etc.) in the embodiments of the present invention are only used to explain the relative position relationship, movement status, etc. between the components under a certain specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indication will also change accordingly.

[0028] In the present invention, unless otherwise clearly specified and limited, the terms "connection", "fixation", etc. should be understood in a broad sense. For example, "fixation" can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, it can be the internal connection of two elements or the interaction relationship between two elements, unless otherwise clearly defined. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0029] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of the present invention, the descriptions of "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In addition, the meaning of "and / or" appearing in the full text includes three parallel schemes. Taking "A and / or B" as an example, it includes scheme A, or scheme B, or a scheme that satisfies both A and B. In addition, the technical solutions between the various embodiments can be combined with each other, but it must be based on the ability of ordinary technicians in the field to implement. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.

[0030] Reference Figure 1-Figure 6 The first aspect of the present invention provides an improved remote sensing image target detection method, the method comprising the following steps: S1: Obtain a remote sensing image target detection dataset and divide the dataset into a training set, a validation set, and a test set in a ratio of 7:2:1; S2: The data set is trained using an improved detection model, wherein the improved detection model includes an input layer, a backbone network, a neck network, and a prediction head. The improved detection model is specifically improved as follows: The designed PyConv4 module is used at the last CBS in the backbone network. The PyConv4 module designed by the present invention obtains target information of different sizes and proportions by using convolution kernels of different sizes, so that the model can better adapt to multi-scale target detection tasks, thereby improving the accuracy and robustness of detection; and the SimAM module is integrated into the PyConv4 module through the PSConv module to suppress the interference of background information and emphasize the importance of targets at different scales; In the neck network, the GAM module is introduced into the C3 module through the CGAM module to enhance the dependency between layers and improve the context understanding ability of the model; Use the decoupling head to classify and locate the final target, further improving the accuracy of remote sensing image target detection and enhancing the generalization ability of the model; S3: Perform target detection on remote sensing images based on the improved detection model.

[0031] As an embodiment of the present application, step S3 includes: S31: The input layer inputs the image into a series of convolution and pooling layers in the backbone network for feature extraction to obtain feature maps of different scales; S32: Inputting feature maps of different scales into the neck network for processing and fusion to obtain a fused feature map; S33: The fused feature map is input into the prediction head for detection, and then the obtained feature map is detected. During the detection, a series of target boxes are generated. The model removes redundant target boxes through the non-maximum suppression algorithm, retains the box with the highest confidence, obtains the target detection result and outputs it.

[0032] In this embodiment, after the above steps, the network structure diagram of the improved detection model constructed by the present invention is as follows: Figure 2 As shown, specifically: The backbone network includes four CBS modules, four C3 modules, one PSConv module and one pyramid pooling layer, which extract features of different scales in turn; The neck network includes six CBS modules, three upsampling modules, four C3 modules, six fusion modules and one CGAM module. The neck network performs split fusion on semantic features of different scales and obtains feature maps of three different scales according to three fusion routes. The prediction head includes three convolution modules and three decoupling heads, which detect feature maps of three different scales respectively.

[0033] As an embodiment of the present application, the three fusion routes are respectively a first fusion route, a second fusion route and a third fusion route, and the structure of the feature fusion network according to the data flow direction is defined as: a first CBS module, a first upsampling module, a first fusion module, a first C3 module, a second CBS module, a second upsampling module, a second fusion module, a second C3 module, a third CBS module, a third fusion module, a third C3 module, a fourth CBS module, a fourth fusion module, a fourth C3 module, a fifth CBS module, a fourth fusion module, a fifth C3 module, a sixth CBS module, a sixth fusion module and a CGAM module; The first fusion route is: The semantic features of the third scale are sequentially input into the first CBS module and the first upsampling module, fused with the semantic features of the second scale in the first fusion module, and then input into the first C3 module and the second CBS module to obtain the first intermediate features. The first intermediate features are input into the second upsampling module, fused with the semantic features of the first scale in the second fusion module, and then input into the second C3 module and the third CBS module to obtain the second intermediate features. The second intermediate features are input into the third upsampling module, and then input into the third C3 module and the fourth CBS module after the third fusion module, and then fused with the second intermediate features in the fourth fusion module, and then input into the fourth C3 module to obtain the first feature map. The second fusion route is: The first fused feature is input into the fifth CBS module, fused with the first intermediate feature in the fifth fusion module, and then input into the third C3 module after fusion to obtain a second feature map; The third fusion route is: The second fusion feature is input into the sixth CBS module, and then into the CGAM module after passing through the sixth fusion module to obtain the third feature map.

[0034] For standard convolution, let the size of the input image be ,in is the width, is the height, is the depth of the input feature map, and the size of the output image is ,in is the width, is the height, is the depth of the output feature map, then the size of the convolution kernel is , Represents the height and width of the convolution kernel, Indicates the depth of the standard convolution kernel, which is consistent with the depth of the input feature map. The standard convolution uses a fixed convolution kernel and step size; however, there are a large number of multi-scale targets in remote sensing images, and the standard convolution using a fixed convolution kernel and step size is insufficient for effective target detection. In order to solve this limitation, the present invention designs a PyConv4 module.

[0035] As an embodiment of the present application, the PyConv4 module structure diagram designed by the present invention is as follows Figure 3 As shown in Figure 1, the PyConv4 module is a four-layer pyramid convolution, which contains four different types of convolution kernels. The convolution kernels of each layer are different in spatial size, and the size of the convolution kernels gradually increases from the bottom to the top of the pyramid. Increase to , and the depth gradually decreases from the first layer to the fourth layer. Pyramid convolution introduces the concept of grouped convolution, which divides the input feature map into multiple different groups, each of which uses a convolution kernel with a different depth. As the number of groups increases, the depth of the convolution kernel gradually decreases, thereby reducing the number of parameters and the amount of calculation.

[0036] Compared with standard convolution, pyramid convolution can expand the size of the convolution kernel without adding additional cost, and can apply convolution kernels of different spatial resolutions and depths in parallel. Therefore, pyramid convolution can process input at multiple filter scales, thereby capturing more detailed information. In addition, for a pyramid convolution with 4 layers, the number of parameters is: ; Its floating point operations (FLOPs) are: ; In contrast, for standard convolution, the number of parameters is: , the floating point operation amount is: . It can be seen that the number of parameters and floating-point operations of pyramid convolution are lower than those of standard convolution; more importantly, under the same number of model parameters, pyramid convolution is more efficient than standard convolution due to its highly parallelized characteristics; in addition, for each layer of the network, pyramid convolution can be flexibly designed. Therefore, pyramid convolution has high flexibility and scalability.

[0037] As an embodiment of the present application, in order to further improve the accuracy of remote sensing image target detection, the present invention designs a PSConv module, such as Figure 4 As shown, the PSConv module includes four layers of pyramid convolution and a SimAM module, and the PSConv module integrates the SimAM module into each layer of the PyConv4 module.

[0038] Specifically, the SimAM module is a simple, parameter-free attention mechanism. After the input feature map passes through SimAM, the square of the deviation between each pixel value and the channel mean is calculated to capture the differences in local features. Then, the noise is suppressed by normalization combined with a small constant to generate preliminary attention weights. Next, the Sigmoid activation function is used to limit the weights and normalize their range to [0,1]. Finally, the generated attention weights are multiplied element by element with the original feature map to enhance the salient areas and suppress the unimportant areas, and the enhanced feature map is output while maintaining the same dimension as the input feature map.

[0039] Compared with the traditional attention mechanism, SimAM does not require learning additional parameters and is therefore easier to integrate into a neural network. It reduces the complexity of training and reasoning while maintaining model accuracy. Therefore, the present invention uses the parameter-free attention mechanism SimAM in pyramid convolution to suppress the interference of background information and emphasize the importance of targets at different scales.

[0040] The traditional C3 module only contains three convolution modules. Although the use of the C3 module can reduce the number of parameters, it lacks an integrated attention mechanism, which limits its ability to emphasize key areas in the feature map. This limitation may have an adverse effect on detection accuracy in scenes with complex backgrounds or occluded targets. To solve this problem, the present invention designs a CGAM module, which introduces a global attention mechanism GAM after the CBM of the second branch of the C3 module to enhance the dependency between layers and improve the contextual understanding ability of the model.

[0041] As an embodiment of the present application, the structural diagram of the CGAM module designed by the present invention is as follows: Figure 5As shown, the CGAM module includes three CBM modules, a residual unit, a GAM module, and a fusion module. The input feature map is divided into two branches through two parallel CBM modules. One branch passes through a residual unit containing multiple Bottleneck modules to extract deep features, and the other branch passes through an attention mechanism GAM module to enhance important features of channel and spatial dimensions. The outputs of the two branches are spliced ​​in the channel dimension through a fusion module, and the fused features are mapped to the target output channel number through a CBM module to generate an enhanced feature map.

[0042] Specifically, after the input feature map passes through the GAM module, the feature map is first processed by the channel attention branch: the feature map channel is expanded, the channel information is compressed and expanded through a two-layer fully connected network, and the channel attention weight is generated, and multiplied with the input feature map channel by channel to complete the channel weighting; then enter the spatial attention branch, extract spatial features through two 7x7 convolutional layers with normalization layers, generate spatial attention weights, and then multiply with the weighted feature map position by position, and finally output a feature map that integrates channel and spatial attention. The present invention introduces the global attention mechanism GAM into the C3 module through the CGAM module, which can enhance the dependency between the layers within the C3 module, thereby improving the contextual understanding ability of the model, improving the model performance and accelerating the reasoning process.

[0043] The traditional YOLOv5 detection head is a coupled detection head, which consists of a convolutional layer, a fully connected layer, and an activation function, and is implemented by fusing and sharing the two branches of classification and regression. However, classification and localization focus on different things: classification focuses on the texture content of the target, while localization focuses more on the edge information of the target. Therefore, the shared design limits the network's ability to express features in different tasks. In addition, the inconsistency of remote sensing image acquisition methods and environments leads to significant differences in the model when detecting different images.

[0044] In order to solve the above problems, the present application uses a decoupling head to classify and locate the final target. The decoupling head used in the present invention is as follows: Figure 6As shown, the decoupling head includes four 1×1 convolutional layers and two 3×3 convolutional layers. The input feature map is processed through multiple convolutional layers to extract features at different levels, and then the category, regression and confidence branches are used for calculation respectively to output the category, bounding box regression information and confidence of each anchor box. According to the anchor box and grid information, the final target detection result, including coordinates, category and confidence, is decoded. The decoupling head adopted by the present invention is more concise and alleviates the additional delay overhead to a certain extent. Unlike the coupled detection head, the decoupling detection head extracts the location and category information of the target respectively, learns through different network branches, and finally fuses the two. This design better meets the needs of classification and positioning tasks, improves the feature expression ability of the model in different tasks, and enhances its ability to adapt to remote sensing images of different types and environments. By learning and fusing the location and category information of the target respectively, the model can more effectively process the texture content and edge information of the target, thereby improving the detection ability and generalization ability of the model.

[0045] As an embodiment of the present application, before step S3, the improved detection model is compared with the classic model and other improved models, and the improved detection model is compared with the classic model and other improved models. and The two evaluation indicators are used to compare and verify the improved detection model. The specific formula is as follows:

[0046]

[0047]

[0048]

[0049] in, It represents the average precision when the IoU threshold is 0.5; It represents the average precision in the range of IoU threshold from 0.5 to 0.95; represents the average precision; It represents the ratio of the number of samples correctly predicted by the algorithm to the total number of samples; Indicates the proportion of samples that are correctly predicted in positive samples; Indicates the number of positive samples that are correctly classified; Indicates the number of positive samples that are misclassified; Indicates the number of negative samples that are misclassified; Indicates the total number of categories.

[0050] As an embodiment of the present application, the improved detection model is compared with the classic model and other improved models. Specifically, in order to accurately verify the performance of the improved detection model in the remote sensing image target detection task, it is compared with a series of YOLOv5, YOLOv3, YOLOv6 and Gold-YOLO models, and the experimental results are shown in Table 1: Table 1. Comparison of experimental results with different models on 6 datasets

[0051] It can be seen from Table 1 that the improved model of the present invention has certain advantages. The improved PCD-YOLOv5s, PCD-YOLOv5m and PCD-YOLOv5x show better performance than YOLOv5s, YOLOv5m and YOLOv5x on five data sets (except the HRSC2016 data set); compared with classic algorithms such as YOLOv3, YOLOv6, Gold-YOLO and YOLOv9c, the improved algorithms show better performance.

[0052] In order to evaluate the performance of each module improved by the present invention on the original model, namely, the PyConv4 module, the PSConv module, the CGAM module and the Decoupled Head, ablation experiments were conducted on six datasets. The results of the ablation experiments are shown in Table 2.

[0053] Table 2. Ablation experiments

[0054] According to Table 2, the improved model achieves better results than the original model (YOLOv5s) on the validation sets of the six datasets. and ,in, They have respectively increased , , , , and , They have respectively increased , , , , and , which further illustrates the feasibility and effectiveness of the improved model.

[0055] The invention is improved based on the YOLOv5 algorithm. Firstly, a new pyramid convolution module PSConv is designed to improve the detection capability of multi-scale targets. Secondly, a PSConv module is designed, which integrates a simple, parameter-free attention mechanism (SimAM) into the pyramid convolution used in the backbone network to increase the attention to important objects of different scales, thereby improving the detection capability of multi-scale targets. Then, a CGAM module is designed to introduce the global attention mechanism GAM into the C3 module used in the neck network, thereby improving the detection capability of small objects by making full use of global context information. Finally, a decoupling head is used to classify and locate the final target, thereby further improving the performance and robustness of remote sensing image target detection, thereby making the model more suitable for the detection needs of targets of different scales.

[0056] The above descriptions are only some preferred embodiments of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the above features are replaced with (but not limited to) technical features with similar functions disclosed in the embodiments of the present disclosure.

Claims

1. An improved remote sensing image target detection method, characterized in that: The method comprises the following steps: S1: Obtain a remote sensing image target detection dataset and divide the dataset into a training set, a validation set, and a test set in a ratio of 7:2:1; S2: The data set is trained using an improved detection model, wherein the improved detection model includes an input layer, a backbone network, a neck network, and a prediction head. The improved detection model is specifically improved as follows: Using the designed PyConv4 module at the last CBS in the backbone network, and integrating the SimAM module into the PyConv4 module through the PSConv module; In the neck network, the GAM module is introduced into the C3 module via the CGAM module; Use the decoupling head to classify and locate the final target; S3: Perform target detection on remote sensing images based on the improved detection model.

2. The improved remote sensing image target detection method according to claim 1, characterized in that: The step S3 comprises: S31: The input layer inputs the image into the backbone network for feature extraction to obtain feature maps of different scales; S32: Inputting feature maps of different scales into the neck network for processing and fusion to obtain a fused feature map; S33: Input the fused feature map into the prediction head for detection, obtain the target detection result and output it.

3. The improved remote sensing image target detection method according to claim 1, characterized in that: The network structure of the improved detection model is: The backbone network includes four CBS modules, four C3 modules, one PSConv module and one pyramid pooling layer, which extract features of different scales in turn; The neck network includes six CBS modules, three upsampling modules, four C3 modules, six fusion modules and one CGAM module. The neck network performs split fusion on semantic features of different scales and obtains feature maps of three different scales according to three fusion routes. The prediction head includes three convolution modules and three decoupling heads, which detect feature maps of three different scales respectively.

4. The improved remote sensing image target detection method as claimed in claim 3, characterized in that: The three fusion routes are the first fusion route, the second fusion route and the third fusion route, and the structure of the feature fusion network according to the data flow direction is defined as: the first CBS module, the first upsampling module, the first fusion module, the first C3 module, the second CBS module, the second upsampling module, the second fusion module, the second C3 module, the third CBS module, the third fusion module, the third C3 module, the fourth CBS module, the fourth fusion module, the fourth C3 module, the fifth CBS module, the fourth fusion module, the fifth C3 module, the sixth CBS module, the sixth fusion module and the CGAM module; The first fusion route is: The semantic features of the third scale are sequentially input into the first CBS module and the first upsampling module, fused with the semantic features of the second scale in the first fusion module, and then input into the first C3 module and the second CBS module to obtain the first intermediate features. The first intermediate features are input into the second upsampling module, fused with the semantic features of the first scale in the second fusion module, and then input into the second C3 module and the third CBS module to obtain the second intermediate features. The second intermediate features are input into the third upsampling module, and then input into the third C3 module and the fourth CBS module after the third fusion module, and then fused with the second intermediate features in the fourth fusion module, and then input into the fourth C3 module to obtain the first feature map. The second fusion route is: The first fused feature is input into the fifth CBS module, fused with the first intermediate feature in the fifth fusion module, and then input into the third C3 module after fusion to obtain a second feature map; The third fusion route is: The second fusion feature is input into the sixth CBS module, and then into the CGAM module after passing through the sixth fusion module to obtain the third feature map.

5. The improved remote sensing image target detection method according to claim 1, characterized in that: The PyConv4 module is a four-layer pyramid convolution, which contains four different types of convolution kernels. The convolution kernels of each layer are different in spatial size, and the size of the convolution kernels gradually increases from the bottom to the top of the pyramid. Increase to , while the depth gradually decreases from the first layer to the fourth layer.

6. The improved remote sensing image target detection method according to claim 1, characterized in that: The PSConv module includes four layers of pyramid convolution and a SimAM module, and the PSConv module is used to integrate the SimAM module into each layer of the PyConv4 module.

7. The improved remote sensing image target detection method according to claim 1, characterized in that: The CGAM module includes three CBM modules, a residual unit, a GAM module, and a fusion module. The input feature map is divided into two branches through two parallel CBM modules. One branch passes through a residual unit containing multiple Bottleneck modules to extract deep features, and the other branch passes through a GAM module to enhance important features in channel and spatial dimensions. The outputs of the two branches are spliced ​​in the channel dimension through a fusion module. The fused features are then mapped to the target output channel number through a CBM module to generate an enhanced feature map.

8. The improved remote sensing image target detection method according to claim 1, characterized in that: The decoupling head includes four 1×1 convolutional layers and two 3×3 convolutional layers. The input feature map is processed by multiple convolutional layers to extract features at different levels, and then calculated by category, regression and confidence branches respectively, and the category, bounding box regression information and confidence of each anchor box are output. According to the anchor box and grid information, the final target detection result, including coordinates, category and confidence, is decoded.

9. The improved remote sensing image target detection method according to claim 1, characterized in that: Before step S3, the improved detection model is compared with the classical model and other improved models, and the and The two evaluation indicators are used to compare and verify the improved detection model. The specific formula is as follows: in, It represents the average precision when the IoU threshold is 0.5; It represents the average precision in the range of IoU threshold from 0.5 to 0.95; represents the average precision; It represents the ratio of the number of samples correctly predicted by the algorithm to the total number of samples; Indicates the proportion of samples that are correctly predicted in positive samples; Indicates the number of positive samples that are correctly classified; Indicates the number of positive samples that are misclassified; Indicates the number of negative samples that are misclassified; Indicates the total number of categories.

10. An improved remote sensing image target detection method as claimed in claim 9, characterized in that: The improved detection model is compared with the classic model and other improved models. Specifically, in order to accurately verify the performance of the improved detection model in the remote sensing image target detection task, it is compared with a series of YOLOv5, YOLOv3, YOLOv6 and Gold-YOLO models respectively.

Citation Information

Patent Citations

  • A remote sensing image detection method, device and storage medium based on yolov5

    CN114998756B

  • SAR image layover region extraction method based on multilayer feature fusion attention mechanism

    CN113469191A

  • Lightweight remote sensing target detection method based on dense feature fusion network

    CN114596488A

  • Deep learning cancer molecular typing prediction method based on multi-scale attention fusion

    CN114841979A

  • Rolov5-based remote sensing image detection method and device and storage medium

    CN114998756A