Remote sensing image target detection method and device based on self-adaptive auxiliary head structure
By incorporating an adaptive auxiliary head structure and a convolutional recalibration multi-scale feature fusion module into the YOLO network, the problems of background clutter and scale differences in target detection of remote sensing images are solved, thereby improving the accuracy and robustness of detection.
Patent Information
- Application Number
- CN202511043633.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-28
- Publication Date
- 2025-11-07
AI Technical Summary
Remote sensing image target detection tasks suffer from problems such as cluttered backgrounds, large scale differences, and feature loss, which are difficult to solve effectively with existing technologies.
An adaptive auxiliary head structure-based remote sensing image target detection method is designed. By adding an adaptive auxiliary head structure between the neck and the detection head of the classic YOLO network structure, and employing a convolutional recalibration multi-scale feature fusion module and a spatial awareness attention module, the semantic representation of targets at different scales and the perceptual ability of the model are enhanced.
It improves the accuracy and effectiveness of remote sensing target detection, reduces information loss caused by fusion of different scales, and enhances the robustness and detection accuracy of the model in complex backgrounds.
Smart Images

Figure CN120912865A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of target detection, and relates to a remote sensing image target detection method and device based on an adaptive auxiliary head structure. BACKGROUND
[0002] In recent years, remote sensing technology has made remarkable growth in related research in many fields such as agriculture, traffic monitoring, environmental monitoring and urban planning. With the advancement of technology, today's remote sensing equipment can easily capture a large number of very high-resolution optical remote sensing images. Combined with deep learning algorithms, remote sensing technology has made some breakthroughs in information extraction, target detection and the like.
[0003] In the field of remote sensing, there are many improved algorithms of YOLO for the characteristics specific to remote sensing images. DCEF2-YOLO proposes a deformable convolution (DFConv) module and an efficient feature fusion structure to effectively detect targets in real time and maximize the use of internal feature information of the target. MSA-YOLO aims to improve the detection performance of the model and proposes a multi-scale band convolution attention mechanism (MSCAM) to reduce the introduction of background noise, combined with multi-scale features, to enhance the model's attention to different size foreground objects. TPH-YOLOv5++ adds an additional prediction head to detect small-scale objects based on YOLOv5. FFCA-YOLO proposes an efficient feature enhancement, fusion and context-aware YOLO detector. In summary, the YOLO architecture has certain scalability advantages in the task of remote sensing image target detection.
[0004] In the task of remote sensing target detection, spatial context aggregation can well cope with the change of spatial resolution. In recent years, spatial attention mechanisms have been integrated into the YOLO structure, such as TPH-YOLOv5 and YOLO-EMS, to cope with the influence of the detection network caused by the inconsistent size of the target due to different heights of the unmanned aerial vehicle. Global-local-global context-aware network (GLGCNet) is used to cope with the detection of prominent targets in complex optical remote sensing images with complex backgrounds and scale changes. An efficient local-global context aggregator (ELGCA) module is introduced for remote sensing image change detection to enhance the comprehensive background information and local spatial scale details, thereby improving the accuracy of detection.
[0005] FPN series multi-scale feature fusion structure is widely used in target detection tasks. The most original FPN structure introduces a horizontal connection from top to bottom path, creating multi-scale high-level semantic feature maps. On this basis, subsequent PAFPN, BiFPN and AFPN structures are proposed, which respectively optimize the transmission efficiency, fusion efficiency and information conflict between information in the target detection task. BSG-FPN is proposed based on the FPN framework, which skillfully combines shallow and deep feature maps and improves multi-scale information to realize target detection in remote sensing images.
[0006] Remote sensing images are affected by many factors, such as shooting conditions, target types, background environments, etc. Adaptive weight adjustment can keep the model stable performance, and enhance the robustness of the model when facing complex data images. Wang et al. proposed an adaptive feature perception method for target detection in remote sensing images to improve the learning ability of the model and reduce the influence of complex background in remote sensing images. Yang et al. proposed an adaptive reinforcement supervision distillation (ARSD) framework to solve the problem of large noise caused by complex background in remote sensing images affecting training performance, and the target size difference is very large. For high-resolution remote sensing images, Sun et al. designed a spatial adaptive feature modulation (SAFM) mechanism that can adaptively select representative features from input features for learning. The Inception series uses convolution kernels of different scales to extract features from input data, so that it can adaptively capture different features in the image.
[0007] Nowadays, target detection tasks are widely used in the field of remote sensing. Although certain research results have been achieved in target detection tasks, there are still problems such as chaotic background, large size difference, feature loss, etc. in remote sensing target detection tasks. SUMMARY
[0008] In view of the problems existing in the above-mentioned traditional method, the present application proposes a remote sensing image target detection method and device based on an adaptive auxiliary head structure, which can improve the effectiveness of remote sensing image target detection.
[0009] In order to achieve the above purpose, the embodiments of the present application adopt the following technical solutions: On the one hand, a remote sensing image target detection method based on an adaptive auxiliary head structure is provided, which comprises: Obtaining a remote sensing image.
[0010] The remote sensing image target detection network based on the adaptive auxiliary head structure is obtained by adding the adaptive auxiliary head structure between the neck and the detection head of the YOLO classic network structure, and replacing the original neck structure with the context aggregation bidirectional connection structure.
[0011] The remote sensing image target detection network based on the adaptive auxiliary head structure is obtained by adding the adaptive auxiliary head structure between the neck and the detection head of the YOLO classic network structure, and replacing the original neck structure with the context aggregation bidirectional connection structure.
[0012] In another aspect, a remote sensing image target detection device based on an adaptive auxiliary head structure is also provided, which comprises: a remote sensing image acquisition unit configured to acquire a remote sensing image; a remote sensing image target detection network construction unit configured to construct a remote sensing image target detection network based on an adaptive auxiliary head structure, wherein the remote sensing image target detection network is obtained by adding the adaptive auxiliary head structure between the neck and the detection head of the YOLO classic network structure, and replacing the original neck structure with the context aggregation bidirectional connection structure; the adaptive auxiliary head structure is used for adaptive structure design with three-layer input, and a convolutional re-scaling multi-scale feature fusion module is used to enhance the semantic representation of targets of different scales; the convolutional re-scaling multi-scale feature fusion module is used to integrate features of different scales through convolutional re-scaling fusion, and then through a branch structure, the features are extracted while the perception ability of the model is enhanced; the context aggregation bidirectional connection structure is a bidirectional FPN structure improved by using the convolutional re-scaling multi-scale feature fusion module and a spatial perception attention module; the spatial perception attention module is used to map features from two dimensions of channels and space by using different convolution branches, and balance the spatial context aggregation. a remote sensing image target detection network training and target detection unit configured to train the remote sensing image target detection network by using remote sensing images, and process a to-be-detected remote sensing image by using the trained remote sensing image target detection network to obtain a target detection result of the to-be-detected remote sensing image.
[0013] One of the above technical solutions has the following advantages and beneficial effects: The remote sensing image target detection method and device based on the adaptive auxiliary head structure have the remote sensing image target detection network based on the adaptive auxiliary head structure, which is obtained by adding the adaptive auxiliary head structure between the neck and the detection head of the YOLO classic network structure, replacing the original neck structure with the context aggregation bidirectional connection structure; the spatial perception attention module, the convolutional re-sampling multi-scale feature fusion, the context aggregation bidirectional connection structure and the adaptive auxiliary head structure are designed; the network has good spatial feature aggregation capability to retain key feature information, and integrates the adaptive weight mechanism to reduce information loss caused by different scale fusion, and performs feature refinement on the detection image, thereby improving the accuracy and effectiveness of remote sensing target detection. BRIEF DESCRIPTION OF DRAWINGS
[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0015] Figure 1 A flowchart of a remote sensing image target detection method based on an adaptive auxiliary head structure in an embodiment; Figure 2 A whole architecture diagram of a remote sensing image target detection network based on an adaptive auxiliary head structure in an embodiment; Figure 3 A block diagram of an adaptive auxiliary head structure in an embodiment; Figure 4 A basic structure diagram of a convolutional re-sampling multi-scale feature fusion module in an embodiment; Figure 5 A CABi-FPN structure diagram in an embodiment; Figure 6 A structure diagram of a spatial perception attention module in an embodiment; Figure 7 A detection result diagram of a random color remote sensing image in a DOTA v1.0 test set in an embodiment; Figure 8 A detection result diagram of a random gray remote sensing image in a DOTA v1.0 test set in an embodiment; Figure 9 A prediction result diagram of 8 randomly selected images in a DOTA v1.0 test set in an embodiment; Figure 10Fig. 1 is a schematic diagram of the visualization result of target features of three remote sensing images randomly selected from the test set of HRRSD in one embodiment. DETAILED DESCRIPTION
[0016] For the purposes of the present application, the technical solutions and advantages thereof are more clearly apparent, the following will be combined with the drawings and embodiments to make further detailed description of the present application. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not intended to limit the present application.
[0017] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in the description of the application herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the application.
[0018] It should be noted that the reference herein to "embodiments" means that the specific features, structures or characteristics described in connection with the embodiments can be included in at least one embodiment of the application. The phrase is exhibited at various places in the specification does not necessarily refer to the same embodiment, nor is it mutually exclusive or alternative to other embodiments. Those skilled in the art can understand that the embodiments described herein can be combined with other embodiments. The term "and / or" used herein refers to any combination of one or more of the associated listed items and all possible combinations, and includes these combinations.
[0019] The remote sensing image target detection network based on the adaptive auxiliary head structure constructed in the method is referred to as PHAS-YOLO.
[0020] The embodiments of the present application will be described in detail below with reference to the accompanying drawings of the embodiments of the present application.
[0021] In one embodiment, as shown in Figure 1 A remote sensing image target detection method based on an adaptive auxiliary head structure is provided, which can include the following processing steps 100 to 104: Step 100: Obtain a remote sensing image.
[0022] Step 102: constructing a remote sensing image target detection network based on an adaptive auxiliary head structure, the remote sensing image target detection network being obtained by adding an adaptive auxiliary head structure between a neck and a detection head of a YOLO classic network structure, replacing an original neck structure with a context aggregation bidirectional connection structure; the adaptive auxiliary head structure is used for adaptive structure design with three-layer input, and a convolutional re-scaling multi-scale feature fusion module is used to enhance semantic representation of different scale targets; the convolutional re-scaling multi-scale feature fusion module is used for convolutional re-scaling fusion, integrating features of different scales, and then enhancing the perception ability of the model through a branch structure while extracting features; the context aggregation bidirectional connection structure is a bidirectional FPN structure improved by the convolutional re-scaling multi-scale feature fusion module and a spatial perception attention module; the spatial perception attention module is used for feature mapping from two dimensions of channels and space by different convolution branches, balancing spatial context aggregation.
[0023] Specifically, an adaptive auxiliary head structure (Adaptive Auxiliary Head Structure, AAHS for short) is added between the original neck and the head, which can further enhance the correlation degree of semantic information between different scales.
[0024] The neck structure of the remote sensing image target detection network (PHAS-YOLO) based on the adaptive auxiliary head structure adopts a context aggregation bidirectional connection structure (Context Aggregation Bidirectional Connection Structure, CABi-FPN), which improves the perception ability of the model and reduces the information loss caused by different scale feature fusion.
[0025] The spatial perception attention module (SAAM) and the convolutional re-scaling multi-scale feature fusion module (Context Aggregation Bidirectional Connection Structure, CRMSFF for short) are two plug-and-play modules designed by the present application. The spatial perception attention module is used to balance spatial context aggregation. The SPP spatial pyramid idea is integrated into the convolutional re-scaling multi-scale feature fusion module to retain key information of features in the multi-scale fusion process, playing a key feature retention role.
[0026] The YOLO classic network can be but is not limited to YOLO v10, YOLO v5, YOLO v11, and YOLO v12 classic networks.
[0027] The overall architecture of the remote sensing image target detection network based on the adaptive auxiliary head structure is as shown in Figure 2
[0028] Step 104: training the remote sensing image target detection network by using the remote sensing image, and processing the remote sensing image to be detected by using the trained remote sensing image target detection network to obtain a target detection result of the remote sensing image to be detected.
[0029] The remote sensing image target detection method based on the adaptive auxiliary head structure has the remote sensing image target detection network based on the adaptive auxiliary head structure. The network is obtained by adding the adaptive auxiliary head structure between the neck and the detection head of the YOLO classic network structure, and replacing the original neck structure with the context aggregation bidirectional connection structure. The spatial perception attention module, the convolution rescaling multi-scale feature fusion, the context aggregation bidirectional connection structure, and the adaptive auxiliary head structure are designed. The network has good spatial feature aggregation capability to retain key feature information, and integrates the adaptive weight mechanism to reduce information loss caused by different scale fusion, and performs feature refinement on the image to be detected, thereby improving the accuracy and effectiveness of remote sensing target detection.
[0030] In one embodiment, the processing of the remote sensing image to be detected by using the trained remote sensing image target detection network in step 104 to obtain the target detection result of the remote sensing image to be detected includes: inputting the remote sensing image to be detected into the backbone network to obtain multi-scale features; inputting the multi-scale features into the context aggregation bidirectional connection structure to obtain three feature maps of different scales; inputting the three feature maps of different scales into the adaptive auxiliary head structure to obtain three semantic enhanced feature maps; and inputting the three semantic enhanced feature maps into the detection head to obtain the target detection result of the remote sensing image to be detected.
[0031] In one embodiment, as Figure 3As shown, the adaptive auxiliary head structure includes three parallel branches; the first branch includes a C2f module, a convolutional re-sampling multi-scale feature fusion module, and a convolutional layer, the second branch and the third branch each control two convolutional re-sampling multi-scale feature fusion modules and a convolutional layer; three feature maps of different scales are input into the adaptive auxiliary head structure to obtain three semantic enhanced feature maps, including: inputting a feature map of a first scale into the C2f module in the first branch to obtain a fine-grained feature; inputting a feature map of a second scale and an up-sampled feature map of a third scale into the first convolutional re-sampling multi-scale feature fusion module in the second branch to obtain a first fusion feature; inputting the feature map of the third scale and a down-sampled feature map of the second scale into the first convolutional re-sampling multi-scale feature fusion module in the third branch to obtain a second fusion feature; taking the fine-grained feature, the first fusion feature, and the second fusion feature as inputs of the convolutional re-sampling multi-scale feature fusion module of the first branch, the second convolutional re-sampling multi-scale feature fusion module of the second branch, and the third branch to obtain a third fusion feature, a fourth fusion feature, and a fifth fusion feature; processing the third fusion feature, the fourth fusion feature, and the fifth fusion feature through the convolutional layer of the first branch, the convolutional layer of the second branch, and the convolutional layer of the third branch, respectively, to obtain three semantic enhanced feature maps.
[0032] Specifically, the commonly used PAFPN and BiFPN as the neck structure of YOLO often ignore the problems of insufficient utilization of bottom features and significant semantic differences, and are difficult to apply in complex and diverse background environments of remote sensing images. In order to solve these problems, the application proposes an adaptive auxiliary head structure, as shown in Figure 3 . Figure 3 The one-to-many head in the middle is a YOLO detection detection head.
[0033] The adaptive auxiliary head structure has three layers of input, corresponding to three scales of input head features to be detected, mainly to further refine the target and enhance the correlation between semantic information. In the first layer, the feature information of Feature1 is first refined, the C2f module is used to learn fine-grained features of the feature map Feature1, locate the target region, and then enter the fusion operation. In the second layer and the third layer, the feature maps Feature2 and Feature3 directly enter the convolutional re-sampling multi-scale feature fusion module (CRMSFF) for feature fusion. This process emphasizes progressive feature fusion and fusion of bottom features, dynamically adjusts the feature map through feature fusion and processing of different levels of features, thereby realizing an adaptive mechanism, and finally transmits the data processed by each layer to the corresponding detection head.
[0034] Compared with the classic YOLO single detection head, the addition of this structure helps to reduce the error of the detection head due to the large gap in semantic information, improve the accuracy and robustness of the model, and the adaptive mechanism can make the model better adapt to the complex environment of remote sensing images.
[0035] In one embodiment, in the convolutional re-dimensioning multi-scale feature fusion module: all input feature maps are spliced to obtain spliced features; each channel of the spliced features is subjected to a global average pooling operation to obtain a vector y; after the vector y is subjected to a CRC operation and a convolutional layer processing, a CRC fused feature map is obtained; the CRC fused feature map is subjected to a segmentation processing from a channel dimension to obtain four segmented features; the four segmented features are processed using a four-branch structure, and the outputs of the four branches are step-by-step fused to obtain the output of the convolutional re-dimensioning multi-scale feature fusion module; the four-branch structure includes a compression module branch, a bottleneck structure module branch, a feature preservation branch, and a spatial pyramid pooling module branch; the compression module branch is used to realize information compression and feature extraction in the channel dimension; the bottleneck structure module branch is used to enhance the generalization ability of the model using a bottleneck structure module; the feature preservation branch does not perform special feature processing on the input feature map, and preserves the original feature, which is combined with the spatial pyramid pooling branch as a kind of residual structure; the spatial pyramid pooling module branch is used to perform pooling operations on the feature maps at different scales, and then perform splicing operations to retain more spatial information.
[0036] In one embodiment, the four segmented features are processed using a four-branch structure, and the outputs of the four branches are step-by-step fused to obtain the output of the convolutional re-dimensioning multi-scale feature fusion module, including: the output features of the compression module branch and the bottleneck structure module branch are added and fused to obtain a first fusion result; the output features of the feature preservation branch and the spatial pyramid pooling module branch are added and fused to obtain a second fusion result; the first fusion result and the second fusion result are respectively subjected to convolution and then spliced to obtain the output features of the convolutional re-dimensioning multi-scale feature fusion module.
[0037] Specifically, multi-scale feature fusion can enhance the semantic representation of targets of different scales, reduce the loss of shallow feature information, and capture more image details. The convolutional re-dimensioning multi-scale feature fusion module (CRMSFF module) proposed in the present application first integrates features of different scales through convolutional re-check (CRC) fusion, and then extracts features while enhancing the perception ability of the model through the branch structure. The basic structure of the convolutional re-dimensioning multi-scale feature fusion module is as shown in Figure 4 . Figure 4The compression module branch (x1 branch) is from reference 1 (https: / / arxiv.org / abs / 1709.01507), the bottleneck structure module (x2 branch) is the Bottleneck built into YOLO, and the feature preservation branch (x3 branch) adopts the SPP built into YOLO.
[0038] The input to the convolutional recalibration multi-scale feature fusion module is N feature maps. ,in t Indicates the input number of the first... t Each feature map For the number of channels, and These represent the height and width of the feature map, respectively. In the convolutional recalibration multi-scale feature fusion module, the first step is a CRC operation to highlight important feature channels in the remote sensing image and suppress less important feature channels. Specifically, all input feature maps... Feature maps are obtained after fusion Then, a global average pooling operation is performed on each channel to obtain the vector y. , The result of the channel compression process is calculated by linearly relating vector y, as shown in the following expression:
[0039]
[0040]
[0041]
[0042] in, The Concat operation fuses the N input feature maps to obtain... . Represents the fused feature map The Middle i The first channel, the first h line, number w The value of the column.
[0043] For the compressed vector Perform a restoration operation to help the model adaptively adjust the weights of different channels.
[0044]
[0045] Calculate the restored channel weight vector The weight vector is then reshaped and features are extracted to obtain the final feature weights. As shown in the following formula:
[0046]
[0047] wherein, represents the extension of the shape by a broadcast mechanism so that the weight vector is restored to the original spatial dimension.
[0048] Considering the problems of excessive background noise, dramatic environmental changes and significant target size differences in remote sensing images, the CRC fused feature map , the present application is first segmented from the channel dimension, and then processed respectively by using a four-branch structure. The compression module branch (x1 branch) mainly realizes information compression and feature extraction in the channel dimension, so that the model can pay more attention to the important features of the target area, rather than relying too much on background noise. The bottleneck structure module (x2 branch) aims to enhance the generalization ability of the model, so that it can be applied to the high-variation environment of remote sensing images. The feature preservation branch (x3 branch) inputs the feature map without special feature processing, preserving the original features, and combines with the spatial pyramid pooling module branch (x4 branch) as a kind of residual structure. The spatial pyramid pooling (SPP) module branch (x4 branch) performs pooling operation on the feature map at different scales, and then performs splicing operation to retain more spatial information, thereby enhancing the model's perception ability to the target. Finally, the four branches are step by step fused, and the useful features are extracted by convolution.
[0049] In one embodiment, the context aggregation bidirectional connection structure includes three spatial perception attention modules, seven convolutional re-sampling multi-scale feature fusion modules, and three convolutional layers; the input features of the context aggregation bidirectional connection structure are multi-scale features extracted by the backbone network; the multi-scale features include P2 features, P4 features, P6 features, P8 features, and P10 features output by Level2, Level4, Level6, Level8, and Level10; in the context aggregation bidirectional connection structure: the P10 features after upsampling and the P8 features are input into the first convolutional re-sampling multi-scale feature fusion module to obtain first multi-scale fusion features; the first multi-scale fusion features after upsampling and the P6 features are input into the second convolutional re-sampling multi-scale feature fusion module to obtain second multi-scale fusion features; the second multi-scale fusion features after upsampling and the P4 features are input into the third convolutional re-sampling multi-scale feature fusion module to obtain third multi-scale fusion features; the third multi-scale fusion features after upsampling and the P2 features are input into the fourth convolutional re-sampling multi-scale feature fusion module to obtain fourth multi-scale fusion features; the P4 features are input into the first spatial perception attention module to obtain first attention features; the fourth multi-scale fusion features after downsampling and the first attention features are input into the fifth convolutional re-sampling multi-scale feature fusion module to obtain a first scale feature map; the P6 features are input into the second spatial perception attention module to obtain second attention features; the fifth multi-scale fusion features after downsampling and the second attention features are input into the sixth convolutional re-sampling multi-scale feature fusion module to obtain a second scale feature map; the P8 features are input into the third spatial perception attention module to obtain third attention features; the sixth multi-scale fusion features after downsampling and the third attention features are input into the seventh convolutional re-sampling multi-scale feature fusion module to obtain a third scale feature map.
[0050] Specifically, in the target detection task of a remote sensing image, targets have scale diversity, and there is a significant difference in information contribution between feature layers. The original YOLO neck structure PAFPN lacks dynamic feature selection, which easily leads to the model being unable to fully utilize all useful feature information extracted by the backbone network. The fusion of multi-scale feature maps can effectively improve this problem, but due to the change of feature map resolution, direct fusion between different scales easily leads to information loss. To solve this problem, the present application proposes a context perception bidirectional connection structure, as shown in Figure 5
[0051] The context-aware bidirectional connection structure is different from the PAFPN and the BiFPN. In the context-aware bidirectional connection structure, the connection module adopts the convolutional re-indexing multi-scale feature fusion module and the spatial perception attention module proposed in the present application. The connection module not only adapts to different scales, but also enhances the perception ability of the model. In addition, the context-aware bidirectional connection structure reduces the connection and output, and focuses the information more concentratedly on the three output feature maps. In the cross-layer connection, the features extracted by the backbone network are integrated into the feature map of the detection region after the attention operation on the original image. At this time, the background noise not processed by the backbone network is removed, and the feature map of the detection region is integrated with more high-quality feature information, thereby effectively improving the performance of the detector. The improved neck structure not only retains the advantages of the BiFPN, but also enables the model to reduce the information loss caused by the fusion process and better cope with the challenges of target recognition.
[0052] In one embodiment, the spatial perception attention module comprises a convolution module, a spatial dimension convolution branch, and a channel dimension convolution branch; the spatial dimension convolution branch comprises a V convolution module and a QK convolution module; the P4 feature is input into the first spatial perception attention module to obtain a first attention feature, comprising: inputting the P4 feature into the convolution module to obtain a convolution feature; inputting the convolution feature into the V convolution module and the QK convolution module of the spatial dimension convolution branch respectively to obtain a V feature tensor and a QK feature tensor;
[0053]
[0054] wherein, is a linear transformation matrix corresponding to the value, is a linear transformation matrix corresponding to the query and the key, is the V feature tensor, is the QK feature tensor.
[0055] The QK feature tensor is used to weight-sum all the V feature tensors, and the convolution operation is performed on the weighted-sum result to obtain a spatial dimension output; the P4 feature is input into the channel dimension convolution branch to obtain a channel dimension output; the spatial dimension output and the channel dimension output are weighted, and the weighted result is fused with the P4 feature to obtain the first attention feature.
[0056] Specifically, in the target detection task of a remote sensing image, due to the diversification of the background, the target and the background are easily confused, thereby causing the expression ability of the feature vector to decrease and information loss to occur. In this regard, the present application proposes a spatial perception attention module to extract features of different dimensions for the remote sensing image target, capture subtle features and spatial relationships in the image, as shown in Figure 6 .
[0057] SAAM uses different convolution branches to map features in both channel and spatial dimensions. The channel dimension aims to extract the attention map of the input image, which helps the model pay more attention to important features in subsequent processing. In the spatial dimension, the vectors of Value, Query and Key are mainly obtained through convolution transformation. The QK branch calculates the similarity scores of Query and all keys, and then normalizes these scores through the softmax function to obtain the distribution of attention weights. The specific expression is as follows:
[0058]
[0059]
[0060]
[0061] where, and represent the convolution operations with the convolution kernel and ; , correspond to the linear transformation matrices of Value, Query and Key, respectively.
[0062] In the transformation process, the feature tensor achieves feature mapping, while the feature tensor is used for spatial context refinement. At the same time, the Softmax function is applied to the spatial dimension to calculate the relative importance of each position, so as to avoid the confusion between target and background as much as possible during recognition.
[0063] Finally, the feature tensor is used to weight-sum all , so as to obtain the context vector of each query position, i.e. the output in the spatial dimension. Then it is fused with the channel branch output to obtain the final output:
[0064]
[0065] where, is the value of any position in the attention map in the channel dimension. is the multiplication of and using a reweighting matrix, aiming to aggregate the extracted features in the input image in the spatial dimension. is the convolution operation with the convolution kernel , which is used for feature extraction to obtain the output in the spatial dimension The spatial dimension output is weighted with the channel dimension output to realize key feature fusion. The weighted result is fused with the input image to obtain the final feature map. The difference from the input image is that the image information is further extracted from the complex environment.
[0066] In one embodiment, the convolution module includes a convolution layer with a convolution kernel of and a convolution layer with a convolution kernel of .
[0067] In some embodiments, experimental examples are also provided, experiments are carried out on the training set, validation set and test set of DIOR, DOTAv1.0 and HRRSD data sets respectively, and the method is compared with 6 methods (YOLO v8n, YOLO v10n, TPH-YOLOv5++, FFCA-YOLO, MAF-YOLO, YOLO v11n). The results in Table 1, Table 2 and Table 3 are the reproducible results of each model on the three data sets. Except for batch and image size, the parameters used by the comparison models are based on the original parameters. The bold values in the table represent the best results in this indicator.
[0068] Table 1 Comparison results on DIOR data set (%)
[0069] Table 1 shows the experimental results of PHAS-YOLO and SOTA model on DIOR data set. Compared with other methods, the method obtains the highest detection accuracy of 89.3% on DIOR data set. And the mAP50:95 data performs best, reaching 66.8%. Compared with baseline YOLO v10n, the mAP50 and mAP50:95 of the method are improved by 0.6% and 2.3% respectively. Compared with other models, it also improves in different degrees. These results show that the performance of PHAS-YOLO is more superior. Especially need to point out that the obvious improvement of mAP50:95 shows that the method has better positioning accuracy and higher detection quality in the task of remote sensing target detection.
[0070] Table 2 Comparison results on DOTAv1.0 data set (%)
[0071] The comparison results on the DOTAv1.0 dataset are shown in Table 2. The four index parameters of the present method all perform well. Compared with the baseline model YOLO v10n, the mAP50 and mAP50:95 of the present method are increased by 4.6% and 6.5%, respectively. Compared with other models, the improvement of the present method is more obvious, indicating that the present method has a significant advantage in detection effect.
[0072] Table 3 Comparison results on the HRRSD dataset (%)
[0073] As shown in Table 3, the results of the present method on the HRRSD dataset show that the mAP50 and mAP50:95 of PHAS-YOLO are optimal, which are increased by 1.6% and 3.5% compared with baseline, respectively, and are also significantly better than other models, indicating that the model has stronger positioning ability and can accurately distinguish positive and negative samples.
[0074] (1) Comparison on each class On the DOTAv1.0 and HRRSD datasets, the mAP50 index of all models on each class in each dataset is compared, and the results are shown in Tables 4 and 5. It can be seen that the average mAP50 of all classes of PHAS-YOLO is increased by 5.6% and 0.6% compared with baseline (YOLOv10), respectively. For each class, it can be found that the present method performs best in 8 classes in DOTAv1.0 and 7 classes in HRRSD. These results show that the present method can effectively enhance the performance of remote sensing image target detection.
[0075] Table 4 Comparison results on each class on the DOTAv1.0 dataset (%)
[0076] Table 5 Comparison results on each class on the HRRSD dataset (%)
[0077] (2) Ablation experiment results To analyze the importance of each module in PHAS-YOLO, the embodiment evaluates six different models from multiple perspectives. The embodiment uses mAP50, mAP50:95 as the main indicators, and Precision, Recall as auxiliary reference indicators. The embodiment gradually applies AAHS, CABi-FPN, and SAAM in the baseline task to verify their effectiveness, and conducts ablation experiments on DIOR, DOTAv1.0, and HRRSD three datasets respectively. Tables 6 to 8 show the influence of increasing or decreasing each module on the evaluation indicators, wherein indicates that the module is used, indicates that the module is not used, and the None option in the table indicates that the SAAM module is not added in CABi-FPN.
[0078] Table 6 Ablation experiment (%) in DIOR dataset
[0079] Table 7 Ablation experiment (%) in DOTA v1.0 dataset
[0080] Table 8 Ablation experiment (%) in HRRSD dataset
[0081] From Tables 6 to 8, it can be seen that: 1) AAHS: As shown in Table 7, adding AAHS structure can significantly improve all evaluation indicators, especially mAP50 (from 0.655 to 0.671) and mAP50:95 (from 0.43 to 0.451). Even in DIOR (Table 6) and HRRSD (Table 8) datasets, the recall rate is slightly lacking, but the other three indicators are improved to some extent, especially the precision (DIOR: from 0.893 to 0.906, HRRSD: from 0.939 to 0.951). This confirms that adding an auxiliary structure of a head can improve the detection effect of the detector, so that the model can accurately distinguish between positive and negative samples.
[0082] 2) CABi-FPN (None): As shown in Tables 6-8, adding CABi-FPN (None) can improve the evaluation indicators, especially in the two indicators of mAP50 and mAP50:95, among which it is particularly evident in the DOTAv1.0 dataset, mAP50 from 0.655 to 0.686, and mAP50:95 from 0.43 to 0.466. These data show that the CABi-FPN designed in the application has obvious performance improvement compared with the original PAFPN neck structure. In the remote sensing target detection task, it can reduce the interference of complex background and multiple target overlaps.
[0083] 3) CABi-FPN (SAAM): This module mainly verifies the role of the SAAM module in the CABi-FPN neck structure. As can be seen from Table 6, the recall rate of CABi-FPN with SAAM (R = 81.8%) is better than that without SAAM (R = 81%), which makes up for the missed detection problem caused by complex background. In the case of adding auxiliary head structure, it is also found that in the DOTAv1.0 and HRRSD two datasets, the four indicators of adding SAAM are higher than those without adding SAAM, and mAP50 can be increased by 3.6% at most, and mAP50:95 can be increased by 3.6% at most.
[0084] (3) Ablation experiment on the module In order to verify the plug-and-play nature of SAAM and CRMSFF modules, this embodiment carries out ablation experiments on the two modules corresponding to the attention mechanism and the connection module.
[0085] Table 9 shows the performance comparison of spatial perception attention mechanism (SAAM) and three common attention mechanisms. It can be seen that the mAP50 and mAP50:95 of SAAM are 87.1% and 63.9% respectively, which are the best, indicating that this attention module has stronger feature expression ability.
[0086] Table 9 Ablation experiment of SAAM in DIOR dataset (%)
[0087] Table 10 shows that compared with the Concat module provided by the torch library, the mAP50 and mAP50:95 of the CRMSFF module proposed in this paper are improved by 0.1% and 1.0% respectively, which can indicate that the CRMSFF module can enhance the expression ability of the model. At the same time, the accuracy is improved by 0.6%, which shows that to some extent, this module can improve the ability of the model in the detection task.
[0088] Table 10 Ablation experiment of CRMSFF in DIOR dataset (%)
[0089] In summary, SAAM and CRMSFF are two high-precision plug-and-play modules that can improve the robustness of the model to a certain extent and ensure the accuracy of target detection in remote sensing target detection tasks.
[0090] (4) Result visualization 1) Detection result visualization based on the most advanced method In order to intuitively analyze the performance of PHAS-YOLO, the model is compared with 9 advanced models in the final prediction results, and 2 prediction pictures are randomly selected from the DOTAv1.0 test set, representing color images and grayscale images respectively. Figure 7 、 Figure 8 The results show that the method can detect more targets, and compared with the other six models, it is more prone to miss detection and false detection. At the same time, the PHAS-YOLO model has higher confidence for the same target, which proves that the network proposed in this paper performs better in remote sensing target detection tasks. Among them, the red dotted box in the matrix represents false detection, and the red dotted oval represents missed detection.
[0091] (2) Baseline detection result visualization Figure 9 The visualization results of the proposed method compared with the baseline network YOLO v10 are shown, where the red dotted box represents false detection, and the red dotted oval represents missed detection. It can be seen that in Figure 9 , the results of the first and second rows show that YOLOv10 has 3 false detections and 2 missed detections, while PHAS-YOLO has no false detection. In Figure 9 , the results of the third and fourth rows show that PHAS-YOLO has fewer false detections and missed detections than YOLO v10, indicating that the PHAS-YOLO model can further extract feature information of remote sensing targets and reduce false detection and missed detection.
[0092] (3) Target feature heat map visualization Figure 10 The visualization results of the target features of the proposed method and the baseline network YOLO v10 and RT-DETR series models are shown. The 3 images of target feature visualization come from the test set of HRRSD. The brighter color area indicates that the model pays more attention to that area. In Figure 10It can be seen that, regardless of the size of the target size, PHAS-YOLO pays more attention to the actual target and less attention to the background. However, for YOLO v10, it can be found that the background part is also bright in the small target feature visualization part, which indicates that YOLO v10 is prone to confusion between the target and the background in the remote sensing task. In addition, YOLO v10 is prone to pay more attention to the background for larger scale targets. The RT-DETR series of models do not pay attention to the target features in the remote sensing image. Compared with the YOLO series, the RT-DETR has certain limitations in the remote sensing target detection task.
[0093] It should be understood that, although the above process Figure 1 The steps are displayed in sequence according to the direction of the arrow, but these steps are not necessarily executed in the order indicated by the arrow. Unless otherwise specified herein, the execution of these steps is not strictly limited in order, and these steps can be executed in other orders. Moreover, at least part of the steps of the above process Figure 1 The steps can include multiple sub-steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be alternately executed with at least part of other steps or sub-steps or stages of other steps.
[0094] In one embodiment, a remote sensing image target detection device based on an adaptive auxiliary head structure is also provided, and the device comprises: A remote sensing image acquisition unit is configured to acquire a remote sensing image.
[0095] A remote sensing image target detection network construction unit is configured to construct a remote sensing image target detection network based on an adaptive auxiliary head structure. The remote sensing image target detection network is obtained by adding an adaptive auxiliary head structure between the neck and the detection head of a YOLO classic network structure, and replacing the original neck structure with a context aggregation bidirectional connection structure. The adaptive auxiliary head structure is configured to adopt a three-layer input adaptive structure design, and adopt a convolution rescaling multi-scale feature fusion module to enhance the semantic representation of different scale targets. The convolution rescaling multi-scale feature fusion module is configured to integrate features of different scales through convolution rescaling fusion, and then enhance the perception ability of the model through a branch structure while extracting features. The context aggregation bidirectional connection structure is a bidirectional FPN structure improved by a convolution rescaling multi-scale feature fusion module and a spatial perception attention module. The spatial perception attention module is configured to perform feature mapping from two dimensions of channels and space using different convolution branches, and balance spatial context aggregation.
[0096] The remote sensing image target detection network training and target detection unit is configured to train a remote sensing image target detection network using remote sensing images, and to process a to-be-detected remote sensing image using the trained remote sensing image target detection network to obtain a target detection result of the to-be-detected remote sensing image.
[0097] In one embodiment, the remote sensing image target detection network training and target detection unit is further configured to input the to-be-detected remote sensing image into a backbone network to obtain multi-scale features, input the multi-scale features into a context aggregation bidirectional connection structure to obtain three feature maps of different scales, input the three feature maps of different scales into an adaptive auxiliary head structure to obtain three semantic enhanced feature maps, and input the three semantic enhanced feature maps into a detection head to obtain the target detection result of the to-be-detected remote sensing image.
[0098] In one embodiment, the adaptive auxiliary head structure includes three parallel branches, the first branch includes a C2f module, a convolutional reparameterization multi-scale feature fusion module, and a convolutional layer, the second branch and the third branch each include two convolutional reparameterization multi-scale feature fusion modules and one convolutional layer, and the remote sensing image target detection network training and target detection unit is further configured to input the first-scale feature map into the C2f module in the first branch to obtain a fine-grained feature, input the second-scale feature map and the up-sampled third-scale feature map into the first convolutional reparameterization multi-scale feature fusion module in the second branch to obtain a first fused feature, input the third-scale feature map and the down-sampled second-scale feature map into the first convolutional reparameterization multi-scale feature fusion module in the third branch to obtain a second fused feature, and use the fine-grained feature, the first fused feature, and the second fused feature as inputs of the convolutional reparameterization multi-scale feature fusion module in the first branch, the second convolutional reparameterization multi-scale feature fusion module in the second branch, and the third convolutional reparameterization multi-scale feature fusion module in the third branch to obtain a third fused feature, a fourth fused feature, and a fifth fused feature, and process the third fused feature, the fourth fused feature, and the fifth fused feature through the convolutional layer in the first branch, the convolutional layer in the second branch, and the convolutional layer in the third branch, respectively, to obtain the three semantic enhanced feature maps.
[0099] In an embodiment, in the convolutional re-scaling multi-scale feature fusion module in the remote sensing image target detection network construction module: all input feature maps are spliced to obtain spliced features; each channel of the spliced features is subjected to a global average pooling operation to obtain a vector y; the vector y is subjected to a CRC operation and a convolutional layer processing to obtain a CRC fused feature map; the CRC fused feature map is segmented in the channel dimension to obtain four segmented features; the four segmented features are processed using a four-branch structure, and the outputs of the four branches are step-by-step fused to obtain the output of the convolutional re-scaling multi-scale feature fusion module; the four-branch structure includes a compression module branch, a bottleneck structure module branch, a feature preservation branch, and a spatial pyramid pooling module branch; the compression module branch is used to realize information compression and feature extraction in the channel dimension; the bottleneck structure module branch is used to enhance the generalization ability of the model using a bottleneck structure module; the feature preservation branch does not perform special feature processing on the input feature map, and preserves the original feature, which is combined with the spatial pyramid pooling branch as a residual-like structure; the spatial pyramid pooling module branch is used to perform pooling operations on the feature maps at different scales, and then perform splicing operations to retain more spatial information.
[0100] In an embodiment, the four segmented features are processed using a four-branch structure, and the outputs of the four branches are step-by-step fused to obtain the output of the convolutional re-scaling multi-scale feature fusion module, including: the output features of the compression module branch and the bottleneck structure module branch are added and fused to obtain a first fusion result; the output features of the feature preservation branch and the spatial pyramid pooling module branch are added and fused to obtain a second fusion result; the first fusion result and the second fusion result are respectively convolved and then spliced to obtain the output features of the convolutional re-scaling multi-scale feature fusion module.
[0101] In one embodiment, the context aggregation bidirectional connection structure includes three spatial perception attention modules, seven convolutional re-sampling multi-scale feature fusion modules, and three convolutional layers; the input features of the context aggregation bidirectional connection structure are multi-scale features extracted by the backbone network; the multi-scale features include P2 features, P4 features, P6 features, P8 features, and P10 features output by Level2, Level4, Level6, Level8, and Level10; in the context aggregation bidirectional connection structure: the P10 features and the P8 features are input into the first convolutional re-sampling multi-scale feature fusion module after being up-sampled, to obtain first multi-scale fusion features; the first multi-scale fusion features and the P6 features are input into the second convolutional re-sampling multi-scale feature fusion module after being up-sampled, to obtain second multi-scale fusion features; the second multi-scale fusion features and the P4 features are input into the third convolutional re-sampling multi-scale feature fusion module after being up-sampled, to obtain third multi-scale fusion features; the third multi-scale fusion features and the P2 features are input into the fourth convolutional re-sampling multi-scale feature fusion module after being up-sampled, to obtain fourth multi-scale fusion features; the P4 features are input into the first spatial perception attention module, to obtain first attention features; the fourth multi-scale fusion features and the first attention features are input into the fifth convolutional re-sampling multi-scale feature fusion module after being down-sampled, to obtain a first scale feature map; the P6 features are input into the second spatial perception attention module, to obtain second attention features; the fifth multi-scale fusion features and the second attention features are input into the sixth convolutional re-sampling multi-scale feature fusion module after being down-sampled, to obtain a second scale feature map; the P8 features are input into the third spatial perception attention module, to obtain third attention features; the sixth multi-scale fusion features and the third attention features are input into the seventh convolutional re-sampling multi-scale feature fusion module after being down-sampled, to obtain a third scale feature map.
[0102] In one embodiment, the spatial perception attention module includes a convolutional module, a spatial dimension convolution branch, and a channel dimension convolution branch; the spatial dimension convolution branch includes a V convolutional module and a QK convolutional module; inputting the P4 features into the first spatial perception attention module to obtain the first attention features includes: inputting the P4 features into the convolutional module to obtain convolutional features; inputting the convolutional features into the V convolutional module and the QK convolutional module of the spatial dimension convolution branch respectively to obtain a V feature tensor and a QK feature tensor;
[0103]
[0104] wherein, is a linear transformation matrix corresponding to the value, is a linear transformation matrix corresponding to the query and the key, V feature tensor, QK feature tensor; The V feature tensors are weighted and summed by using the QK feature tensor, convolution operation is performed on the weighted and summed result, and a spatial dimension output is obtained; the P4 feature is input into a channel dimension convolution branch, and a channel dimension output is obtained; the spatial dimension output and the channel dimension output are weighted, and the weighted result is fused with the P4 feature, and a first attention feature is obtained.
[0105] In one embodiment, the convolution module includes a convolution layer with a convolution kernel of and a convolution layer with a convolution kernel of .
[0106] It can be understood that the specific explanations and descriptions of the remote sensing image target detection device based on the adaptive auxiliary head structure can refer to the corresponding explanations and descriptions of the remote sensing image target detection method based on the adaptive auxiliary head structure in the above embodiments, and will not be repeated here. Each module in the above remote sensing image target detection device based on the adaptive auxiliary head structure can be realized by software, hardware and combinations thereof. The above modules can be embedded in or independent of a device with data processing function in hardware form, or can be stored in the memory of the device in software form, so that the processor can call and execute the operations corresponding to each module. The device can be, but is not limited to, various types of data processing computer devices in the art.
[0107] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, all possible combinations of the technical features in the above embodiments are not described, but as long as the combinations of the technical features do not exist, they should be considered as the scope of the present application.
[0108] The above embodiments only express several implementation manners of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the protection scope of the present application. It should be pointed out that for ordinary skilled in the art, without departing from the concept of the present application, some modifications and improvements can be made, which all belong to the protection scope of the present application.
Claims
1. A remote sensing image target detection method based on adaptive auxiliary head structure, characterized in that, The method comprises: acquiring a remote sensing image; constructing a remote sensing image target detection network based on an adaptive auxiliary head structure, wherein the remote sensing image target detection network is obtained by adding an adaptive auxiliary head structure between a neck and a detection head of a YOLO classic network structure, replacing an original neck structure with a context aggregation bidirectional connection structure; the adaptive auxiliary head structure is used for adaptive structure design with three layers of input, and a convolutional re-scaling multi-scale feature fusion module is used to enhance semantic representation of different scale targets; the convolutional re-scaling multi-scale feature fusion module is used for convolutional re-scaling fusion to integrate features of different scales, and then through a branch structure, the model's perception ability is enhanced while extracting features; the context aggregation bidirectional connection structure is a bidirectional FPN structure improved by a convolutional re-scaling multi-scale feature fusion module and a spatial perception attention module; the spatial perception attention module is used for feature mapping from two dimensions of channels and space by different convolution branches, and balances spatial context aggregation; training the remote sensing image target detection network by using the remote sensing image, and processing a to-be-detected remote sensing image by using the trained remote sensing image target detection network to obtain a target detection result of the to-be-detected remote sensing image. 2.The remote sensing image target detection method based on adaptive auxiliary head structure according to claim 1, wherein, processing a to-be-detected remote sensing image by using the trained remote sensing image target detection network to obtain a target detection result of the to-be-detected remote sensing image, comprising: inputting the to-be-detected remote sensing image into a backbone network to obtain multi-scale features; inputting the multi-scale features into a context aggregation bidirectional connection structure to obtain three feature maps of different scales; inputting the three feature maps of different scales into the adaptive auxiliary head structure to obtain three semantic enhanced feature maps; inputting the three semantic enhanced feature maps into a detection head to obtain a target detection result of the to-be-detected remote sensing image. 3.The method of claim 2, wherein, The adaptive auxiliary head structure comprises three parallel branches; a first branch comprises a C2f module, a convolutional re-scaling multi-scale feature fusion module, and a convolutional layer, and a second branch and a third branch each comprise two convolutional re-scaling multi-scale feature fusion modules and a convolutional layer; inputting the three feature maps of different scales into the adaptive auxiliary head structure to obtain three semantic enhanced feature maps, comprising: inputting a first scale feature map into a C2f module in the first branch to obtain a fine-grained feature; inputting a second scale feature map and an up-sampled third scale feature map into a first convolutional re-scaling multi-scale feature fusion module in the second branch to obtain a first fusion feature; inputting a third scale feature map and a down-sampled second scale feature map into a first convolutional re-scaling multi-scale feature fusion module in the third branch to obtain a second fusion feature; inputting the fine-grained feature, the first fusion feature, and the second fusion feature as inputs of a convolutional re-scaling multi-scale feature fusion module of the first branch, a second convolutional re-scaling multi-scale feature fusion module of the second branch, and a second convolutional re-scaling multi-scale feature fusion module of the third branch to obtain a third fusion feature, a fourth fusion feature, and a fifth fusion feature; The third fusion feature, the fourth fusion feature and the fifth fusion feature are respectively processed by a convolution layer of the first branch, a convolution layer of the second branch and a convolution layer of the third branch to obtain three semantic enhanced feature maps. 4.The method of claim 1, wherein, In the convolutional re-sampling multi-scale feature fusion module: all the input feature maps are spliced to obtain spliced features; each channel of the spliced features is subjected to a global average pooling operation to obtain a vector y; the vector y is subjected to a CRC operation and a convolution layer processing to obtain a CRC fused feature map; the CRC fused feature map is segmented in the channel dimension to obtain four segmented features; the four segmented features are processed by a four-branch structure, and the outputs of the four branches are step-by-step fused to obtain the output of the convolutional re-sampling multi-scale feature fusion module; the four-branch structure includes a compression module branch, a bottleneck structure module branch, a feature preservation branch and a spatial pyramid pooling module branch; the compression module branch is used to realize information compression and feature extraction in the channel dimension; the bottleneck structure module branch is used to enhance the generalization ability of the model by using a bottleneck structure module; the feature preservation branch does not perform special feature processing on the input feature map, and retains the original feature, which is combined with the spatial pyramid pooling branch as a residual-like structure; the spatial pyramid pooling module branch is used to perform pooling operations on the feature maps in different scales, and then perform splicing operations to retain more spatial information.
5. The method of claim 4, wherein the method further comprises: The four segmented features are processed by a four-branch structure, and the outputs of the four branches are step-by-step fused to obtain the output of the convolutional re-sampling multi-scale feature fusion module, including: the output features of the compression module branch and the bottleneck structure module branch are added and fused to obtain a first fusion result; the output features of the feature preservation branch and the spatial pyramid pooling module branch are added and fused to obtain a second fusion result; the first fusion result and the second fusion result are respectively convolved and spliced to obtain the output feature of the convolutional re-sampling multi-scale feature fusion module. 6.The method of claim 1, wherein, The context aggregation bidirectional connection structure includes three spatial perception attention modules, seven convolutional re-sampling multi-scale feature fusion modules and three convolution layers. The input features of the context aggregation bidirectional connection structure are multi-scale features extracted by the backbone network; the multi-scale features include P2 features, P4 features, P6 features, P8 features and P10 features output by Level2, Level4, Level6, Level8 and Level10; In the context aggregation bidirectional connection structure: the P10 features are upsampled and input into the first convolutional re-sampling multi-scale feature fusion module to obtain a first multi-scale fusion feature; the first multi-scale fusion feature is upsampled and input into the second convolutional re-sampling multi-scale feature fusion module with the P6 features to obtain a second multi-scale fusion feature; the second multi-scale fusion feature is upsampled and input into the third convolutional re-sampling multi-scale feature fusion module with the P4 features to obtain a third multi-scale fusion feature; The third multi-scale fusion feature is up-sampled and input into a fourth convolutional re-sampling multi-scale feature fusion module together with the P2 feature to obtain a fourth multi-scale fusion feature; The P4 feature is input into a first spatial perception attention module to obtain a first attention feature; The fourth multi-scale fusion feature is down-sampled and input into a fifth convolutional re-sampling multi-scale feature fusion module together with the first attention feature to obtain a first scale feature map; The P6 feature is input into a second spatial perception attention module to obtain a second attention feature; The fifth multi-scale fusion feature is down-sampled and input into a sixth convolutional re-sampling multi-scale feature fusion module together with the second attention feature to obtain a second scale feature map; The P8 feature is input into a third spatial perception attention module to obtain a third attention feature; The sixth multi-scale fusion feature is down-sampled and input into a seventh convolutional re-sampling multi-scale feature fusion module together with the third attention feature to obtain a third scale feature map.
7. The method of claim 1, wherein the method further comprises: The spatial perception attention module comprises a convolution module, a spatial dimension convolution branch and a channel dimension convolution branch; the spatial dimension convolution branch comprises a V convolution module and a QK convolution module; The P4 feature is input into a first spatial perception attention module to obtain a first attention feature, comprising: The P4 feature is input into the convolution module to obtain a convolution feature; The convolution feature is input into the V convolution module and the QK convolution module of the spatial dimension convolution branch respectively to obtain a V feature tensor and a QK feature tensor; wherein, is a linear transformation matrix corresponding to the values, is a linear transformation matrix corresponding to the query and the key, is a V eigentensor, is a QK eigentensor; The QK feature tensor is used to perform weighted summation on all the V feature tensors, and the weighted summation result is subjected to convolution operation to obtain a spatial dimension output; The P4 feature is input into the channel dimension convolution branch to obtain a channel dimension output; The spatial dimension output and the channel dimension output are subjected to weighted processing, and the weighted result is fused with the P4 feature to obtain the first attention feature. 8.The method of claim 1, wherein, The convolution module comprises a convolution layer with a convolution kernel of and a convolution layer with a convolution kernel of .
9. A remote sensing image target detection device based on adaptive auxiliary head structure, characterized in that, The device comprises: a remote sensing image acquisition unit configured to acquire a remote sensing image; a remote sensing image target detection network construction unit configured to construct a remote sensing image target detection network based on an adaptive auxiliary head structure, wherein the remote sensing image target detection network is obtained by adding an adaptive auxiliary head structure between a neck and a detection head of a YOLO classic network structure, and replacing an original neck structure with a context aggregation bidirectional connection structure; the adaptive auxiliary head structure is configured to adopt a three-layer input adaptive structure design, and adopt a convolutional re-sampling multi-scale feature fusion module to enhance semantic representation of different scale targets; the convolutional re-sampling multi-scale feature fusion module is configured to integrate features of different scales through convolutional re-checking fusion, and then enhance perception ability of the model while extracting the features through a branch structure; the context aggregation bidirectional connection structure is a bidirectional FPN structure improved by the convolutional re-sampling multi-scale feature fusion module and a spatial perception attention module; the spatial perception attention module is configured to map features from two dimensions of a channel and a space by different convolution branches, and balance spatial context aggregation; The remote sensing image target detection network training and target detection unit is configured to train the remote sensing image target detection network by using remote sensing images, and process a to-be-detected remote sensing image by using the trained remote sensing image target detection network to obtain a target detection result of the to-be-detected remote sensing image.
Citation Information
Cited By
Power plant type detection method based on deep learning
CN121353809A
A power plant type detection method based on deep learning
CN121353809B