Small target detection network based on sparse feature enhancement fusion
Through the sparse feature enhancement module and a progressive gradient enhancement detection head, the network performance is optimized, and the feature extraction difficulties and information loss problems of small object detection in remote sensing images are solved, and high-precision and efficient small object detection are achieved.
Patent Information
- Application Number
- CN202510595191.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-08-15
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art has problems such as difficulty in extracting target feature, loss of information, insufficient fusion, and insufficient gradient information in small-scale object detection tasks in visible light remote sensing images, resulting in insufficient detection accuracy.
The sparse feature enhancement module, hierarchical feature fusion module and progressive gradient enhancement detection head are adopted. Through sparse convolution, global feature focus enhancement module and hierarchical enhancement fusion module, combined with the programmable gradient information mechanism, network performance is optimized and detection accuracy is improved.
It significantly improves the accuracy and accuracy of small object detection, enhances the target distinction ability in complex backgrounds, reduces the calculation amount and time overhead, and improves the real-time and detection efficiency of the model.
Smart Images

Figure CN120495922A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of small target detection in remote sensing images, and in particular to a small target detection network based on sparse feature enhancement and fusion. Background Art
[0002] The main difficulty in detecting small targets in visible light remote sensing images lies in how to effectively extract target features. The inherent blurriness of visible light remote sensing images, the scarcity of detailed information about tiny targets, the lack of prominent appearance features, and the low degree of discrimination from the background all pose challenges to the task of detecting small targets. The traditional solution is to improve the detection accuracy of the model by introducing additional shallow features and additional shallow information about the target in each layer. Although this method improves the detection performance of the model to a certain extent, it will bring about difficulties in feature extraction of small targets, loss or insufficient fusion of small target information, insufficient feature extraction in small target detection, and insufficient gradient information, which limits the application of the model in this scenario. Summary of the Invention
[0003] The purpose of the present invention is to provide a small target detection network based on sparse feature enhancement and fusion. By proposing a sparse feature enhancement module and a hierarchical feature fusion module, a progressive gradient enhancement detection head is designed, and a programmable gradient information mechanism is introduced to improve and optimize the detection head to solve the above problems.
[0004] The technical solution adopted by a small target detection network based on sparse feature enhancement fusion disclosed in the present invention is:
[0005] A small target detection network based on sparse feature enhancement fusion, including:
[0006] Backbone network, extracts multi-level features of different scales through the backbone network;
[0007] The sparse feature enhancement module performs deconvolution operations on shallow features to maximize image detail recovery. By integrating features from global to local, it further improves the detection accuracy of small targets and enhances the ability to distinguish small targets in complex backgrounds.
[0008] The progressive gradient enhancement detection head effectively optimizes the performance of the network and improves the detection effect by dynamically adjusting the way information is propagated during training. It significantly improves the detection accuracy without adding additional parameters, ensuring that the network can accurately detect targets and perform accurate classification.
[0009] As an optimal solution, the sparse feature enhancement module includes sparse convolution, global feature focus enhancement module and hierarchical enhancement fusion module. First, the multi-level features of the backbone network are processed by sparse convolution and then fused and passed as input to the global feature focus enhancement module to further enrich the low-level information in the shallow features. Then, the hierarchical enhancement fusion module is used to progressively fuse the features of different levels layer by layer. Finally, the enhanced fusion features are input into the progressive gradient enhancement detection head for final inference.
[0010] As a preferred solution, the sparse convolution is a sparse convolution improved based on depthwise separable convolution. Depthwise separable convolution splits the standard convolution into two independent operations: depthwise convolution and pointwise convolution. Depthwise separable convolution first performs convolution operations on each input channel independently using convolution kernels of different sizes, and realizes the fusion between channels through pointwise convolution. For each input channel C, the mathematical expression of depthwise convolution is as follows:
[0011]
[0012] Among them, X i+m,j+n,c is the pixel value of the input image in channel c, W m,n,c is the weight of the corresponding convolution kernel, and the weight of each input channel is independent;
[0013] Point-by-point convolution performs information fusion on the output features of depth convolution through a 1×1 convolution operation, integrating and merging the features of multiple channels. The formula is as follows:
[0014]
[0015] Among them, X i,j,c is the value of the output feature map after depth convolution in channel c, W 1,1,c,k It is the weight of the 1×1 convolution kernel, which is used to fuse information from different channels. The final output feature map has a value of Y in the kth channel. i,j,k By decomposing the standard convolution into these two steps, the depth-wise separable convolution significantly reduces the amount of computation, especially when the number of channels is large, which can significantly improve the computational efficiency.
[0016] As a preferred solution, a sparsification mechanism is further introduced to expand and optimize the convolution operation, further reducing the amount of calculation in the convolution operation. Through the mask mechanism, some non-critical weights are discarded, while ensuring the effective extraction of key information and eliminating redundant calculations. Sparse convolution can be represented by a binary mask matrix. Assuming that the sparse mask matrix is S, the calculation formula of sparse convolution is as follows:
[0017]
[0018] Among them, S m,n,c is a binary sparse mask that determines which weights are set to zero. If S m,n,c The value is 0, indicating that the convolution at the corresponding position will not participate in the calculation.
[0019] Similarly, in the process of point-by-point convolution, a sparsity mechanism is introduced to optimize the amount of calculation. When performing point-by-point convolution, the channel sparse mask S is used to optimize the amount of calculation. c To filter the channels involved in information fusion, the calculation formula of sparse point-by-point convolution is as follows:
[0020]
[0021] Among them, S c It is a sparse mask that determines which channels participate in the point-by-point convolution calculation and takes a value of 0 or 1. 1,1,c,k is the weight of the point-by-point convolution.
[0022] As a preferred solution, the global feature enhancement module focuses on and optimizes the fine features of small targets, enabling the network to better handle small target detection tasks in complex backgrounds. The global feature enhancement module provides a flexible and efficient mechanism that can improve the ability of feature expression through the interactive fusion of features at different levels, thereby significantly improving the detection performance of small targets.
[0023] As a preferred solution, the hierarchical enhancement fusion module aims to effectively fuse multi-level features. The feature fusion module is improved by using an inverted residual structure combined with a cascade mechanism:
[0024] The features are processed using depthwise separable convolution operations to greatly reduce the amount of computation and retain the original information through residual links. To further improve the efficiency of feature fusion, a cascaded group attention mechanism is used to decompose the attention calculation into multiple levels of local calculations and gradually enhance the feature information through the cascade structure. The working process is as follows:
[0025]
[0026] Among them, Q and K are the query and key matrices after grouping, respectively, and K T The key matrix is transposed. The attention results of each group are transmitted and fused layer by layer in a cascade manner. The output of each layer not only depends on the calculation results of the previous layer, but also further integrates the features of the previous layer. The process is as follows:
[0027]
[0028] Where G is the number of groups, V i Represents the value matrix of each group, representing the feature information of the feature map,
[0029] Through the above optimization, while reducing the amount of computation and memory consumption, the recognizability of small targets is enhanced, the information transmission strategy is optimized, and the performance of the model is effectively improved.
[0030] As a preferred solution, the progressive gradient enhancement detection head enhances the gradient flow during training by adding two key modules: multi-level auxiliary information and auxiliary reversible branches, thereby reducing information loss in back propagation and improving the performance of the detection network.
[0031] As a preferred solution, the multi-level auxiliary information is used to integrate features of different scales at different levels, so that low-level information of the image can be retained in feature maps at higher levels, thus achieving the complementarity of appearance information and semantic information:
[0032] Among them, F MAI is the fused multi-level feature map, W i is the weight of features at different levels, F i Represents feature maps at different levels.
[0033] As a preferred solution, the auxiliary reversible branch ensures the effectiveness and stability of information transmission in the network, allowing the network to better learn accurate features, the training process can converge faster, and the gradient flow is maintained through skip connections, reverse calculations, and other behaviors;
[0034] When the main branch processes the feature map through the network module, the auxiliary branch also processes the input data in parallel to obtain the output feature map Z of the auxiliary branch. arb , in the back-propagation phase, the auxiliary branch improves the gradient of the main network through reverse operation:
[0035]
[0036] Among them, L is the loss function, G arb It is the gradient of the auxiliary branch. After combining the gradient of the main branch with the gradient of the auxiliary branch, the main network model parameters are updated;
[0037]
[0038] Among them, η is the learning rate, G update is the updated gradient after the main branch and auxiliary branch are combined, θ old is the parameter of the network model before updating, θ new are the updated parameters of the network model in this iteration:
[0039] The auxiliary reversible branch significantly improves the stability of gradient transfer by injecting additional robust gradient information into the main branch, making the update efficiency of each model iteration higher, thereby improving the overall training speed and performance of the network.
[0040] The beneficial effects of a small target detection network based on sparse feature enhancement and fusion disclosed by the present invention are as follows: in response to the difficulty of feature extraction in YOLOv8 and the problem of information loss during the transmission process, a sparse feature enhancement module is designed to effectively extract low-level features of the target by utilizing spatial sparsity, and to enhance low-level features by focusing on global features to highlight important areas and suppress background noise. Then, a hierarchical enhancement fusion module is used to achieve organic fusion of features with the help of a cascade group attention mechanism, reducing information loss during the layer-by-layer transmission of features. Then, a progressive gradient enhancement detection head is used to optimize the training phase of the model by introducing programmable gradient information. By adaptively programming and adjusting the programmable gradient information, the model can more flexibly select key areas in the image during training. In addition, richer gradient information is introduced in the back propagation of information to help the model improve its sensitivity to small feature changes and further optimize feature learning. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 This is the overall network block diagram of a small target detection network based on sparse feature enhancement fusion in the present invention.
[0042] Figure 2 This is a structural diagram of a sparse feature enhancement module of a small target detection network based on sparse feature enhancement fusion of the present invention.
[0043] Figure 3 This is a structural diagram of a global feature focusing enhancement module of a small target detection network based on sparse feature enhancement fusion in the present invention.
[0044] Figure 4 This is a structural diagram of a global feature enhancement module of a small target detection network based on sparse feature enhancement fusion in the present invention.
[0045] Figure 5 This is a structural diagram of a hierarchical enhancement fusion module of a small target detection network based on sparse feature enhancement fusion in the present invention.
[0046] Figure 6 This is a structural diagram of a progressive gradient enhancement detection head of a small target detection network based on sparse feature enhancement fusion in the present invention.
[0047] Figure 7 This is an example graph of AITOD data.
[0048] Figure 8 This is a visualization of the SEFE-YOLO experimental results. DETAILED DESCRIPTION
[0049] The present invention will be further described and explained below in conjunction with specific embodiments and accompanying drawings:
[0050] Please refer to Figure 1 ,A small object detection network based on sparse feature enhancement fusion (SFEF-YOLO), including: a backbone network (CSPDarkNet), which extracts multi-level features of different scales;
[0051] The Sparse Feature Enhancement Module (SFEM) performs deconvolution operations on shallow features to restore image details to the greatest extent possible. By integrating features from global to local, it further improves the detection accuracy of small targets and enhances the ability to distinguish small targets in complex backgrounds.
[0052] Please refer to Figure 2 The sparse feature enhancement module includes sparse convolution, global feature focus enhancement module and hierarchical enhancement fusion module. First, the multi-level features of the backbone network are processed by sparse convolution and then fused. They are passed as input to the global feature focus enhancement module to further enrich the low-level information in the shallow features. Then, the hierarchical enhancement fusion module is used to fuse the features of different levels layer by layer. Finally, the enhanced fusion features are input to the progressive gradient enhancement detection head for final inference.
[0053] The Progressive Gradient Enhancement Detection Head (PGEDH) effectively optimizes the performance of the network and improves the detection effect by dynamically adjusting the way information is propagated during training. It significantly improves the detection accuracy without adding additional parameters, ensuring that the network can accurately detect and classify targets.
[0054] Figure 2 P2, P3, P4 and P5 refer to the backbone network through the input image I inMulti-level features are obtained after downsampling by 4x, 8x, 16x, and 32x, respectively. Traditional small object detection networks perform poorly on the P3, P4, and P5 layers. A common solution is to directly add features extracted from the P2 layer, improving small object detection capabilities by transferring shallow features. However, the large image size of the P2 layer significantly affects the number of model parameters and computational complexity, significantly impacting the real-time performance of detection. To address this issue, a sparse feature enhancement module is proposed based on the extraction and fusion of small object features. First, the features of the P2 layer are processed through sparse convolution and then fused with the features of the P3 layer. The sparse convolution is then passed as input to the global feature focus enhancement module (GDFEM), further enriching the low-level information in the shallow features. Next, a hierarchical enhancement fusion module is used to progressively fuse features from different layers. Finally, the enhanced fused features are input to the progressive gradient enhancement detection head for final inference.
[0055] Traditional convolution operations are performed by performing calculations on every input channel and every position of each convolution kernel. This comprehensive approach ensures the integrity and accuracy of feature extraction, but it also comes with a significant amount of computation, especially when the input data size is large, where computational complexity and time overhead increase significantly. The general convolution operation formula is shown below.
[0056]
[0057] Here, X is the input feature map, of size H × W, where H and W are the height and width of the input image, respectively. In this chapter, the output image is 800 × 800. W is the convolution kernel, of size k × k. Y is the output feature map. The subscripts i and j are the coordinate indices of the output feature map, and m and n are the indices of the convolution kernel, indicating the sliding window of the convolution kernel. This comprehensive calculation ensures comprehensive feature extraction, but it also suffers from the problem of excessive computational effort. To address this issue, depthwise separable convolution was developed.
[0058] Sparse convolution is an improved sparse convolution based on depthwise separable convolution. Depthwise separable convolution splits the standard convolution into two independent operations: depthwise convolution and pointwise convolution. Depthwise separable convolution first performs convolution operations on each input channel independently using convolution kernels of different sizes, and achieves fusion between channels through pointwise convolution. For each input channel C, the mathematical expression of depthwise convolution is as follows:
[0059]
[0060] Among them, X i+m,j+n,c is the pixel value of the input image in channel c, W m,n,cis the weight of the corresponding convolution kernel, and the weight of each input channel is independent;
[0061] Point-by-point convolution performs information fusion on the output features of depth convolution through a 1×1 convolution operation, integrating and merging the features of multiple channels. The formula is as follows:
[0062]
[0063] Among them, X i,j,c is the value of the output feature map after depth convolution in channel c, W 1,1,c,k It is the weight of the 1×1 convolution kernel, which is used to fuse information from different channels. The final output feature map has a value of Y in the kth channel. i,j,k By decomposing the standard convolution into these two steps, the depth-wise separable convolution significantly reduces the amount of computation, especially when the number of channels is large, which can significantly improve the computational efficiency.
[0064] The sparsification mechanism is further introduced to expand and optimize the convolution operation, further reducing the amount of calculation in the convolution operation. Through the mask mechanism, some non-critical weights are discarded, while ensuring the effective extraction of key information and eliminating redundant calculations. Sparse convolution can be represented by a binary mask matrix. Assuming that the sparse mask matrix is S, the calculation formula of sparse convolution is as follows:
[0065]
[0066] Among them, S m,n,c is a binary sparse mask that determines which weights are set to zero. If S m,n,c The value is 0, indicating that the convolution at the corresponding position will not participate in the calculation.
[0067] Similarly, in the process of point-by-point convolution, a sparsity mechanism is introduced to optimize the amount of calculation. When performing point-by-point convolution, the channel sparse mask S is used to optimize the amount of calculation. c To filter the channels involved in information fusion, the calculation formula of sparse point-by-point convolution is as follows:
[0068]
[0069] Among them, S c It is a sparse mask that determines which channels participate in the point-by-point convolution calculation and takes a value of 0 or 1. 1,1,c,k is the weight of point-by-point convolution, which can further reduce redundant calculations and improve the efficiency of the model.
[0070] For the small target features extracted by the P2 layer, the process of extracting information enhancement features P'2 through sparse convolution can be expressed as follows:
[0071] P'2=fpointwise (f depthwise (P2)
[0072] Among them, f pointwise (·) represents the mapping function of sparse point-wise convolution, f depthwise (·) represents the mapping function of sparse depthwise convolution.
[0073] The appearance features of small targets usually exist in the shallow layer and are easily lost during the downsampling process. In order to further enhance the extracted features, the feature P after the fusion of P2 and P3 is merge2_3 Enhance to obtain features richer in small target appearance information. merge2_3 The input is sent to the global feature focus enhancement module, which uses multi-branch enhancement to enhance the important information area and suppress background noise. This makes the details and appearance information of small targets clearer and enhances the robustness of features under complex backgrounds. The structure of the global feature focus enhancement module is as follows: Figure 3 shown.
[0074] The global feature enhancement module is the core structure of the global feature focus enhancement module, which can effectively fuse and enhance features at multiple levels and different scales. It focuses on and optimizes the fine features of small targets, enabling the network to better handle small target detection tasks in complex backgrounds. The global feature enhancement module provides a flexible and efficient mechanism that can improve the ability of feature expression through the interactive fusion of features at different levels, thereby significantly improving the detection performance of small targets. The structure of the global feature enhancement module is as follows: Figure 4 shown.
[0075] in, Figure 4 The Local, Large, and Global fields of the OKM module correspond to the local branch, the large core branch, and the global branch, respectively.
[0076] The global feature enhancement module merge2_3 After convolution processing, P merge2_3 Passed to three branches for processing respectively.
[0077] The local branch focuses on supplementing the local details of the image, suppressing background noise and highlighting important parts to emphasize the appearance information of small targets. The process can be expressed as follows.
[0078]
[0079] The large kernel branch uses large convolution kernels to increase the receptive field and capture more contextual information. To reduce the number of parameters and avoid weakening the focus on small objects due to overly large convolution kernels, the large convolution kernel size is adjusted to 31×31. A strip self-attention mechanism is also introduced, using convolution kernels of 1×31 and 31×1 sizes to capture strip-shaped contextual information.
[0080]
[0081] The feature maps after three different convolution operations are fused to obtain feature maps enhanced based on contextual information in various directions.
[0082]
[0083] Among them, α, β, and γ are learnable weights that control the degree of fusion of convolution results in different directions. The final fused feature map It will contain contextual information from all directions, providing richer features for subsequent feature processing and detection tasks.
[0084] The global branch complements the model's overall scene understanding. Because the input image is large, a 31×31 receptive field cannot fully cover the image. Therefore, the Dynamic Convolutional Attention Module (DCAM) and the Feature Scale Attention Module (FSAM) introduce attention mechanisms in the spatial and channel dimensions. Furthermore, frequency domain information is used to help the model better understand the spatial information in the image, thereby improving its inference performance.
[0085]
[0086] Among them, GAP(·) represents global average pooling, IF(·) is the mapping function of the inverse fast Fourier transform, and F(·) is the mapping function of the Fourier transform.
[0087] refer to Figure 5 , the hierarchical enhancement fusion module aims to effectively fuse multi-level features and improves the feature fusion module by using an inverted residual structure combined with a cascade mechanism:
[0088] The depth-wise separable convolution operation is used to process features, greatly reducing the amount of computation and retaining the original information through residual links. To further improve the efficiency of feature fusion, the cascaded group attention mechanism is used to decompose the attention calculation into multiple levels of local calculations and gradually enhance the feature information through the cascade structure. The working process is as follows:
[0089]
[0090] Among them, Q and K are the query and key matrices after grouping, respectively, and K T The key matrix is transposed. The attention results of each group are transmitted and fused layer by layer in a cascade manner. The output of each layer not only depends on the calculation results of the previous layer, but also further integrates the features of the previous layer. The process is as follows:
[0091]
[0092] Where G is the number of groups, V i Represents the value matrix of each group, representing the feature information of the feature map,
[0093] Through the above optimization, while reducing the amount of computation and memory consumption, the recognizability of small targets is enhanced, the information transmission strategy is optimized, and the performance of the model is effectively improved.
[0094] The progressive gradient enhancement detection head enhances the gradient flow during training by adding two key modules: multi-level auxiliary information and auxiliary reversible branches, reducing information loss in back propagation and improving the performance of the detection network. The structure of the progressive gradient enhancement detection head is as follows: Figure 6 shown.
[0095] Multi-level Auxiliary Information (MAI) is used to integrate features of different scales at different levels, so that low-level information of the image can be retained in feature maps at higher levels, achieving the complementarity of appearance information and semantic information:
[0096]
[0097] Among them, F MAI is the fused multi-level feature map, W i is the weight of features at different levels, F i Represents feature maps at different levels.
[0098] To maintain gradient flow during training and make it more robust, PGEDH also builds an auxiliary reversible branch to ensure the effectiveness and stability of information transmission within the network. This allows the network to better learn precise features and the training process to converge faster. The auxiliary reversible branch receives the network's intermediate output during training and is independent of the main branch. It maintains gradient flow through skip connections, reverse calculations, and other actions.
[0099] When the main branch processes the feature map through the network module, the auxiliary branch also processes the input data in parallel to obtain the output feature map Z of the auxiliary branch.arb , in the back-propagation phase, the auxiliary branch improves the gradient of the main network through reverse operation:
[0100]
[0101] Among them, L is the loss function, G arb It is the gradient of the auxiliary branch. After combining the gradient of the main branch with the gradient of the auxiliary branch, the main network model parameters are updated;
[0102]
[0103] Among them, η is the learning rate, G update is the updated gradient after the main branch and auxiliary branch are combined, θ old is the parameter of the network model before updating, θ new are the updated parameters of the network model in this iteration:
[0104] The auxiliary reversible branch significantly improves the stability of gradient transfer by injecting additional robust gradient information into the main branch, making the update efficiency of each model iteration higher, thereby improving the overall training speed and performance of the network.
[0105] To validate the performance of the proposed SFEF-YOLO, a series of experiments are conducted. First, the dataset, evaluation metrics, and experimental setup are introduced. SFEF-YOLO is then compared with the baseline network YOLOv8 and current mainstream object detection algorithms. Finally, ablation experiments demonstrate the effectiveness of the improved global feature focus enhancement module and the layered enhancement fusion module.
[0106] Dataset
[0107] In order to more fully verify the performance of the SFEF-YOLO network for small target detection tasks, we chose to use the self-made dataset VLSTD and the open source remote sensing image small target dataset AITOD for performance evaluation experiments to analyze the performance of the algorithm in different application scenarios.
[0108] The AITOD dataset is a large-scale dataset for detecting tiny objects in the air proposed by Wang et al. from Wuhan University in 2021. The dataset contains 28,036 aerial images, covering 700,621 object examples in eight categories including airplanes, bridges, oil tanks, ships, swimming pools, vehicles, people, and windmills. Compared with existing target detection datasets, the average size of targets in the AITOD dataset is 12.8 pixels, which is much smaller than other datasets. Figure 7 Shows some data from the AITOD dataset.
[0109] Evaluation indicators
[0110] To more intuitively demonstrate the effectiveness of the proposed algorithm, we use the MS COCO evaluation matrix as an evaluation metric, combined with the small-object metrics proposed in AITOD, to uniformly evaluate the detection accuracy, recall, detection capability, and inference speed of the proposed SFEF-YOLO algorithm and the algorithmic composite model used for comparison. The relevant metrics used are described in detail below.
[0111] In target detection, IoU is used to determine the degree of match between the predicted box and the true box. By presetting a threshold μ, a qualitative judgment is made on the detection result based on the IoU value between the predicted box and the true labeled box. The calculation formula of IoU is shown below.
[0112]
[0113] Among them, A and B represent the predicted box and the true annotation box respectively, ∩ represents the intersection, that is, the size of the overlapping area of the predicted box and the true annotation box, and ∪ represents the union, that is, the size of the joint area of the predicted box and the true annotation box.
[0114] When the IoU value between the predicted box and the ground-truth box is greater than or equal to the threshold μ, it is considered a correct detection result, called a positive example, and denoted as TP (True Positive). If the IoU value is less than the threshold μ and the predicted box exists, it is a negative example, denoted as FP (False Positive). If the ground-truth box is not detected, it is considered a missed detection, denoted as FN (False Negative). Based on the above three detection results, the model's detection precision (Precision, P) and recall (Recall, R) are calculated respectively. The formulas are shown below.
[0115]
[0116] In practical applications, recall represents the model's recall capability, while precision represents its precision capability. These two metrics can comprehensively evaluate a model's performance. However, given the common conflict between these two metrics, Average Precision (AP) is introduced to comprehensively evaluate model performance. AP represents the average value of P under different R values. The specific calculation process for AP is shown below.
[0117] Among them, p(r i ) represents the precision P with respect to the recall rate r ifunction, N is the number of sampling points of recall rate, which is usually 101. In the evaluation criteria of MS COCO, AP is the core indicator, which is further subdivided into indicators under different IoU thresholds and target scales to comprehensively measure the performance of the model in tasks of different scales. According to the size of the target, AP is further subdivided into: AP 50 、AP 75 、AP s 、AP m etc. AP s The AP is used to evaluate the average precision of the model in detecting targets with a size smaller than 32×32 pixels. 50 It represents the overall accuracy when the IoU threshold is 0.5. In view of the fact that the targets in the experimental data are mainly small and medium-sized targets, AP is used. s and AP 50 As a key indicator to measure the small target detection performance of the model.
[0118] The proposed SFEF-YOLO is compared with the baseline network YOLOv8. The experimental results are shown in Table 1-1 and Table 1-2.
[0119] Table 1-1 Comparative experimental results of SFEF-YOLO and Baseline on the AITOD dataset.
[0120]
[0121] As shown in Table 1-1, the performance of the SFEF-YOLO network has been significantly improved in the AITOD dataset. 50 and AP s Compared with the baseline network YOLOv8, the performance is improved by 8.2% and 5.9% respectively. 75 The AP also increased by 4.1% and 4.2% respectively, indicating that the network is more accurate in locating the target. This is mainly due to the introduction of the sparse feature enhancement module, which enhances the shallow features of the target, making the details and appearance information of the small target more clear and prominent, and minimizing the information loss during the inter-layer information transmission process, making the features of the small target more significant and improving the efficiency of information utilization. At the same time, the AP for medium-sized target detection performance is improved. m There is also a 6.8% improvement. This is because the global feature focus enhancement module can effectively suppress noise and strengthen areas of important information, enhance and fuse features of multiple levels and different scales, and improve the performance of multi-scale target detection tasks.
[0122] Table 1-2 Comparative experimental results of SFEF-YOLO and Baseline on the VLSTD dataset.
[0123]
[0124] As shown in Table 1-2, the performance of the SFEF-YOLO network has been significantly improved in the AITOD dataset. 50 and AP s Compared with the baseline network YOLOv8, the performance is improved by 7.4% and 7.9% respectively. From the experimental results, SFEF-YOLO has excellent generalization ability and outstanding detection performance in different application scenarios.
[0125] In order to more intuitively demonstrate the effectiveness of the SFEF-YOLO method proposed in this chapter, some experimental results are visualized. The visualization of the results is as follows: Figure 8 shown.
[0126] Figure 8 The detection visualization results of four remote sensing images are selected. There are four columns from left to right in the figure, representing the original image, the true label, the prediction result of YOLOv8 and the prediction result of SFEF-YOLO proposed in this chapter. It is not difficult to see from the comparison of the pictures that Figure 8 In (a), in addition to the target ship, there are also many reefs in the ocean background. Although YOLOv8 successfully detects the ship target, it is interfered by the reefs and misidentifies the reefs as ships. SFEF-YOLO successfully detects the ship and distinguishes between the ship and the reefs. Figure 8 In (b), due to the extremely low distinction between the background and the target and the relatively small size of the target itself, YOLOv8 has difficulty in extracting small targets and therefore fails to detect any small targets. SFEF-YOLO, despite some errors, still successfully detects a small target. Figure 8 In (c), SFEF-YOLO not only correctly detects the target, but also has a higher confidence level in target detection, which shows that SFEF-YOLO can better identify the target and confirm the target more strongly than YOLOv8. Figure 8 In (d), SFEF-YOLO correctly detects the target when YOLOv8 misses it.
[0127] In summary, SFEF-YOLO has superior capabilities for extracting small object features, better understands image content, clearly distinguishes between foreground and target, and is highly capable of detecting small objects in complex backgrounds. Through a series of optimization methods, SFEF-YOLO significantly improves various metrics in small object detection in remote sensing images compared to the baseline network YOLOv8.
[0128] SFEF-YOLO was compared with other mainstream object detection methods. The experimental results are shown in Tables 1-3 and 1-4. The proposed time of each algorithm is reflected in the table. The main comparisons were with other YOLO series algorithms, including YOLOv6, DAMO-YOLO, and YOLOX, as well as other common object detection algorithms: FCOS, DetectoRS, Cascade R-CNN, and RTMDet. The experimental results show that compared with other YOLO series algorithms and common object detection algorithms, YOLOv10, proposed in 2024, is the latest object detection model. The experimental results of SFEF-YOLO on the AITOD dataset and the VLSTD dataset are shown in Tables 1-3 and 1-4.
[0129] Table 1-3 Comparative experimental results of SFEF-YOLO and other methods on the AITOD dataset.
[0130]
[0131] From the experimental results in Table 1-3, we can see that on the AITOD dataset, SFEF-YOLO achieved the best performance in all performance evaluation indicators. Among them, AP, one of the key evaluation indicators for small target detection, 50 Compared with the second-best DetectoRS algorithm, it has improved by 1.8%. Another key indicator AP s It also achieves a 6% performance improvement over the latest object detection algorithm, YOLOv10, fully demonstrating the outstanding performance of the SFEF-YOLO algorithm proposed in this chapter in small object detection tasks in remote sensing images. Although SFEF-YOLO primarily concentrates computational resources in shallow networks to enhance the features of small objects, its effective noise suppression, prominent processing of target areas, and the use of large kernel convolutions in the global feature focus enhancement module enable the network to better understand image content. This allows the small object detection network based on sparse feature enhancement fusion proposed in this chapter to continue its outstanding performance in medium-sized object detection tasks, maintaining a 6.2% performance lead over RTMDet, which is suitable for natural object detection.
[0132] Table 1-4 Comparative experimental results of SFEF-YOLO and other methods on the VLSTD dataset.
[0133]
[0134]
[0135] From the experimental results in Table 1-4, we can see that on the VLSTD dataset, SFEF-YOLO has the advantages of AP, AP 50 、AP 75 、AP s and AP m All metrics achieved optimal performance. Furthermore, compared to images in the AITOD dataset, the simulated VLSTD dataset features higher discrimination between target and background, less occlusion of the target by the surrounding environment, and a relatively constant target size. This significantly reduces the difficulty of detection and results in generally higher detection accuracy for the network. Despite this, SFEF-YOLO still achieved performance improvements.
[0136] Implementation results demonstrate that the algorithm proposed in this chapter fully utilizes extracted features through effective enhancement of low-level target features, rational allocation of computational resources, and optimized information transmission strategies. Furthermore, through a programmable gradient information mechanism, it helps the model detect targets faster and more accurately, comprehensively improving the network's detection performance and demonstrating strong competitiveness compared to numerous object detection algorithms.
[0137] To verify whether the modules proposed in this chapter optimize the network, we conducted ablation experiments to explore the impact of the global feature focus enhancement module and the progressive gradient boosting detection head on network performance. The methods in the experimental results are highlighted in bold, and the best and second-best results for each metric are highlighted in bold and underlined, respectively.
[0138] Verification experiment of SFEM
[0139] To comprehensively evaluate the global feature focus enhancement module's effectiveness in enhancing small object features, we conducted validation experiments on the sparse convolution layer (SPDConv) in SFEM, the global feature focus enhancement module (GDFEM), and the hierarchical enhancement fusion module (HEFM). The effectiveness of sparse convolution was verified by replacing it with a standard convolution. The effectiveness of GDFEM was also verified by replacing it with the native PAN-FPN (Yolov8), retaining the sparse convolutions used in the pyramid for extracting low-level features, and continuing to use the hierarchical enhancement module for feature fusion. Finally, the effectiveness of HEFM was verified by replacing the feature fusion module in SFEM with the C2F module. The experimental results are shown in Tables 1-5.
[0140] Table 1 - Ablation experiment results of 5GDFEM.
[0141]
[0142] Ablation experiments demonstrate that the global feature focus enhancement module plays a crucial role in SFEM. This is primarily due to the difficulty in extracting highly expressive features when detecting small objects. GDFEM effectively enhances features and transmits these optimized, information-rich features to the deep network, preserving the object's appearance and thus improving network performance. The sparse convolutions in GDFEM, specifically designed for extracting low-level features, fully exploit the spatial sparsity of objects. While improving computational efficiency, they also effectively suppress noise, enhancing object recognition and further enhancing model performance.
[0143] The hierarchical enhancement fusion module introduces a cascaded group attention mechanism, dividing the overall computation into multiple subgroups and fusing the results of each subgroup separately. This module primarily improves feature fusion efficiency and reduces information loss. However, due to the limited capabilities of the baseline network for extracting small object features, directly introducing an attention mechanism can lead to overfitting, thereby reducing model performance.
[0144] In summary, SFEM helps the network better understand image content and obtain higher-quality feature representation by effectively enhancing target features, suppressing noise, and optimizing information transmission strategies, thereby significantly improving the detection efficiency of the model.
[0145] The present invention provides a small target detection network based on sparse feature enhancement and fusion. In response to the difficulty of feature extraction in YOLOv8 and the problem of information loss during the transmission process, a sparse feature enhancement module is designed. By utilizing spatial sparsity, the low-level features of the target are effectively extracted. The low-level features are enhanced by global feature focusing to highlight important areas and suppress background noise. Then, a hierarchical enhancement fusion module is used to achieve organic fusion of features with the help of a cascade group attention mechanism, reducing the loss of information during the layer-by-layer transmission of features. Then, a progressive gradient enhancement detection head is used to optimize the training phase of the model by introducing programmable gradient information. By adaptively programming and adjusting the programmable gradient information, the model can more flexibly select key areas in the image during training. In addition, richer gradient information is introduced in the back propagation of information to help the model improve its sensitivity to small feature changes and further optimize feature learning.
[0146] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the scope of protection of the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the essence and scope of the technical solutions of the present invention.
Claims
1. A small target detection network based on sparse feature enhancement fusion, characterized by: include Backbone network, extracts multi-level features of different scales through the backbone network; The sparse feature enhancement module performs deconvolution operations on shallow features to maximize image detail recovery. By integrating features from global to local, it further improves the detection accuracy of small targets and enhances the ability to distinguish small targets in complex backgrounds. The progressive gradient enhancement detection head effectively optimizes the performance of the network and improves the detection effect by dynamically adjusting the way information is propagated during training. It significantly improves the detection accuracy without adding additional parameters, ensuring that the network can accurately detect targets and perform accurate classification.
2. A small target detection network based on sparse feature enhancement fusion according to claim 1, characterized in that: The sparse feature enhancement module includes sparse convolution, global feature focus enhancement module and hierarchical enhancement fusion module. First, the multi-level features of the backbone network are processed by sparse convolution and then fused and passed as input to the global feature focus enhancement module to further enrich the low-level information in the shallow features. Then, the hierarchical enhancement fusion module is used to progressively fuse the features of different levels layer by layer. Finally, the enhanced fused features are input into the progressive gradient enhancement detection head for final inference.
3. A small target detection network based on sparse feature enhancement fusion according to claim 2, characterized in that: The sparse convolution is an improved sparse convolution based on depthwise separable convolution. Depthwise separable convolution splits the standard convolution into two independent operations: depthwise convolution and pointwise convolution. Depthwise separable convolution first performs convolution operations on each input channel independently using convolution kernels of different sizes, and achieves fusion between channels through pointwise convolution. For each input channel C, the mathematical expression of depthwise convolution is as follows: Among them, X i+m,j+n,c is the pixel value of the input image in channel c, W m,n,c is the weight of the corresponding convolution kernel, and the weight of each input channel is independent; Point-by-point convolution performs information fusion on the output features of depth convolution through a 1×1 convolution operation, integrating and merging the features of multiple channels. The formula is as follows: Among them, X i,j,c is the value of the output feature map after depth convolution in channel c, W 1,1,c,k It is the weight of the 1×1 convolution kernel, which is used to fuse information from different channels. The final output feature map has a value of Y in the kth channel. i,j,k By decomposing the standard convolution into these two steps, the depth-wise separable convolution significantly reduces the amount of computation, especially when the number of channels is large, which can significantly improve the computational efficiency.
4. A small target detection network based on sparse feature enhancement fusion according to claim 3, characterized in that: The sparsification mechanism is further introduced to expand and optimize the convolution operation, further reducing the amount of calculation in the convolution operation. Through the mask mechanism, some non-critical weights are discarded, while ensuring the effective extraction of key information and eliminating redundant calculations. Sparse convolution can be represented by a binary mask matrix. Assuming that the sparse mask matrix is S, the calculation formula of sparse convolution is as follows: Among them, S m,n,c is a binary sparse mask that determines which weights are set to zero. If S m,n,c The value is 0, indicating that the convolution at the corresponding position will not participate in the calculation. Similarly, in the process of point-by-point convolution, a sparsity mechanism is introduced to optimize the amount of calculation. When performing point-by-point convolution, the channel sparse mask S is used to optimize the amount of calculation. c To filter the channels involved in information fusion, sparse Among them, S c It is a sparse mask that determines which channels participate in the point-by-point convolution calculation and takes a value of 0 or 1. 1,1,c,k is the weight of the point-by-point convolution.
5. The small target detection network based on sparse feature enhancement fusion according to claim 2, characterized in that: The global feature enhancement module focuses on and optimizes the fine features of small targets, enabling the network to better handle small target detection tasks in complex backgrounds. The global feature enhancement module provides a flexible and efficient mechanism that can improve the ability of feature expression through the interactive fusion of features at different levels, thereby significantly improving the detection performance of small targets.
6. A small target detection network based on sparse feature enhancement fusion according to claim 2, characterized in that: The hierarchical enhancement fusion module aims to effectively fuse multi-level features. It improves the feature fusion module by using an inverted residual structure combined with a cascade mechanism: The features are processed using depthwise separable convolution operations to greatly reduce the amount of computation and retain the original information through residual links. To further improve the efficiency of feature fusion, a cascaded group attention mechanism is used to decompose the attention calculation into multiple levels of local calculations and gradually enhance the feature information through the cascade structure. The working process is as follows: Among them, Q and K are the query and key matrices after grouping, respectively, and K T The key matrix is transposed. The attention results of each group are transmitted and fused layer by layer in a cascade manner. The output of each layer not only depends on the calculation results of the previous layer, but also further integrates the features of the previous layer. The process is as follows: Where G is the number of groups, V i Represents the value matrix of each group, representing the feature information of the feature map, Through the above optimization, while reducing the amount of computation and memory consumption, the recognizability of small targets is enhanced, the information transmission strategy is optimized, and the performance of the model is effectively improved.
7. The small target detection network based on sparse feature enhancement fusion according to claim 1, characterized in that: The proposed progressive gradient enhancement detection head enhances the gradient flow during training by adding two key modules: multi-level auxiliary information and auxiliary reversible branches, thereby reducing information loss in backpropagation and improving the performance of the detection network.
8. The small target detection network based on sparse feature enhancement fusion according to claim 7, characterized in that: The multi-level auxiliary information is used to integrate features of different scales at different levels, so that low-level information of the image can be retained in the feature maps of higher levels, thus achieving the complementarity of appearance information and semantic information: Among them, F MAI is the fused multi-level feature map, W i is the weight of features at different levels, F i Represents feature maps at different levels.
9. The small target detection network based on sparse feature enhancement fusion according to claim 7, characterized in that: The auxiliary reversible branch ensures the effectiveness and stability of information transmission in the network, allowing the network to better learn accurate features and the training process to converge faster. It maintains the flow of gradients through skip connections, reverse calculations, and other behaviors. When the main branch processes the feature map through the network module, the auxiliary branch also processes the input data in parallel to obtain the output feature map Z of the auxiliary branch. arb , in the back-propagation phase, the auxiliary branch improves the gradient of the main network through reverse operation: Among them, L is the loss function, G arb It is the gradient of the auxiliary branch. After combining the gradient of the main branch with the gradient of the auxiliary branch, the main network model parameters are updated; Among them, η is the learning rate, G update is the updated gradient after the main branch and auxiliary branch are combined, θ old is the parameter of the network model before updating, θ new are the updated parameters of the network model in this iteration: The auxiliary reversible branch significantly improves the stability of gradient transfer by injecting additional robust gradient information into the main branch, making the update efficiency of each model iteration higher, thereby improving the overall training speed and performance of the network.
Citation Information
Cited By
Bridge demolition real-time safety monitoring method based on multi-source information fusion
CN121524969A
A bridge demolition real-time safety monitoring method based on multi-source information fusion
CN121524969B
Remote sensing small target detection method based on sparse feature enhancement and related equipment
CN121788808A