Small target feature processing method and system based on multi-scale convolution
By combining multi-scale convolution and attention mechanism, the accuracy and robustness problems of small target detection in high-voltage transmission line inspection are solved, and efficient and compact detection effects are achieved.
Patent Information
- Application Number
- CN202511140656.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-15
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-08-15
AI Technical Summary
Existing detection models have a weak response to small-scale, low signal-to-noise ratio targets during high-voltage transmission line inspections, resulting in a high missed detection rate and inaccurate positioning. This makes it difficult to meet the needs of high-reliability detection, and the model's high complexity makes it unsuitable for deployment in resource-constrained scenarios.
Multi-scale convolution and small target feature processing methods are adopted, combined with multi-scale shared convolution, spatial to depth convolution, cross-stage polymorphic convolution and dual attention level detection head, embedded SimAM attention processing and composite loss function, and a lightweight feature fusion and detection module is designed.
It achieves accurate perception of tiny and dense defect targets in complex scenes, improves detection accuracy and model compactness, reduces the number of model parameters, and improves detection accuracy and robustness.
Smart Images

Figure CN120726440A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image analysis, and in particular relates to a small target feature processing method and system based on multi-scale convolution. Background Art
[0002] As the high-voltage transmission network continues to expand, safety inspections of transmission lines are crucial to ensuring stable grid operation. Numerous transmission lines traverse complex terrains, including mountains, plateaus, and deserts, subjecting them to prolonged exposure to harsh environments such as strong winds, lightning strikes, ice, and snow. This can lead to structural damage to key components such as insulators, anti-vibration hammers, and spacers, including cracks, breakage, and detachment. Furthermore, foreign objects such as bird nests and plastic film may become attached to the lines. If not detected promptly, these can cause serious faults such as flashovers and short circuits, threatening grid security.
[0003] In recent years, drone-based and deep learning-based target detection technology has been widely used in power inspections, significantly improving inspection efficiency and coverage. However, in actual aerial imagery, most defects and foreign objects are small, low-contrast, and have blurred textures. They are often affected by lighting variations, background interference, and occlusion, making them typical weakly visible targets. Existing detection models lack the ability to extract features and perform multi-scale fusion. They respond poorly to small-scale, low-signal-to-noise ratio targets, resulting in high missed detection rates and inaccurate positioning, making them difficult to meet the requirements of high-reliability inspections.
[0004] To further improve detection performance, some methods enhance feature representation by introducing deep network structures or complex attention mechanisms. However, this often results in increased model parameters and structural redundancy, hindering deployment and application in resource-constrained scenarios. Therefore, how to effectively enhance the feature perception of small, low-contrast targets while controlling model complexity, thereby improving detection accuracy and robustness, remains a pressing technical challenge. Summary of the Invention
[0005] In order to solve the problems raised in the background technology, the present invention provides a small target feature processing method and system based on multi-scale convolution.
[0006] The technical solutions of the present invention are as follows: The present invention provides a small target feature processing method based on multi-scale convolution, comprising: S1: Acquire a transmission line image and perform multiple downsampling to obtain the corresponding first, second, and third extracted feature maps; The third extracted feature map is processed by multi-scale shared convolution to obtain the first feature map; S2: After the first feature map and the third extracted feature map are spliced, a multi-scale feature fusion process is performed to obtain a first fused feature map; The first extracted feature map is subjected to space-to-depth convolution to obtain the second fused feature map, specifically: The first extracted feature map is divided into four non-overlapping sub-blocks in the spatial dimension according to different row and column indices. The four non-overlapping sub-blocks are cascaded along the channel dimension to generate an enhanced feature map. After that, local context features are extracted through grouped convolution and weighted fusion is performed in combination with a lightweight spatial attention mechanism to obtain the second fused feature map. The second fused feature map, the first fused feature map, and the second extracted feature map are spliced together, and after cross-stage polymorphic convolution processing, multi-scale feature fusion processing is performed to obtain a third fused feature map, which is up-sampled and spliced with the first extracted feature map to obtain the second feature map; S3: The second feature map is subjected to multi-scale feature fusion cascade processing, specifically: The second feature map is processed by multi-scale feature fusion and SimAM attention in sequence to obtain the first-scale feature map; After the first-scale feature map and the third-scale fusion feature map are spliced, multi-scale feature fusion processing and SimAM attention processing are performed in sequence to obtain the second-scale feature map; After the second-scale feature map is spliced with the first fusion feature map, multi-scale feature fusion processing and SimAM attention processing are performed in sequence to obtain the third-scale feature map; The first scale feature map, the second scale feature map and the third scale feature map constitute a third feature map group; S4: Each feature map in the third feature map group is processed by the dual attention level detection head to obtain the fourth feature map group, which is then processed by the regression head and the classification head for detection.
[0007] The multi-scale shared convolution processing described in S1 is specifically as follows: The third extracted feature map is subjected to channel compression by convolution to obtain the intermediate feature map, and then multi-scale context features are extracted through three dilation branches with different expansion rates and one ghost convolution branch, while retaining the third extracted feature map as the residual branch; after the features of the dilation branch, ghost convolution branch and residual branch are spliced in the channel dimension, feature fusion and compression are performed through convolution to obtain the first feature map.
[0008] The multi-scale feature fusion processing described in S2 is specifically as follows: After the features to be processed are extracted through reparameterized convolution, they are processed by spatial attention and depth-wise separable convolution in sequence to obtain the main branch enhanced features; The main branch enhanced features and the residual branch features are channel-wise spliced and fused through convolution.
[0009] The cross-stage polymorphic convolution processing described in S2 is specifically as follows: The splicing features are parallelized to perform four different shapes of depth convolution operations to obtain multi-shape convolution features; After the spliced features are processed by frequency domain channel attention, spatial attention processing is performed to obtain spatially weighted frequency domain features; The spatially weighted frequency domain features are fused with the multi-shape convolution features, and the fused features are channel-joined with the residual branch features of the splicing features. After the number of channels is compressed by convolution, channel attention processing is performed.
[0010] Each feature map in the third feature map group described in S4 is processed by a dual attention level detection head, specifically: Each scale feature of the third feature map group is first subjected to channel compression by convolution, and then scale attention and spatial attention are applied in sequence, and feature refinement is completed in shared convolution.
[0011] Furthermore, the method further includes: calculating loss, where the loss is composed of classification loss and regression loss.
[0012] It also includes: a double distillation mechanism, including a block-correlation knowledge distillation strategy loss based on calculating the Kullback–Leibler divergence for each category, and a feature alignment loss using channel weight distance.
[0013] It also includes: building a graphical user interface that supports image input, model loading, inspection result visualization, defect statistical analysis and result export.
[0014] The present invention also provides a small target feature processing system based on multi-scale convolution, comprising: Feature extraction module: used to obtain the transmission line image and obtain the corresponding first, second and third extracted feature maps after multiple downsampling; The third extracted feature map is processed by multi-scale shared convolution to obtain the first feature map; Feature fusion module: After the first feature map and the third extracted feature map are spliced, a multi-scale feature fusion process is performed to obtain a first fused feature map; The first extracted feature map is subjected to space-to-depth convolution to obtain the second fused feature map, specifically: The first extracted feature map is divided into four non-overlapping sub-blocks in the spatial dimension according to different row and column indices. The four non-overlapping sub-blocks are cascaded along the channel dimension to generate an enhanced feature map. After that, local context features are extracted through grouped convolution and weighted fusion is performed in combination with a lightweight spatial attention mechanism to obtain the second fused feature map. The second fused feature map, the first fused feature map, and the second extracted feature map are spliced together, and after cross-stage polymorphic convolution processing, multi-scale feature fusion processing is performed to obtain a third fused feature map, which is up-sampled and spliced with the first extracted feature map to obtain the second feature map; Cascade processing module: The second feature map performs multi-scale feature fusion cascade processing, specifically: The second feature map is processed by multi-scale feature fusion and SimAM attention in sequence to obtain the first-scale feature map; After the first-scale feature map and the third-scale fusion feature map are spliced, multi-scale feature fusion processing and SimAM attention processing are performed in sequence to obtain the second-scale feature map; After the second-scale feature map is spliced with the first fusion feature map, multi-scale feature fusion processing and SimAM attention processing are performed in sequence to obtain the third-scale feature map; The first scale feature map, the second scale feature map and the third scale feature map constitute a third feature map group; Detection module: Each feature map in the third feature map group is processed by the dual attention level detection head to obtain the fourth feature map group, which is then processed by the regression head and classification head for detection.
[0015] Beneficial effects The present invention utilizes multi-scale shared convolution processing, space-to-depth convolution processing, cross-stage polymorphic convolution processing, and combines multi-scale feature fusion processing and dual attention hierarchical detection heads to achieve accurate perception of tiny and dense defect targets in complex scenes, and has comprehensive advantages in detection accuracy and model compactness.
[0016] This paper embeds SimAM attention processing during feature fusion and designs a composite loss function that combines SlideLoss, IoU, and normalized Gaussian Wasserstein distance to optimize model performance from two dimensions: "feature focus" and "loss modeling." SimAM, as a lightweight attention module, significantly improves saliency expression without introducing trainable parameters. Furthermore, the loss function comprehensively considers classification boundary smoothness, target localization accuracy, and shape matching capabilities, effectively accelerating model convergence, improving accuracy and stability, and suppressing overfitting in multi-scale scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 The test results of the model using the method of the present application are compared with those of the model not using the method of the present application. DETAILED DESCRIPTION
[0018] The following examples are intended to illustrate the present invention rather than to further limit the present invention.
[0019] The present invention provides a small target feature processing method based on multi-scale convolution, comprising: S1: Obtain a transmission line image and perform multiple downsampling to obtain the corresponding first, second, and third extracted feature maps (the number of downsampling times is not less than 3); The third extracted feature map is processed by multi-scale shared convolution to obtain the first feature map.
[0020] First, the acquired transmission line images were preprocessed, including image enhancement. The preprocessed images were annotated and divided into training, validation, and test sets. For example, the preprocessed images were annotated using Roboflow for object detection. Finally, the annotated images were divided into a training set (4,141 images), a validation set (591 images), and a test set (1,184 images) in a 7:1:2 ratio.
[0021] Analysis of the labeled data shows that there are a significant number of small objects in the image, and their distribution is uneven. The vast majority of objects have a normalized size less than 0.2, and small objects are primarily concentrated in the upper region of the image, exhibiting complex multi-scale spatial distribution characteristics.
[0022] In the downsampling stage, the first, second and third extracted feature maps can be 4, 8 and 16 times the feature maps, respectively, and are correspondingly recorded as P2 feature map, P3 feature map and P4 feature map.
[0023] Preferably, the multi-scale shared convolution processing (denoted as MDSC) is specifically as follows: The third extracted feature map is subjected to channel compression through convolution (such as 1×1) to obtain the intermediate feature map ; Multi-scale context features are extracted through three dilated convolution branches with different expansion rates and one ghost convolution (Ghost Convolution, Ghost-Conv) branch, while retaining the third extracted feature map as the residual branch; the features of the dilated convolution branch, ghost convolution branch and residual branch are spliced in the channel dimension, and then feature fusion and compression are performed through convolution (such as 1×1) to obtain the first feature map . Satisfies the following expression: ; Where, 、 、 Represent the dilation rates of 1, 3, and 5 for the dilation branch features, respectively; is the 5×5 ghost convolution branch feature; The third extracted feature map for the residual connection.
[0024] S2: After the first feature map and the third extracted feature map are spliced, a multi-scale feature fusion process is performed to obtain a first fused feature map; The first extracted feature map is subjected to space-to-depth convolution to obtain the second fused feature map, specifically: The first extracted feature map is divided into four non-overlapping sub-blocks in the spatial dimension according to different row and column indices. The four non-overlapping sub-blocks are cascaded along the channel dimension to generate an enhanced feature map. After that, local context features are extracted through grouped convolution and weighted fusion is performed in combination with a lightweight spatial attention mechanism to obtain the second fused feature map. The second fused feature map, the first fused feature map and the second extracted feature map are spliced together, and after cross-stage polymorphic convolution processing, multi-scale feature fusion processing is performed to obtain a third fused feature map. After upsampling, it is spliced with the first extracted feature map to obtain the second feature map.
[0025] First, the first feature map After nearest neighbor upsampling, it is spliced with the 16-fold downsampled feature map (P4 feature map, i.e., the third extracted feature map) in the feature extraction stage, and preliminary semantic integration is performed through multi-scale feature fusion processing (denoted as MSFA-Block).
[0026] The multi-scale feature fusion process adopts a lightweight multi-path convolution structure and attention fusion mechanism, including a main branch and a residual branch. Specifically: Features to be processed Features are extracted through reparameterized convolution (denoted as RepConv), and then lightweight spatial attention is applied to enhance the response of significant areas. Then, deep separable convolution (DSConv) is used for spatial feature extraction to obtain the main branch enhanced features. ; can be expressed as: ; Where, represents spatial attention processing, is a depth-wise separable convolution operation, is the activation function.
[0027] Main branch enhancements and residual branch features Perform channel splicing and fusion through 1×1 convolution.
[0028] Next, regarding the space-to-depth convolution processing (denoted as SPDConv), the downsampled 4x feature map (P2 feature map, i.e., the first extracted feature map) in the feature extraction stage is divided into four non-overlapping sub-blocks in a 2×2 grid in the spatial dimension according to different row and column indices. Specifically, all pixels with even row and column indices are extracted to form the first sub-block, all pixels with odd row indices and even column indices are extracted to form the second sub-block, all pixels with even row indices and odd column indices are extracted to form the third sub-block, and all pixels with odd row and column indices are extracted to form the fourth sub-block. Then, these four sub-blocks are cascaded along the channel dimension to generate an enhanced feature map. , the expression is as follows: ; Where, Indicates a cascade operation along the channel dimension, Represents four spatial sub-blocks obtained by dividing according to different row and column indices.
[0029] The enhanced feature map is subjected to grouped 3×3 convolution to extract local context features, and is further combined with a lightweight spatial attention mechanism for weighted fusion. The resulting second fused feature map is denoted as The expression is as follows: ; Where, represents a convolution operation with grouping, is the global average pooling, is the spatial attention weight convolution kernel, is the activation function, Represents element-by-element multiplication. This structure enhances the local receptive field while improving spatial sensitivity, making it suitable for small target feature modeling.
[0030] The second fused feature map is then concatenated with the first fused feature map after nearest neighbor upsampling and the shallow feature map downsampled 8 times (P3 feature map, i.e., the second extracted feature map) to undergo a cross-stage polymorphic convolution (CSPOKM). The CSPOKM employs a typical CSP (Cross Stage Partial) architecture, dividing the features into a main branch and a residual branch. Specifically: The main branch first performs four different depthwise convolution operations on the spliced features in parallel, with kernel sizes of 1×31, 31×1, 31×31, and 1×1, respectively. This extracts multi-shape directionally aware features to enhance the spatial receptive field and directional discrimination capabilities. To further enhance the representation of frequency-domain features, a frequency-domain channel attention mechanism (FCA) is introduced. Specifically, a Fast Fourier Transform (FFT) is performed on the spliced features, and channel attention weights are generated using a 1×1 convolution and a sigmoid activation function to adaptively enhance frequency-domain amplitude information. Subsequently, spatial attention processing (SCA) is used to enhance the saliency of spatial locations.
[0031] The spatially weighted frequency domain features are fused with the multi-shape convolutional features. The fused features are then channel-concatenated with the residual branch of the concatenated features, and the number of channels is compressed using convolution (e.g., 1×1). Finally, a lightweight channel attention (CA) is applied to further enhance important features while preserving the integrity of semantic information. This results in CSPOKM-processed features.
[0032] Finally, the features processed by CSPOKM are processed by multi-scale feature fusion, multi-path convolution and residual fusion, and then upsampled and spliced with the downsampled 4-fold feature map (P2 feature map, i.e. the first extracted feature map) from the feature extraction stage to obtain the second feature map. .
[0033] S3: The second feature map is subjected to multi-scale feature fusion cascade processing, specifically: The second feature map is processed by multi-scale feature fusion and SimAM attention in sequence to obtain the first-scale feature map; After the first-scale feature map and the third-scale fusion feature map are spliced, multi-scale feature fusion processing and SimAM attention processing are performed in sequence to obtain the second-scale feature map; After the second-scale feature map is spliced with the first fusion feature map, multi-scale feature fusion processing and SimAM attention processing are performed in sequence to obtain the third-scale feature map; The first scale feature map, the second scale feature map, and the third scale feature map constitute a third feature map group.
[0034] Among them, the SimAM module is a parameter-free attention mechanism, the core idea of which is to weight each pixel based on its similarity with the entire image to enhance the response of significant areas. is the weighted attention feature map, which is the third feature map group .
[0035] S4: Each feature map in the third feature map group is processed by the dual attention level detection head to obtain the fourth feature map group, which is then processed by the regression head and the classification head for detection.
[0036] Preferably, the third characteristic map group Each feature map in is processed by a dual attention hierarchical detection head (DAHD), specifically: Each scale feature of the third feature map group is first subjected to channel compression by convolution (such as 1×1), and then scale attention and spatial attention are applied in sequence, and feature refinement is completed in shared convolution.
[0037] ; Where, represents global average pooling; represents spatial attention processing; represents element-wise multiplication; Represents the convolution in shared convolution extraction; Represents the feature maps of different scales in the third feature map group; the fusion feature map of each scale processed by DAHD is recorded as , the fourth feature map group finally obtained Used for final detection prediction.
[0038] The DAHD branch generates box offsets and category confidences through the regression head and classification head respectively, and then splices them in the channel dimension to form predictions at each scale. In the inference stage, the three-scale results are combined to obtain the final detection , the specific operations are as follows: ; ; Where, Indicates channel dimension splicing; represents the scaling factor; is the bounding box regression branch, For classification branches; represents the function that transforms the predicted distribution into the actual bounding box coordinates.
[0039] The present invention utilizes multi-scale shared convolution processing, space-to-depth convolution processing, cross-stage polymorphic convolution processing, and combines multi-scale feature fusion processing and dual attention hierarchical detection heads to achieve accurate perception of tiny and dense defect targets in complex scenes, and has comprehensive advantages in detection accuracy and model compactness.
[0040] The present invention also includes adopting a combined loss function in the training stage to improve the target detection accuracy and robustness, where the combined loss consists of a classification loss and a regression loss.
[0041] Specifically, the classification loss adopts SlideLoss, which is a loss function that calculates the prediction confidence. Adaptive modulation weights are applied to smooth difficult samples, which are calculated as follows: ; Where, Confidence in model predictions; is the true label; represents the standard cross entropy loss function; is the IoU adaptive threshold (usually 0.5); It is an indicator function that takes the value of 1 when the condition is met, otherwise it takes the value of 0. This mechanism improves the sensitivity of the model to difficult-to-classify samples by dynamically adjusting the loss weight.
[0042] The regression loss uses a weighted combination of IoU loss and Normalized Gaussian Wasserstein Distance (NWD) loss to improve the accuracy and stability of bounding box prediction. This combination enhances the model's ability to perceive scale changes and objects with blurred boundaries while maintaining positioning accuracy.
[0043] This paper embeds the SimAM parameter-free attention mechanism during the feature fusion stage and designs a composite loss function that combines SlideLoss, IoU, and NWD to optimize model performance from two dimensions: "feature focus" and "loss modeling." On the one hand, SimAM, as a lightweight attention module, significantly improves saliency expression without introducing trainable parameters. On the other hand, the loss function comprehensively considers classification boundary smoothness, target localization accuracy, and shape matching capabilities, effectively accelerating model convergence, improving accuracy and stability, and suppressing overfitting in multi-scale scenarios.
[0044] Furthermore, this paper introduces a knowledge distillation strategy to mitigate the performance degradation of small models during compression. Specifically, this strategy employs a dual distillation mechanism, including a Block Correlation Knowledge Distillation (BCKD) loss that calculates the Kullback–Leibler divergence for each category, and a feature alignment loss that uses the Channel Weighted Distance (CWD) loss.
[0045] The present invention introduces the BCKD distillation strategy during the training process, so that the student model can obtain detection performance close to that of the teacher model while maintaining its reasoning ability. The detection performance is close to that of the teacher model, which significantly alleviates the problem of "insufficient accuracy of small models" in edge deployment scenarios.
[0046] The present invention also includes the construction of a graphical user interface that supports image input, model loading, inspection result visualization, defect statistical analysis, and result export. For example, a graphical visualization interface built on PyQt5 integrates functions such as image input, model loading, inspection result display, defect classification and quantity statistics, and result storage. The operation process is simple and interactive, providing a good foundation for subsequent embedding into existing inspection business systems or access to edge terminal platforms, significantly improving the inspection efficiency and on-site decision support capabilities of power operation and maintenance personnel.
[0047] The present invention also provides a small target feature processing system based on multi-scale convolution, comprising: Feature extraction module: used to obtain the transmission line image and obtain the corresponding first, second and third extracted feature maps after multiple downsampling; The third extracted feature map is processed by multi-scale shared convolution to obtain the first feature map; Feature fusion module: After the first feature map and the third extracted feature map are spliced, a multi-scale feature fusion process is performed to obtain a first fused feature map; The first extracted feature map is subjected to space-to-depth convolution to obtain the second fused feature map, specifically: The first extracted feature map is divided into four non-overlapping sub-blocks in the spatial dimension according to different row and column indices. The four non-overlapping sub-blocks are cascaded along the channel dimension to generate an enhanced feature map. After that, local context features are extracted through grouped convolution and weighted fusion is performed in combination with a lightweight spatial attention mechanism to obtain the second fused feature map. The second fused feature map, the first fused feature map, and the second extracted feature map are spliced together, and after cross-stage polymorphic convolution processing, multi-scale feature fusion processing is performed to obtain a third fused feature map, which is up-sampled and spliced with the first extracted feature map to obtain the second feature map; Cascade processing module: The second feature map performs multi-scale feature fusion cascade processing, specifically: The second feature map is processed by multi-scale feature fusion and SimAM attention in sequence to obtain the first-scale feature map; After the first-scale feature map and the third-scale fusion feature map are spliced, multi-scale feature fusion processing and SimAM attention processing are performed in sequence to obtain the second-scale feature map; After the second-scale feature map is spliced with the first fusion feature map, multi-scale feature fusion processing and SimAM attention processing are performed in sequence to obtain the third-scale feature map; The first scale feature map, the second scale feature map and the third scale feature map constitute a third feature map group; Detection module: Each feature map in the third feature map group is processed by the dual attention level detection head to obtain the fourth feature map group, which is then processed by the regression head and classification head for detection.
[0048] Experimental results During the experimental phase, the teacher model used a YOLOv8s backbone network with fixed weights during training. The student model used an improved version of the network structure (based on the method of this invention), with a smaller parameter size and stronger ability to process small object features. To enhance semantic alignment, feature distillation employed a multi-layer supervision mechanism. Three groups of intermediate layers in the teacher and student networks (e.g., layers 18, 21, and 24 correspond to layers 21, 25, and 29) were selected for channel-wise alignment. The difference in their channel distributions was measured using the Channel Weighted Distance (CWD) loss, guiding the student network to more effectively learn multi-level semantic features.
[0049] Table 1. Comparison of processing results of the present invention and the original model on the self-built data set
[0050] Table 1 shows the performance comparison between the proposed method and YOLOv8n on the self-built transmission line defect dataset. The evaluation indicators include the average precision of small targets (AP S ), mAP@0.5, mAP@0.5:0.95 and model parameters (Params). Among them, mAP is the arithmetic mean of the average precision (AP) of each category, reflecting the overall detection performance of the model; mAP@0.5 represents the mAP calculated when the intersection over union (IoU) threshold is not less than 0.5, which is used to evaluate the detection ability of the model under normal positioning accuracy; mAP@0.5:0.95 represents the average value of mAP calculated at 10 levels of IoU threshold from 0.5 to 0.95 with a step size of 0.05, which comprehensively reflects the positioning robustness of the model; AP S It measures the detection performance of small objects with an area smaller than 32×32 pixels. The IoU is the ratio of the overlapping area of the predicted and ground-truth boxes to their union area, and is used to determine whether the detection is correct.
[0051] As shown in Table 1, the proposed method has good performance on the self-built transmission line dataset. S The mAP@0.5 score increased by 15.4 percentage points, reaching 81.9%, a 9.9 percentage point increase over YOLOv8n, a 5.4 percentage point increase in mAP@0.5:0.95, and a 33.3% reduction in model parameters. Experimental results show that this method significantly improves the detection capability of small defects and low-visibility objects while maintaining a small model size.
[0052] Figure 1The comparison of the detection results of the proposed method and the original YOLOv8 model (which does not use the proposed method) on a self-built transmission line defect detection dataset is shown. The figure includes several typical scenarios, focusing on the improvement of the proposed method in identifying insulators and anti-vibration hammers. Figure 1 In the original YOLOv8 model, the background area was mistakenly identified as an insulator, resulting in a false detection. In contrast, the proposed method can effectively distinguish between the target and the background, successfully avoiding this false detection and improving the reliability of the detection results.
[0053] These results show that the proposed method is more robust when dealing with small targets and complex backgrounds, and can significantly improve detection accuracy and reliability.
[0054] Table 2. Comparison of processing results of the present invention and the original model in VisDrone 2019
[0055] Table 2 shows a performance comparison between the proposed method and the YOLOv8 model on the VisDrone2019 dataset. This dataset, constructed by the Machine Learning and Data Mining Laboratory of University A in collaboration with multiple institutions, is a benchmark dataset for object detection in drone aerial photography scenarios. It covers complex backgrounds such as urban roads, transportation hubs, and dense crowds. Its small size, severe occlusion, and varying viewpoints make it an effective tool for evaluating models' detection capabilities in complex environments.
[0056] Experimental results show that the proposed method has a high detection accuracy of small targets (AP S ), mAP@0.5 increased by 3.0 percentage points, and mAP@0.5:0.95 increased by 1.8 percentage points. Simultaneously, the number of model parameters decreased from 3.0M to 2.0M, a 33.3% reduction. These results demonstrate the effectiveness of this invention in improving detection accuracy and enhancing small object recognition capabilities.
Claims
1. A small target feature processing method based on multi-scale convolution, characterized in that: include: S1: Acquire a transmission line image and perform multiple downsampling to obtain the corresponding first, second, and third extracted feature maps; The third extracted feature map is processed by multi-scale shared convolution to obtain the first feature map; S2: After the first feature map and the third extracted feature map are spliced, a multi-scale feature fusion process is performed to obtain a first fused feature map; The first extracted feature map is subjected to space-to-depth convolution to obtain the second fused feature map, specifically: The first extracted feature map is divided into four non-overlapping sub-blocks in the spatial dimension according to different row and column indices. The four non-overlapping sub-blocks are cascaded along the channel dimension to generate an enhanced feature map. After that, local context features are extracted through grouped convolution and weighted fusion is performed in combination with a lightweight spatial attention mechanism to obtain the second fused feature map. The second fused feature map, the first fused feature map, and the second extracted feature map are spliced together, and after cross-stage polymorphic convolution processing, multi-scale feature fusion processing is performed to obtain a third fused feature map, which is up-sampled and spliced with the first extracted feature map to obtain the second feature map; S3: The second feature map is subjected to multi-scale feature fusion cascade processing, specifically: The second feature map is processed by multi-scale feature fusion and SimAM attention in sequence to obtain the first-scale feature map; After the first-scale feature map and the third-scale fusion feature map are spliced, multi-scale feature fusion processing and SimAM attention processing are performed in sequence to obtain the second-scale feature map; After the second-scale feature map is spliced with the first fusion feature map, multi-scale feature fusion processing and SimAM attention processing are performed in sequence to obtain the third-scale feature map; The first scale feature map, the second scale feature map and the third scale feature map constitute a third feature map group; S4: Each feature map in the third feature map group is processed by the dual attention level detection head to obtain the fourth feature map group, which is then processed by the regression head and the classification head for detection.
2. The small target feature processing method based on multi-scale convolution according to claim 1 is characterized in that: The multi-scale shared convolution processing described in S1 is specifically as follows: The third extracted feature map is subjected to channel compression by convolution to obtain the intermediate feature map, and then multi-scale context features are extracted through three dilation branches with different expansion rates and one ghost convolution branch, while retaining the third extracted feature map as the residual branch; after the features of the dilation branch, ghost convolution branch and residual branch are spliced in the channel dimension, feature fusion and compression are performed through convolution to obtain the first feature map.
3. The small target feature processing method based on multi-scale convolution according to claim 1 is characterized in that: The multi-scale feature fusion processing described in S2 is specifically as follows: After the features to be processed are extracted through reparameterized convolution, they are processed by spatial attention and depth-wise separable convolution in sequence to obtain the main branch enhanced features; The main branch enhanced features and the residual branch features are channel-wise spliced and fused through convolution.
4. The small target feature processing method based on multi-scale convolution according to claim 1 is characterized in that: The cross-stage polymorphic convolution processing described in S2 is specifically as follows: The splicing features are parallelized to perform four different shapes of depth convolution operations to obtain multi-shape convolution features; After the spliced features are processed by frequency domain channel attention, spatial attention processing is performed to obtain spatially weighted frequency domain features; The spatially weighted frequency domain features are fused with the multi-shape convolution features, and the fused features are channel-joined with the residual branch features of the splicing features. After the number of channels is compressed by convolution, channel attention processing is performed.
5. The small target feature processing method based on multi-scale convolution according to claim 1 is characterized in that: Each feature map in the third feature map group described in S4 is processed by a dual attention level detection head, specifically: Each scale feature of the third feature map group is first subjected to channel compression by convolution, and then scale attention and spatial attention are applied in sequence, and feature refinement is completed in shared convolution.
6. The small target feature processing method based on multi-scale convolution according to claim 1 is characterized in that: Also includes: Calculate the loss, which consists of classification loss and regression loss.
7. The small target feature processing method based on multi-scale convolution according to claim 1 is characterized in that: Also includes: A dual distillation mechanism is adopted, including a block-correlation knowledge distillation strategy loss that calculates the Kullback–Leibler divergence for each category, and a feature alignment loss that uses channel weight distance.
8. The small target feature processing method based on multi-scale convolution according to claim 1 is characterized in that: Also includes: Build a graphical user interface that supports image input, model loading, inspection result visualization, defect statistical analysis, and result export.
9. A small target feature processing system based on multi-scale convolution, characterized in that: include: Feature extraction module: used to obtain the transmission line image and obtain the corresponding first, second and third extracted feature maps after multiple downsampling; The third extracted feature map is processed by multi-scale shared convolution to obtain the first feature map; Feature fusion module: After the first feature map and the third extracted feature map are spliced, a multi-scale feature fusion process is performed to obtain a first fused feature map; The first extracted feature map is subjected to space-to-depth convolution to obtain the second fused feature map, specifically: The first extracted feature map is divided into four non-overlapping sub-blocks in the spatial dimension according to different row and column indices. The four non-overlapping sub-blocks are cascaded along the channel dimension to generate an enhanced feature map. After that, local context features are extracted through grouped convolution and weighted fusion is performed in combination with a lightweight spatial attention mechanism to obtain the second fused feature map. The second fused feature map, the first fused feature map, and the second extracted feature map are spliced together, and after cross-stage polymorphic convolution processing, multi-scale feature fusion processing is performed to obtain a third fused feature map, which is up-sampled and spliced with the first extracted feature map to obtain the second feature map; Cascade processing module: The second feature map performs multi-scale feature fusion cascade processing, specifically: The second feature map is processed by multi-scale feature fusion and SimAM attention in sequence to obtain the first-scale feature map; After the first-scale feature map and the third-scale fusion feature map are spliced, multi-scale feature fusion processing and SimAM attention processing are performed in sequence to obtain the second-scale feature map; After the second-scale feature map is spliced with the first fusion feature map, multi-scale feature fusion processing and SimAM attention processing are performed in sequence to obtain the third-scale feature map; The first scale feature map, the second scale feature map and the third scale feature map constitute a third feature map group; Detection module: Each feature map in the third feature map group is processed by the dual attention level detection head to obtain the fourth feature map group, which is then processed by the regression head and classification head for detection.
Citation Information
Patent Citations
Insulator defect detection method based on improved Center Net
CN114359153A
Aerial image target detection method based on multi-scale cavity convolution
CN116824413A
Insulator defect detection method based on multi-scale characteristics and channel perception
CN119295828A
Defect detection method, system and equipment for power transmission line and storage medium
CN119649031A
Multi-scale double-flow fusion real-time semantic segmentation method for road scene
CN119888222A
Cited By
Medical image segmentation method based on frequency domain enhancement and multi-scale cavity convolution
CN122115483A