Power scene small target defect detection method and system based on dynamic task alignment
By adopting a dynamic task-aligned method for small target defect detection in power scenarios, combined with the HRNet backbone network and various modules, the problem of poor small target detection in power inspection is solved, achieving efficient and accurate small target detection and improving the efficiency and accuracy of power equipment inspection.
Patent Information
- Application Number
- CN202511055239.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-11-14
AI Technical Summary
Existing technologies have poor performance in detecting small targets during power line inspections, especially due to the rapid loss of small target features during convolution, spatial misalignment of prediction boxes in classification and regression tasks, and the difficulty of adapting a single convolution kernel to the feature extraction requirements of multi-scale targets.
A method for detecting small target defects in power scenarios based on dynamic task alignment is adopted. By combining the HRNet backbone network with attention-guided task collaboration module, switchable multi-scale convolutional blocks and flexible inverted bottleneck module, multi-resolution parallel branches and cross-stage feature fusion are realized. Task feature alignment is performed by using cross-task attention mechanism and dynamic anchor boxes, and feature extraction and fusion are performed by switchable multi-scale convolutional blocks and flexible inverted bottleneck module.
It improves the efficiency and accuracy of small target inspection in the field of power line inspection, significantly enhances the precision and recall rate of small target detection, and reduces the amount of computation and the false detection rate.
Smart Images

Figure CN120953582A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power scene detection technology, and more specifically to a method and system for detecting small target defects in power scenes based on dynamic task alignment. Background Technology
[0002] With the continuous increase in demand for power transmission and the rapid expansion of the power system, the pressure on its safe and stable operation is growing daily, and the requirements for the stable operation of power equipment are constantly increasing. Equipment maintenance tasks are becoming increasingly arduous, and the demand for modern, automated, and efficient inspection technologies is becoming increasingly urgent. Robotic inspection is one of the important technologies in this field. Compared with traditional manual inspection, robotic inspection solves the shortcomings of high labor intensity, low inspection efficiency, and insufficient digitalization, and has gradually become an emerging method for equipment inspection in the power industry. Especially in recent years, with the continuous development and in-depth application of drone technology, data processing technology, and software technology in the field of inspection, the power grid industry has built an intelligent inspection business model of "helicopter / drone line inspection + lidar scanning + data processing and analysis + data application and visualization." This model can effectively reduce labor intensity, improve inspection efficiency, expand inspection coverage, and realize the digital display of inspection results, which is of great significance for enhancing the safety, stability, and operational efficiency of the power grid. In the process of robotic inspection, target detection is a key technical task. Traditional target detection algorithms perform poorly in identifying and detecting targets of different scales, especially small targets, in inspection tasks. This is because small targets occupy a small pixel area in the overall image (usually less than 5%), resulting in low information content and high signal-to-noise ratio. In power systems, equipment defects (such as cracked insulators in transmission lines, worn hardware in substations, and overheated cable joints) often exist in the form of "small targets." Therefore, research on small target detection algorithms can solve common problems in the field of power inspection and improve the efficiency and accuracy of robotic inspections.
[0003] The power system encompasses diverse components such as substations, transmission lines, and distribution facilities. Small target defect detection technology, by combining equipment criticality with the probability of defect occurrence, provides core support for the intelligent allocation of inspection resources. Existing technologies face three major challenges: rapid loss of small target features during convolution, spatial misalignment of prediction boxes in classification and regression tasks, and the difficulty of adapting a single convolution kernel to the feature extraction needs of multi-scale targets.
[0004] Therefore, in view of the shortcomings of the existing technology, how to provide a method and system for detecting small target defects in power scenarios based on dynamic task alignment is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0005] In view of this, the present invention provides a method and system for detecting small target defects in power scenarios based on dynamic task alignment, which improves the efficiency and accuracy of small target inspection in the field of power inspection.
[0006] To achieve the above objectives, the present invention adopts the following technical solution: a method for detecting small target defects in power scenarios based on dynamic task alignment, comprising:
[0007] Acquire power scenario data;
[0008] A defect detection model for power scenarios is constructed; the defect detection model adopts HRNet as the backbone network, and sequentially connects an attention-guided task collaboration module, a switchable multi-scale convolutional block, and a flexible inverted bottleneck module;
[0009] The power scenario data is input into the defect detection model to perform defect detection on the power scenario.
[0010] Preferably, the backbone network outputs a joint feature map through multi-resolution parallel branching and cross-stage feature fusion;
[0011] Based on the joint feature map, the task collaboration module generates regression prediction maps and classification prediction maps for target detection through task decomposition, and uses cross-task attention mechanism and dynamic anchor boxes to align task features.
[0012] The switchable multi-scale convolutional block groups the input features by channel grouping, adaptively selects convolutional kernels of different scales through a gating selection mechanism, and uses graph convolution to fuse cross-group features.
[0013] The flexible inverted bottleneck module integrates inverted bottleneck, ConvNext, and depthwise separable convolution variants, and dynamically optimizes the computation path through neural architecture search.
[0014] Preferably, the cross-stage feature fusion achieves multi-scale information deep fusion through a cross-stage bidirectional feature exchange mechanism, including:
[0015] Low-resolution branches pass semantic context information to high-resolution branches;
[0016] High-resolution branches transmit spatial detail features to low-resolution branches;
[0017] In this process, each network stage performs dynamic feature interaction through cross-branch feature projection using a 1×1 convolution, expressed as follows:
[0018]
[0019] In the formula, For resolution adaptation functions, F represents the joint feature. i in Let i represent the input feature, i represent the feature level, and k represent the target level.
[0020] Preferably, based on the joint feature map, the task collaboration module generates a regression prediction map and a classification prediction map for target detection through task decomposition, including:
[0021] Using the joint feature map as input, task decoupling is performed on the joint feature map: the regression branch generates a target box regression prediction map, and the classification branch generates a defect classification prediction map; at the same time, regression correction features and classification correction features are derived.
[0022] Preferably, task feature alignment is performed using a cross-task attention mechanism and dynamic anchor boxes, including:
[0023] In the dynamic anchor frame path, the regression correction features are global average pooled and then normalized parameters are output through two fully connected networks and activation functions to dynamically adjust the preset anchor frame size so that the minimum anchor frame fits the target of the power scenario.
[0024] In the cross-attention path, the spatial attention branch generates a heatmap of the target region through convolution-normalization operations, while the channel attention branch generates feature channel weights through a fully connected layer.
[0025] Preferably, the switchable multi-scale convolutional block groups the input features by channel, adaptively selects convolutional kernels of different scales through a gating selection mechanism, and performs cross-group feature fusion using graph convolution, including:
[0026] The input feature map is evenly divided into three channel subsets, which are then fed into convolutional kernels of three different scales: 1×1, 3×3, and 5×5, for parallel processing. Each scale branch uses depthwise separable convolution.
[0027] An attention-based gating selection mechanism is introduced, and dynamic weights are generated through a two-level fully connected network;
[0028] Feature fusion is performed using a channel attention-enhanced weighted summation method;
[0029] A lightweight graph convolutional network is used to construct a dynamic adjacency matrix based on cosine similarity to model the relationships between channels.
[0030] Preferably, the flexible inverted bottleneck module integrates inverted bottleneck, ConvNext, and depthwise separable convolution variants, and dynamically optimizes the computation path through neural architecture search, including:
[0031] The flexible inverted bottleneck module employs two depthwise separable convolutional paths and automatically converges to the optimal configuration by searching and optimizing the gating parameters through a differentiable architecture.
[0032] The extended layer employs a hybrid operation strategy, supporting three feature transformation modes simultaneously, and adaptively adjusts the receptive field through spatial pyramid convolution.
[0033] Preferably, the depth-separable convolutional path includes a pre-path and a post-path; the pre-path is located before the expansion layer, and the post-path is located between the expansion layer and the projection layer.
[0034] Preferably, the small target defect detection system for power scenarios based on dynamic task alignment includes: a data acquisition module for acquiring power scenario data;
[0035] The model building module is used to build a defect detection model for power scenarios. The defect detection model adopts HRNet as the backbone network, and is connected in sequence to an attention-guided task collaboration module, a switchable multi-scale convolutional block, and a flexible inverted bottleneck module.
[0036] The target detection module is used to input the power scene data into the defect detection model to perform defect detection on the power scene.
[0037] As can be seen from the above technical solutions, compared with the prior art, the present invention discloses a method and system for detecting small target defects in power scenarios based on dynamic task alignment. (1) Model backbone network: Through multi-resolution parallel branches and cross-stage feature fusion, high-resolution details are maintained while extracting deep semantic information, and joint feature maps are output. (2) Attention-guided task collaboration module: Receives joint features from the backbone network, generates regression prediction maps and classification prediction maps through task decomposition, and achieves task feature alignment by using cross-task attention mechanism and dynamic anchor box generation. (3) Switchable multi-scale convolutional block: Channel grouping is performed on the input features, and different scale convolutional kernels (1×1 / 3×3 / 5×5) are adaptively selected through gating selection mechanism, and cross-group feature fusion is achieved by using graph convolution. (4) Flexible inverted bottleneck module: As an extensible unit of the backbone network, inverted bottleneck, ConvNext and depth separable convolution variants are integrated, and the neural architecture search optimization module enables the strategy to introduce them into the backbone network to improve model adaptability and improve the efficiency and accuracy of small target inspection in the field of power inspection. Attached Figure Description
[0038] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0039] Figure 1 This is a schematic diagram of the model backbone network provided in an embodiment of the present invention.
[0040] Figure 2 This is a schematic diagram of an attention-guided task collaboration module provided in an embodiment of the present invention.
[0041] Figure 3 This is a schematic diagram of a task alignment module provided in an embodiment of the present invention.
[0042] Figure 4 This is a schematic diagram of the power scenario risk identification results provided in an embodiment of the present invention.
[0043] Figure 5 This is a schematic diagram of the process for detecting small target defects in power scenarios based on dynamic task alignment, provided in an embodiment of the present invention. Detailed Implementation
[0044] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0045] This invention discloses a method for detecting small target defects in power scenarios based on dynamic task alignment, such as... Figure 5 As shown, it includes:
[0046] Acquire power scenario data;
[0047] A defect detection model for power scenarios is constructed; the defect detection model adopts HRNet as the backbone network, and sequentially connects an attention-guided task collaboration module, a switchable multi-scale convolutional block, and a flexible inverted bottleneck module;
[0048] The power scenario data is input into the defect detection model to perform defect detection on the power scenario.
[0049] Specifically, the backbone network outputs a joint feature map through multi-resolution parallel branching and cross-stage feature fusion;
[0050] Based on the joint feature map, the task collaboration module generates regression prediction maps and classification prediction maps for target detection through task decomposition, and uses cross-task attention mechanism and dynamic anchor boxes to align task features.
[0051] The switchable multi-scale convolutional block groups the input features by channel grouping, adaptively selects convolutional kernels of different scales through a gating selection mechanism, and uses graph convolution to fuse cross-group features.
[0052] The flexible inverted bottleneck module integrates inverted bottleneck, ConvNext, and depthwise separable convolution variants, and dynamically optimizes the computation path through Neural Architecture Search (NAS).
[0053] Specifically, the cross-stage feature fusion achieves deep fusion of multi-scale information through a cross-stage bidirectional feature exchange mechanism, including:
[0054] Low-resolution branches pass semantic context information to high-resolution branches;
[0055] High-resolution branches transmit spatial detail features to low-resolution branches;
[0056] In this process, each network stage performs dynamic feature interaction through cross-branch feature projection using a 1×1 convolution, expressed as follows:
[0057]
[0058] In the formula, For resolution adaptation functions, F represents the joint feature. i in Let i represent the input feature, i represent the feature level, and k represent the target level.
[0059] like Figure 1 As shown, the backbone network used in this embodiment of the invention completely solves the problem of small target feature loss caused by layer-by-layer downsampling in traditional detection models through a revolutionary multi-resolution parallel architecture. Its core lies in constructing four levels of parallel feature processing branches (the main branch always maintains the original resolution of 1024×1024, and the secondary branches are downsampled to 512×512, 256×256, and 128×128 respectively), and achieving deep fusion of multi-scale information through a cross-stage bidirectional feature exchange mechanism: the low-resolution branch transmits semantic context information to the high-resolution branch (through bilinear upsampling fusion), and the high-resolution branch transmits spatial detail features to the low-resolution branch (through 3×3 convolution downsampling connections with a stride of 2). Specifically, each network stage uses a 1×1 convolution cross-branch feature projection formula:
[0060]
[0061] Here, T is the resolution adaptation function, which performs dynamic feature interaction to ensure that minor electrical defects such as insulator cracks (minimum 8×8 pixels) and bolt detachment retain their complete spatial structure in the deep network. To address the specific needs of densely populated power equipment scenarios, the feature exchange frequency was optimized (one cross-branch fusion is performed every two residual blocks, a 50% improvement over the original architecture). Simultaneously, lightweight modifications were implemented: depthwise separable convolutions were used instead of standard convolutions in the 128×128 branch, and the number of channels in the feature exchange layer was compressed by 50%, reducing the total number of parameters from 21.6M to 8.3M.
[0062] Specifically, based on the joint feature map, the task collaboration module generates regression prediction maps and classification prediction maps for object detection through task decomposition, including:
[0063] Using the joint feature map as input, task decoupling is performed on the joint feature map: the regression branch generates a target box regression prediction map, and the classification branch generates a defect classification prediction map; at the same time, regression correction features and classification correction features are derived.
[0064] Specifically, task feature alignment is achieved using cross-task attention mechanisms and dynamic anchor boxes, including:
[0065] In the dynamic anchor frame path, the regression correction features are global average pooled and then normalized parameters are output through two fully connected networks and activation functions to dynamically adjust the preset anchor frame size so that the minimum anchor frame fits the target of the power scenario.
[0066] In the cross-attention path, the spatial attention branch generates a heatmap of the target region through convolution-normalization operations, while the channel attention branch generates feature channel weights through a fully connected layer.
[0067] like Figure 2 As shown, the joint features in this embodiment of the invention are first processed by the task decomposition module to generate two key prediction maps for object detection: regression prediction map A and classification prediction map B. These two prediction maps are used for object location regression and category classification tasks, respectively. Simultaneously, two new feature maps C and D are generated in the branch of the task decomposition module. The main function of these two feature maps is to further align and correct regression prediction map A and classification prediction map B. Through alignment and correction operations, the feature representations between the two tasks can be better coordinated, thereby improving the overall performance of the model.
[0068] The attention-guided task collaboration module proposed in this invention innovatively integrates three core technologies: dynamic anchor box generation, cross-attention mechanism, and task alignment correction. Through multi-level feature interaction, it significantly optimizes the performance of small power target detection. The module utilizes high-resolution feature maps output by HRNet. As input, task decoupling is first achieved through dual-path 3×3 convolution: the regression branch generates the target bounding box coordinate prediction map. (Including four positioning parameters: x, y, w, h), the classification branch generates a defect category prediction map. (Corresponding to 8 types of power risk probabilities); Meanwhile, from F hr Derived regression corrected feature C = δ(Conv1×1) alignreg (F hr )) and classification correction feature D=δ(Conv1×1 aligncls (F hr ))(δ is the ReLU activation function, and the feature dimensions are all 1. In the dynamic anchor box optimization path, the regression correction feature C is compressed into a 64-dimensional vector through global average pooling, and then passed through a two-layer fully connected network FC. 128 → Normalized parameters output by FC4 and Sigmoid activation functions:
[0069]
[0070] Where, Δ scale ∈[0.5,2] controls the scaling of the anchor frame, Δ ratio Adjust the aspect ratio ∈[0.7,1.4] to dynamically adjust the preset anchor frame size accordingly:
[0071]
[0072] This allows the smallest anchor frame to fit bolt-level targets of 8×8 pixels (traditional fixed anchor frames have a deviation rate of 38.2% in power scenarios).
[0073] In the cross-attention path, the spatial attention branch generates a heatmap A of the target region through convolution-normalization operations. space =Softmax(Conv3×3) sp (D)), the channel attention branch generates feature channel weights through a fully connected layer. The two methods work synergistically to enhance regression features. This operation increases the feature response intensity at insulator string connections by 3.7 times in transmission line scenarios. Finally, a learnable parameter matrix is used to align the classification and regression tasks.
[0074]
[0075] The residual connections preserve the original feature information, and the task coupling matrix dynamically adjusts the weight allocation for classification and regression. Experiments show that this module reduces the average localization error of targets smaller than 32×32 pixels from 6.8 pixels in traditional methods to 4.6 pixels (a reduction of 32.4%) in power-intensive scenarios, while also reducing the false detection rate of high-risk targets such as insulator cracks by 21.5%.
[0076] Specifically, the switchable multi-scale convolutional block groups the input features by channel, adaptively selects convolutional kernels of different scales through a gating selection mechanism, and performs cross-group feature fusion using graph convolution, including:
[0077] The input feature map is evenly divided into three channel subsets, which are then fed into convolutional kernels of three different scales: 1×1, 3×3, and 5×5, for parallel processing. Each scale branch uses depthwise separable convolution.
[0078] An attention-based gating selection mechanism is introduced, and dynamic weights are generated through a two-level fully connected network;
[0079] Feature fusion is performed using a channel attention-enhanced weighted summation method;
[0080] A lightweight graph convolutional network is used to construct a dynamic adjacency matrix based on cosine similarity to model the relationships between channels and improve the feature representation capability.
[0081] like Figure 3 As shown, the Switchable Multi-Scale Convolution Block is an innovative dynamic feature extraction structure that significantly improves the ability of neural networks to detect multi-scale defects in power equipment through multi-scale parallel processing and adaptive feature fusion mechanism.
[0082] The core design of this module includes three key technical aspects:
[0083] First, input feature map F in ∈R (H×W×C) The channel is evenly divided into three subsets, which are then fed into convolutional kernels of different scales (1×1, 3×3, and 5×5) for parallel processing. Each scale branch employs depthwise separable convolution to reduce computational complexity. The mathematical expression for this is:
[0084]
[0085] Where DWConv represents depthwise convolution, and PWConv represents pointwise convolution. This indicates a feature splicing operation.
[0086] Secondly, the module introduces an attention-based gating selection mechanism, generating dynamic weights through a two-stage fully connected network: the first stage uses the globally average pooled feature GAP(F) to perform the gating. in )∈R C Mapped to the intermediate layer z = ReLU(W1·GAP(F) in The second-stage output normalized gating weights are G = Softmax(W2·z + b2) ∈ R. 3 , where W1∈R (C / 4×C) W2∈R (3×C / 4) These are trainable parameters. The final feature fusion uses a channel attention-enhanced weighted summation method: Where ⊙ represents channel-level multiplication.
[0087] To further enhance feature representation capabilities, the module also includes a lightweight graph convolutional network (GCN), which constructs a dynamic adjacency matrix A based on cosine similarity. ij =cos(F i ,F j The relationship between channels is modeled using F + I, and its update formula is F. final =LayerNorm(σ(D) (-1 / 2) AD (-1 / 2) F merged W gcn )), where W gcn ∈R (C×C) Let σ be the learnable transformation matrix, and let σ represent the SiLU activation function.
[0088] Experiments show that, in power equipment defect detection tasks, compared with the traditional fixed convolution kernel design, this structure improves the detection accuracy of insulator cracks (average size 15×15 pixels) from 78.3% to 92.6% while maintaining the same number of parameters, and increases the recall rate of loose bolts (8×8 pixels) by 34.5%. Simultaneously, the computational load during inference is reduced by approximately 28% through a dynamic gating selection mechanism. This design is particularly suitable for the characteristics of power inspection scenarios, such as large differences in equipment defect scale and complex backgrounds, providing an effective solution for small target detection.
[0089] Specifically, the flexible inverted bottleneck module integrates inverted bottleneck, ConvNext, and depthwise separable convolution variants. It dynamically optimizes the computational path through neural architecture search to refine the features of each branch of HRNet. Simultaneously, it adapts the NAS-optimized computational path to the multi-resolution feature distribution of HRNet, complementing the switchable multi-scale convolution. This includes:
[0090] The flexible inverted bottleneck module employs two depthwise separable convolutional paths and automatically converges to the optimal configuration by searching and optimizing the gating parameters through a differentiable architecture.
[0091] The extended layer employs a hybrid operation strategy, supporting three feature transformation modes simultaneously, and adaptively adjusts the receptive field through spatial pyramid convolution.
[0092] Specifically, the depth-separable convolutional path includes a pre-path and a post-path; the pre-path is located before the expansion layer, and the post-path is located between the expansion layer and the projection layer (1*1 convolutional layer).
[0093] The Flexible Inverted Bottleneck Block is a highly configurable neural network building block. Its core innovation lies in dynamically optimizing computational paths through Neural Architecture Search (NAS). Based on reinforcement learning, it designs modules to search for the optimal network structure sequence in the search space and uses a gating mechanism to dynamically select computational paths during inference, perfectly balancing computational efficiency and feature representation capabilities. This module makes three groundbreaking improvements on the classic inverted bottleneck structure of MobileNetV2: multimodal extension layers generate diverse features; switchable depthwise convolutional paths dynamically select processing channels based on feature characteristics; and dynamic receptive field adjustment scales the results of the first two steps. These three elements form a closed loop of "feature generation → processing → calibration" through the NAS gating mechanism.
[0094] Switchable depthwise convolution path:
[0095] The module contains two optional depthwise separable convolutional paths: a pre-path (before the expansion layer) and a post-path (between the expansion and projection layers), whose activation states are controlled by gating coefficients g1, g2 ∈ {0, 1}.
[0096]
[0097] Where k is the dynamically selected kernel size (3 / 5 / 7), F in For input features, F expand To extend features, F pre For input features, F post Features after input.
[0098] The gating parameters are optimized by differentiable architecture search (DARTS). Specifically, the gating parameters are transformed into a continuous probability distribution, and the probability distribution of network weights and gating parameters is optimized simultaneously by gradient descent. Finally, the optimal gating state is determined based on the probability to complete the parameter optimization. In the power equipment defect detection task, it automatically converges to the optimal configuration (measured g1 = 0.87, g2 = 0.42).
[0099] Multimodal extension layer:
[0100] The extension phase employs a hybrid operation strategy, supporting three feature transformation modes simultaneously. For local image features (such as object edges and textures), ConvNext is called first; for global semantics (such as scene classification), FFN can be enabled. When computational resources are limited, MBConv can serve as an efficient alternative.
[0101] F expand =α·ConvNext(F in )+β·FFN(F in )+γ·MBConv(F in );
[0102] The weight coefficients α, β, and γ are dynamically generated by the hypernetwork (NAS reinforcement learning algorithm) and satisfy α + β + γ = 1. Experiments show that in substation equipment images, the model automatically prefers the ConvNext mode (α = 0.63), while in transmission line scenarios it tends to favor MBConv (γ = 0.71).
[0103] Dynamic receptive field modulation:
[0104] Adaptive adjustment of the receptive field through spatial pyramid convolution (SPConv):
[0105]
[0106] Expansion rate d and weight w i Determined by the feature content: First, calculate the global context vector c = GAP(F) in Then, parameters are generated through a lightweight decision network:
[0107] (d,w1,w2,w3)=MLP(c;θ).
[0108] In one specific embodiment of the present invention, a small target defect detection system for power scenarios based on dynamic task alignment includes:
[0109] The data acquisition module is used to acquire power scenario data;
[0110] The model building module is used to build a defect detection model for power scenarios. The defect detection model adopts HRNet as the backbone network, and is connected in sequence to an attention-guided task collaboration module, a switchable multi-scale convolutional block, and a flexible inverted bottleneck module.
[0111] The target detection module is used to input the power scene data into the defect detection model to perform defect detection on the power scene.
[0112] In one specific embodiment of the present invention, such as Figure 1-3 As shown in the figure, the small target defect detection method for power scenarios based on dynamic task alignment proposed in this embodiment of the invention achieves accurate detection of small target defects through a three-level framework of high-resolution feature preservation, task collaborative alignment, and dynamic scale adaptation. The specific technical solution is as follows:
[0113] I. Backbone Network: Multi-resolution Parallel Branching and Cross-Stage Feature Fusion
[0114] The backbone network adopts the HRNet architecture. Figure 1 This approach preserves small target features through four levels of parallel branches with varying resolutions. The main branch maintains the original image resolution to ensure that spatial details of minute defects such as insulator cracks (8×8 pixels) are not lost. The secondary branches gradually extract deep semantic information through downsampling, while simultaneously achieving high- and low-resolution feature fusion through cross-branch interactions indicated by upsampling and downsampling arrows.
[0115] Specifically, each branch dynamically exchanges features through 1×1 convolutional feature projection. Figure 1 The 1×1 convolutional feature projection is used, and cross-branch fusion is performed once for every two residual blocks. This allows the high-resolution branch to obtain semantic information while the low-resolution branch retains spatial details. For lightweight requirements, the 128×128 branch uses depthwise separable convolution, and the final output is a joint feature map containing small object details and semantic information.
[0116] II. Attention-Guided Task Collaboration Module: Task Alignment and Dynamic Anchor Box Optimization
[0117] This module ( Figure 2 Using the joint features output by the backbone network as input, the synergy between classification and regression tasks is optimized through three steps: task decomposition, feature correction, and dynamic alignment.
[0118] First, the joint features are processed by the task decomposition module to generate the core outputs: regression prediction map A (location regression), classification prediction map B (category classification), and auxiliary feature maps C (regression correction) and D (classification correction). Auxiliary feature map C is extracted using 1×1 convolution and used for dynamic anchor box generation—after global average pooling compression, the anchor box scale and aspect ratio parameters are output through a fully connected network, allowing the anchor box to adapt to small targets such as 8×8 pixel bolts (solving the deviation problem of traditional fixed anchor boxes). Auxiliary feature map D enhances its interaction with classification features through a cross-attention mechanism.
[0119] Ultimately, regression prediction map A is optimized using dynamic anchor boxes, and classification prediction map B is optimized using feature fusion. Spatial alignment is achieved through task alignment correction.
[0120] III. Switchable Multi-Scale Convolutional Blocks: Dynamic Scale Adaptation and Feature Enhancement (corresponding to) Figure 3 )
[0121] This module ( Figure 3 For multi-scale defect detection, feature extraction optimization is achieved through multi-branch parallel processing, dynamic weight selection, and cross-channel fusion.
[0122] The input features are first uniformly divided into three subsets by channels, and then fed into 1×1, 3×3, and 5×5 convolutional kernels, respectively. Each branch uses depthwise separable convolutions to reduce computation. Subsequently, dynamic weights (gating mechanism) are generated through fully connected layers and Softmax, and channel attention is used to weight features at different scales.
[0123] Finally, the features are fused across channels via graph convolution (GCN) to strengthen channel correlation and output enhanced features.
[0124] This invention proposes a detection framework that integrates high-resolution feature preservation, task-based collaborative alignment, and dynamic scale adaptation. First, HRNet is used as the backbone network, employing multi-resolution parallel branches and cross-stage feature fusion to capture semantic information while preserving spatial details of small targets. Second, an attention-guided task collaboration module is designed, utilizing a cross-task attention mechanism to enhance feature interaction between classification and regression tasks, and dynamically generating anchor box parameters adapted to small targets through conditional convolution. Finally, switchable multi-scale convolutional blocks are constructed, using a gating selection mechanism to adaptively select the optimal convolutional kernel scale, combined with graph convolution to achieve efficient fusion of cross-channel features.
[0125] Experimental results show that after only 12 training cycles on the power scenario dataset, the method achieves an AP@.5 score of 90.2%, a 6.2% improvement over the original YOLOv8, with a 40% reduction in parameters. Compared with mainstream models such as Gold-Yolo and FADC, this embodiment demonstrates superior performance in both small target detection accuracy and model lightweighting. In particular, it reduces the false negative rate by more than 35% in detecting typical defects such as insulator damage and floating debris on conductors, providing a solution that combines efficiency and accuracy for intelligent inspection of power equipment.
[0126] The proposed small target detection method for power systems based on dynamic task alignment in this invention has been systematically validated on a dedicated power scenario dataset (containing eight types of defects, including insulator breakage and foreign objects in conductors). Training was performed using YOLOv8 as the baseline framework, with a batch size of 32 and 12 training epochs. Experimental results are as follows: Figure 4 As shown in Table 1, this embodiment demonstrates significant advantages in both accuracy and efficiency compared to mainstream detection models.
[0127] Table 1 Comparison Experiment of Backbone Networks
[0128] algorithm AP@.5 AP@[.5,.95] Params(M) CSPDarknet 84.0% 66.8% 3.01 ResNet50-FPN 79.3% 58.2% 4.15 Ours 90.2% 51.3% 2.08
[0129] The method of this embodiment was compared with the state-of-the-art (SOTA) method in power scenarios on a test set, and the detection results are shown in Table 2. Compared with other methods, the model of this embodiment achieves superior detection performance.
[0130] Table 2 Comparison of State-of-the-Art (SOTA) in Power Scenarios
[0131]
[0132] Typical scenario metrics:
[0133] 1) Insulator defect detection
[0134] Crack detection rate: 92.6% (78.3% using traditional methods);
[0135] Minimum detectable size: 8×8 pixels (originally 16×16);
[0136] 2) The inspection results of the transmission lines are shown in Table 3.
[0137] Table 3 Comparison of Inspection Results
[0138] Defect types Improved recall rate Decreasing false positive rate Floating objects on wires +34.5% -28.7% Tower corrosion +22.1% -19.3% Insulator damage +38.2% -31.5%
[0139] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.
[0140] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for detecting small target defects in power scenarios based on dynamic task alignment, characterized in that, include: Acquire power scenario data; Construct a defect detection model for power scenarios; The defect detection model uses HRNet as the backbone network, and sequentially connects an attention-guided task collaboration module, a switchable multi-scale convolutional block, and a flexible inverted bottleneck module. The power scenario data is input into the defect detection model to perform defect detection on the power scenario.
2. The method for detecting small target defects in power scenarios based on dynamic task alignment according to claim 1, characterized in that, The backbone network outputs a joint feature map through multi-resolution parallel branching and cross-stage feature fusion. Based on the joint feature map, the task collaboration module generates regression prediction maps and classification prediction maps for target detection through task decomposition, and uses cross-task attention mechanism and dynamic anchor boxes to align task features. The switchable multi-scale convolutional block groups the input features by channel grouping, adaptively selects convolutional kernels of different scales through a gating selection mechanism, and uses graph convolution to fuse cross-group features. The flexible inverted bottleneck module integrates inverted bottleneck, ConvNext, and depthwise separable convolution variants, and dynamically optimizes the computation path through neural architecture search.
3. The method for detecting small target defects in power scenarios based on dynamic task alignment according to claim 2, characterized in that, The cross-stage feature fusion achieves deep fusion of multi-scale information through a cross-stage bidirectional feature exchange mechanism, including: Low-resolution branches pass semantic context information to high-resolution branches; High-resolution branches transmit spatial detail features to low-resolution branches; In this process, each network stage performs dynamic feature interaction through cross-branch feature projection using a 1×1 convolution, expressed as follows: In the formula, For resolution adaptation functions, Indicates joint features, Let i represent the input feature, i represent the branch level, and k represent the target level.
4. The method for detecting small target defects in power scenarios based on dynamic task alignment according to claim 2, characterized in that, Based on the joint feature map, the task collaboration module generates regression prediction maps and classification prediction maps for object detection through task decomposition, including: Using the joint feature map as input, task decoupling is performed on the joint feature map: the regression branch generates a target box regression prediction map, and the classification branch generates a defect classification prediction map; at the same time, regression correction features and classification correction features are derived.
5. The method for detecting small target defects in power scenarios based on dynamic task alignment according to claim 4, characterized in that, Task feature alignment is achieved using cross-task attention mechanisms and dynamic anchor boxes, including: In the dynamic anchor frame path, the regression correction features are global average pooled and then normalized parameters are output through two fully connected networks and activation functions to dynamically adjust the preset anchor frame size so that the minimum anchor frame fits the target of the power scenario. In the cross-attention path, the spatial attention branch generates a heatmap of the target region through convolution-normalization operations, while the channel attention branch generates feature channel weights through a fully connected layer.
6. The method for detecting small target defects in power scenarios based on dynamic task alignment according to claim 1, characterized in that, The switchable multi-scale convolutional block groups the input features by channel, adaptively selects convolutional kernels of different scales through a gating selection mechanism, and performs cross-group feature fusion using graph convolution, including: The input feature map is evenly divided into three channel subsets, which are then fed into convolutional kernels of three different scales: 1×1, 3×3, and 5×5, for parallel processing. Each scale branch uses depthwise separable convolution. An attention-based gating selection mechanism is introduced, and dynamic weights are generated through a two-level fully connected network; Feature fusion is performed using a channel attention-enhanced weighted summation method; A lightweight graph convolutional network is used to construct a dynamic adjacency matrix based on cosine similarity to model the relationships between channels.
7. The method for detecting small target defects in power scenarios based on dynamic task alignment according to claim 1, characterized in that, The flexible inverted bottleneck module integrates inverted bottleneck, ConvNext, and depthwise separable convolution variants, and dynamically optimizes the computation path through neural architecture search, including: The flexible inverted bottleneck module employs two depthwise separable convolutional paths and automatically converges to the optimal configuration by searching and optimizing the gating parameters through a differentiable architecture. The extended layer employs a hybrid operation strategy, supporting three feature transformation modes simultaneously, and adaptively adjusts the receptive field through spatial pyramid convolution.
8. The method for detecting small target defects in power scenarios based on dynamic task alignment according to claim 7, characterized in that, The depth-separable convolutional path includes a front path and a back path; the front path is located before the expansion layer, and the back path is located between the expansion layer and the projection layer.
9. A small target defect detection system for power scenarios based on dynamic task alignment, employing the small target defect detection method for power scenarios based on dynamic task alignment as described in any one of claims 1-8, characterized in that, include: The data acquisition module is used to acquire power scenario data; The model building module is used to build defect detection models for power scenarios; The defect detection model uses HRNet as the backbone network, and sequentially connects an attention-guided task collaboration module, a switchable multi-scale convolutional block, and a flexible inverted bottleneck module. The target detection module is used to input the power scene data into the defect detection model to perform defect detection on the power scene.