Insulator defect detection method based on mixed attention and multi-scale features
By combining residual neural network, hybrid attention mechanism and improved Transformer model, the problem of insufficient multi-scale feature extraction in insulator defect detection is solved, and high-precision and stable small-scale defect detection is achieved, which is suitable for smart grid systems.
Patent Information
- Application Number
- CN202510182037.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-07-18
AI Technical Summary
The prior art is difficult to effectively extract multi-scale features in insulator defect detection, resulting in low detection accuracy of small defects and easy to miss or miss detection in complex backgrounds, making it difficult to meet the high accuracy and real-time requirements of smart grid systems.
The bottom-up residual neural network is used to combine with the top-down reverse residual neural network, and the multi-head mixed attention mechanism and improved Transformer model, combined with the Hungarian bipartite matching algorithm, the comprehensive utilization of multi-scale features and accurate defect detection are achieved.
It improves the accuracy and robustness of small defect detection, can maintain high accuracy and stability in complex backgrounds, and meets the real-time detection needs of smart grid systems.
Smart Images

Figure CN120339160A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of electric power inspection, and in particular relates to an insulator defect detection method based on hybrid attention and multi-scale features. Background Art
[0002] Traditional insulator defect detection methods usually rely on manual inspection, image processing algorithms, and machine learning models based on convolutional neural networks (CNNs). These methods perform relatively well in detecting conventional and large defects, but often face problems of insufficient feature extraction and low detection accuracy when processing insulator images containing small target defects. In addition, existing methods are easily disturbed when dealing with complex backgrounds, lighting changes, and perspective diversity, resulting in missed detections or false detections. With the development of smart grid systems, higher requirements are placed on the accuracy and real-time performance of insulator detection. Existing technologies still face great challenges in balancing the accuracy of small defect detection and model computational efficiency. Therefore, how to improve feature extraction capabilities, improve the model's ability to process multi-scale information, and maintain high computational efficiency while performing high-precision detection has become a key issue that needs to be urgently addressed in the field of insulator defect detection. Summary of the invention
[0003] The purpose of the present invention is to overcome the defects of the above-mentioned prior art and to provide an insulator defect detection method based on hybrid attention and multi-scale features.
[0004] The purpose of the present invention can be achieved by the following technical solutions:
[0005] The present invention provides an insulator defect detection method based on hybrid attention and multi-scale features, comprising the following steps:
[0006] Acquire an image of an insulator of a transmission line to be inspected, and preprocess the image of the insulator;
[0007] A multi-scale feature extraction module is used to perform feature extraction and feature fusion on the preprocessed image, wherein the multi-scale feature extraction module includes a bottom-up residual neural network and a top-down reverse residual neural network, wherein the bottom-up residual neural network extracts feature maps of different levels through multiple residual blocks, and the feature maps of different levels include low-level texture feature maps, local space feature maps, mid-level semantic feature maps, and global semantic feature maps; the top-down reverse residual neural network horizontally fuses the high-level feature map with the low-level feature map through an upsampling operation, and sequentially outputs a first fused feature map, a second fused feature map, a third fused feature map, and a fourth fused feature map;
[0008] The fused feature maps are input into a multi - head hybrid attention mechanism for processing. The multi - head hybrid attention mechanism includes local attention and global attention. Among them, local attention processes the first fused feature map, the second fused feature map, and the third fused feature map, and global attention processes the fourth fused feature map;
[0009] The feature maps processed by the multi - head hybrid attention mechanism are input into an improved Transformer model for processing, and the probability distribution of the defect categories and the bounding box coordinates of the insulator image are output. The defect detection of the insulator is realized according to the output probability distribution of the defect categories and the bounding box coordinates.
[0010] Furthermore, the bottom - up residual neural network is a multi - scale feature extraction module based on the pre - trained ResNet - 50 architecture, including multiple residual blocks. Each residual block gradually extracts features from the input image through convolutional layers, including:
[0011] The Conv2 stage: In this stage, a convolutional kernel is used to perform a convolutional operation on the input image to extract low - level texture features. The size of the convolutional kernel for the convolutional operation is 3×3, the stride is 1, and a low - level texture feature map with a resolution of H / 4×W / 4 and 256 channels is generated;
[0012] The Conv3 stage: This stage extracts local spatial features and generates a local spatial feature map with a resolution of H / 8×W / 8 and 512 channels;
[0013] The Conv4 stage: This stage extracts intermediate semantic features and generates an intermediate semantic feature map with a resolution of H / 16×W / 16 and 1024 channels;
[0014] The Conv5 stage: This stage extracts global semantic features and generates a global semantic feature map with a resolution of H / 32×W / 32 and 2048 channels.
[0015] Furthermore, each residual block gradually extracts features from the input image through convolutional layers. The feature extraction process of the residual block is described as:
[0016] F i (X) = R i (F i-1 (X)) = σ(W i *F i-1 (X)+b i )+F i-1 (X)
[0017] where F i (X) represents the feature map output by the residual block in the i - th Convi stage, and F i-1(X) represents the feature map output by the residual block in the Convi-1 stage during the Convi-1 stage, R i represents the residual block in the Convi stage, σ is the activation function, W i and b i are the weight and bias of the convolutional layer respectively, * represents the convolution operation, and the value range of i is 2, 3, 4, and 5, corresponding to the feature extraction processes of the Conv2, Conv3, Conv4, and Conv5 stages in the bottom-up residual neural network respectively.
[0018] Furthermore, the top-down inverse residual neural network performs horizontal fusion of high-level feature maps and low-level feature maps through upsampling operations, specifically including:
[0019] Generate the fourth fusion feature map by performing convolution operation and batch normalization on the global semantic feature map;
[0020] Adjust the size of the global semantic feature map through upsampling operation and horizontally fuse it with the intermediate semantic feature map to generate the third fusion feature map of H / 16×W / 16, with the number of channels being 256;
[0021] Adjust the size of the intermediate semantic feature map through upsampling operation and horizontally fuse it with the local spatial feature map to generate the second fusion feature map of H / 8×W / 8, with the number of channels being 256;
[0022] Adjust the size of the local spatial feature map through upsampling operation and horizontally fuse it with the low-level texture feature map to generate the first fusion feature map of H / 4×W / 4, with the number of channels being 256.
[0023] Furthermore, the top-down inverse residual neural network performs horizontal fusion of high-level feature maps and low-level feature maps through upsampling operations, described as:
[0024] G i (F i (X)) = ReLU(W up *U i+1 (F i+1 (X)) + b up +F i (X))
[0025] Among them, G i (F i (X)) is the feature map after fusion of the i-th layer, representing the feature output after upsampling, fusion, and activation, F i (X) represents the feature map output by the residual block in the Convi stage during the Convi stage, W up *U i+1 represents the upsampling operation, b upis the bias term in the upsampling operation, and ReLU represents the ReLU activation function.
[0026] Further, the local attention processes the first fused feature map, the second fused feature map, and the third fused feature map, which specifically includes:
[0027] Set a local window of a fixed size on each input feature map, and determine the position and range of each window according to the feature point coordinates in the input feature map;
[0028] Using the self-attention mechanism, by calculating the correlation between pixels within each window, obtain the attention score of each pixel, and obtain the attention score matrix within each window;
[0029] By weighting the pixels within each local region in each feature map with the obtained attention score matrix, obtain the feature map after weighted processing;
[0030] Merge the feature maps after weighted processing of all local windows to obtain the feature map that fuses the local attention features, and obtain the first fused feature map, the second fused feature map, and the third fused feature map after local attention processing.
[0031] Further, the global attention processes the fourth fused feature map, which specifically includes:
[0032] Map the number of channels of the fourth fused feature map to the corresponding dimension through convolution operation or linear transformation to obtain the query matrix Q, the key matrix K, and the value matrix V;
[0033] Perform a dot product operation on the query matrix Q and the key matrix K to calculate the similarity between each pair of positions, and obtain the attention score matrix;
[0034] Use the Softmax function to normalize the attention score matrix, and perform weighted summation of the normalized attention weight matrix and the value matrix V to obtain the weighted feature representation of each position, and output the fourth fused feature map after global attention processing.
[0035] Further, the improved Transformer model includes an encoder, a decoder, and a prediction unit, and the prediction unit includes 50 preset learnable object query vectors.
[0036] Further, inputting the feature map processed by the multi-head hybrid attention mechanism into the improved Transformer model for processing specifically includes:
[0037] Input the first fused feature map, the second fused feature map, the third fused feature map, and the fourth fused feature map processed by the multi-head hybrid attention mechanism into the encoder, perform non-linear transformation through a feed-forward neural network, and process through residual connection and LayerNorm operations to output a fused feature vector. The formula is as follows:
[0038] E fused = σ(W f ·[E1, E2, E3, E4] + b f )
[0039] where E fused is the fused feature vector output by the encoder, E1, E2, E3, and E4 are the feature vectors of the first fused feature map, the second fused feature map, the third fused feature map, and the fourth fused feature map after non-linear transformation through a feed-forward neural network and processing through residual connection and LayerNorm operations, [E1, E2, E3, E4] represents the concatenation operation of the four feature vectors, W f is a learnable fusion parameter matrix, b f is the bias term, and σ is the activation function;
[0040] Input the fused feature vector E fused output by the encoder into the multi-head self-attention mechanism of the decoder for processing, input the feature vector processed by the multi-head self-attention mechanism into the cross-attention mechanism for processing. The cross-attention mechanism interacts with the key and value vectors of the feature vector processed by the multi-head self-attention mechanism through 50 preset learnable object query vectors, calculates a new feature representation through the cross-attention mechanism, and inputs the feature vector processed by the cross-attention mechanism into a feed-forward neural network for non-linear transformation and activation function processing, and outputs the defect category probability distribution and bounding box coordinates of the insulator image through residual connection and layer normalization operations.
[0041] Furthermore, the multi-scale feature extraction module and the improved Transformer model are trained based on the Hungarian bipartite matching algorithm by combining the intersection over union between the predicted box and the ground truth box. Specifically, it includes:
[0042] During the training process, each time a training image is input, the improved Transformer model generates multiple predicted boxes. Using the Hungarian bipartite matching algorithm, by calculating the matching cost between each predicted box and all ground truth boxes, the optimal matching relationship is selected, and the loss function is calculated according to the optimal matching relationship. By calculating the loss function, the parameters of the multi-scale feature extraction module and the improved Transformer model are adjusted through the gradient descent method to achieve the training of the multi-scale feature extraction module and the improved Transformer model. The formula is as follows:
[0043]
[0044] Among them, L is the loss function, N is the number of generated predicted bounding boxes, M is the number of ground truth bounding boxes in the input training images, cost(i, j) is the matching cost, which is used to measure the difference between the i-th predicted bounding box and the j-th ground truth bounding box, and x ij is a binary variable indicating whether the i-th predicted bounding box matches the j-th ground truth bounding box, and x ij = 1 indicates that the i-th predicted bounding box matches the j-th ground truth bounding box, and x ij = 0 indicates that the i-th predicted bounding box does not match the j-th ground truth bounding box. is the intersection over union between the i-th predicted bounding box b i and the j-th ground truth bounding box , and θ is a preset parameter.
[0045] Compared with the prior art, the present invention has the following advantages:
[0046] (1) The present invention adopts a multi-scale feature extraction module that combines a bottom-up residual neural network and a top-down inverse residual neural network. This module can effectively extract different levels of features in the image, from low-level texture features to global semantic features, and then through lateral fusion of high-level and low-level features, realizes the comprehensive utilization of multi-scale information, can more comprehensively capture the multi-scale features in the insulator image, effectively improves the detection accuracy of small defects, and maintains high robustness in complex backgrounds.
[0047] (2) The present invention introduces a multi-head hybrid attention mechanism that combines local attention and global attention. Among them, local attention can focus on local region features, thereby enhancing the model's ability in small defect detection; global attention enhances the attention to global semantic features and improves the model's ability to process complex backgrounds and global information. By focusing on the joint learning of local and global features, the present invention can optimize feature extraction at different scales, thereby improving the accuracy of defect detection, especially showing more stability when dealing with complex environments such as perspective changes and uneven lighting.
[0048] (3) The present invention utilizes an improved Transformer model, including an encoder, a decoder, and a prediction unit, to process the input multi-scale fusion feature map through effective feature fusion and multi-head self-attention mechanism, and finally outputs the probability distribution of defect categories and the coordinates of the bounding boxes. This design ensures the accurate positioning and classification of defects.
[0049] (4) During the model training process, the present invention uses the Hungarian bipartite matching algorithm to perform optimal matching by calculating the intersection over union between the predicted bounding boxes and the ground truth bounding boxes, thereby accurately calculating the loss function and optimizing the model parameters through the gradient descent method. This algorithm can optimize the matching of the target bounding boxes during the training process, ensuring the correct classification and localization of each predicted bounding box by the model, thereby improving the accuracy and stability of detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 It is the overall model diagram of the present invention;
[0051] Figure 2 It is the multi-scale feature extraction module diagram of the present invention;
[0052] Figure 3 It is the schematic diagram of the multi-head hybrid attention mechanism of the present invention;
[0053] Figure 4 It is the improved Transformer model diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0054] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0055] Embodiment 1:
[0056] This embodiment provides an insulator defect detection method based on hybrid attention and multi-scale features. The specific implementation steps are as follows:
[0057] Obtain the transmission line insulator image to be detected. The image undergoes preprocessing, including operations such as noise removal, illumination normalization, and size adjustment, to ensure that the input image meets the requirements of subsequent processing. This step can effectively remove image noise and eliminate the influence of illumination changes on the image quality, laying a good foundation for subsequent feature extraction and defect detection.
[0058] After the image preprocessing is completed, a multi-scale feature extraction module is used to extract and fuse features from the preprocessed image. This multi-scale feature extraction module includes a bottom-up residual neural network and a top-down inverse residual neural network. In the bottom-up residual neural network, first, multiple residual blocks are used to extract low-level texture features, local spatial features, intermediate semantic features, and global semantic feature maps. Specifically, in the Conv2 stage, low-level texture features are extracted through convolutional kernels to generate a low-level texture feature map with a resolution of H / 4×W / 4; in the Conv3 stage, local spatial features are extracted to generate a local spatial feature map with a resolution of H / 8×W / 8; in the Conv4 stage, intermediate semantic features are extracted to generate an intermediate semantic feature map with a resolution of H / 16×W / 16; in the Conv5 stage, global semantic features are extracted to generate a global semantic feature map with a resolution of H / 32×W / 32. This multi-scale feature extraction can perform a detailed analysis of the image at different levels, effectively enhancing the detection ability of small defects, and at the same time having strong robustness to complex backgrounds. The feature extraction process of the residual block is described as:
[0059] F i (X) = R i (F i-1 (X)) = σ(W i *F i-1 (X) + b i ) + F i-1 (X)
[0060] Among them, F i (X) represents the feature map output by the residual block in the i-th Convi stage, and F i-1 (X) represents the feature map output by the residual block in the (i - 1)-th Convi-1 stage. R i represents the residual block in the i-th Convi stage. σ is the activation function. W i and b i are the weights and biases of the convolutional layer respectively. * represents the convolution operation. The value range of i is 2, 3, 4, and 5, corresponding to the feature extraction processes in the Conv2, Conv3, Conv4, and Conv5 stages of the bottom-up residual neural network respectively.
[0061] A top-down reverse residual neural network is adopted to horizontally fuse high-level feature maps and low-level feature maps through upsampling operations. In this process, first, the global semantic feature map is processed through convolution operations and batch normalization to generate the fourth fused feature map; then, the global semantic feature map and the intermediate semantic feature map are horizontally fused through upsampling operations to generate the third fused feature map of H / 16×W / 16; the intermediate semantic feature map and the local spatial feature map are horizontally fused through upsampling operations to generate the second fused feature map of H / 8×W / 8; finally, the local spatial feature map and the low-level texture feature map are fused to generate the first fused feature map of H / 4×W / 4. This fusion operation can effectively retain feature information at different levels, enabling high-level and low-level features to complement each other fully, further enhancing the ability to identify target defects, described as:
[0062] G i (F i (X)) = ReLU(W up *U i+1 (F i+1 (X)) + b up +F i (X))
[0063] Among them, G i (F i (X)) is the feature map after fusion at the i-th layer, representing the feature output after upsampling, fusion, and activation. F i (X) represents the feature map output by the residual block in the Convi stage at the Convi stage. W up *U i+1 represents the upsampling operation, and b up is the bias term in the upsampling operation. ReLU represents the ReLU activation function.
[0064] The fused feature map is input into the multi-head hybrid attention mechanism for processing. The multi-head hybrid attention mechanism includes local attention and global attention. Among them, local attention processes the first fused feature map, the second fused feature map, and the third fused feature map, while global attention processes the fourth fused feature map. Local attention can focus on small defect areas, enhancing the feature identification ability of the model in local areas. Especially when dealing with small or complex defects, it can effectively enhance the sensitivity of the model. Global attention focuses on the global information of the image, helping to improve the performance of the model when dealing with complex backgrounds or multiple interferences.
[0065] The feature map processed by the multi-head hybrid attention mechanism is input into the improved Transformer model for further processing. The model includes an encoder, a decoder, and a prediction unit. The input feature map is first processed by the encoder. In this process, each input feature map undergoes a non-linear transformation through a feed-forward neural network and is processed through residual connections and LayerNorm operations, outputting a fused feature vector.
[0066] E fused = σ(W f ·[E1, E2, E3, E4] + b f )
[0067] where E fused is the fused feature vector output by the encoder, and E1, E2, E3, and E4 are the feature vectors after the first fused feature map, the second fused feature map, the third fused feature map, and the fourth fused feature map undergo non-linear transformation through the feed-forward neural network and are processed through residual connections and LayerNorm operations respectively. [E1, E2, E3, E4] represents the concatenation operation of the four feature vectors. W f is a learnable fusion parameter matrix, b f is the bias term, and σ is the activation function;
[0068] This fused feature vector includes local and global feature information and can better reflect the detailed features and overall features in the input image. Then, through the multi-head self-attention mechanism and cross-attention mechanism of the decoder, the accuracy of feature representation is further improved. The cross-attention mechanism effectively enhances the recognition ability of various defect targets through interaction with 50 preset learnable object query vectors, ensuring the accuracy of defect classification.
[0069] During the model training process, based on the Hungarian bipartite matching algorithm, the prediction boxes and the ground truth boxes are matched, the matching cost is calculated, and the optimal matching relationship is selected. According to the matching relationship, the loss function is calculated, and the parameters of the multi-scale feature extraction module and the improved Transformer model are optimized through the gradient descent method, thereby realizing the training of the model. This step can further improve the accuracy of the model in defect localization by accurately matching the prediction boxes and the ground truth boxes, ensuring the accuracy of the detection results. The formula is:
[0070]
[0071] where L is the loss function, N is the number of generated prediction boxes, M is the number of ground truth boxes of the target in the input training image, cost(i, j) is the matching cost, used to measure the difference between the i-th prediction box and the j-th ground truth box, and x ij is a binary variable indicating that the i-th prediction box is matched with the j-th ground truth box, and xij = 1 indicates that the i-th predicted bounding box matches the j-th ground truth bounding box, x ij = 0 indicates that the i-th predicted bounding box does not match the j-th ground truth bounding box, is the i-th predicted bounding box b i and the j-th ground truth bounding box The intersection over union between them, and θ is a preset parameter.
[0072] Through the above steps, the present invention combines a hybrid attention mechanism, multi-scale feature extraction, and an improved Transformer model to achieve precise detection of insulator defects in complex backgrounds. This method has strong robustness and real-time performance, can effectively improve the detection accuracy of small defects, and can maintain stable performance in complex environments. At the same time, the present invention can significantly improve the computational efficiency and meet the requirements of the smart grid system for real-time detection and high-precision analysis.
[0073] Example 2:
[0074] The present invention realizes precise detection of insulator defects in transmission lines through improved deep learning technology. The data comes from the insulator detection records of a certain power system, with a total of n samples, and the data sampling frequency is once per hour. The dataset includes images of normal, damaged, and flashover insulators, which are used to train and test the proposed multi-scale feature fusion Transformer (MSFFT) model. The overall structure of the MSFFT insulator defect detection model is as Figure 1 shown. The specific implementation plan is as follows:
[0075] Construction of the multi-scale feature extraction module:
[0076] In order to establish an effective insulator defect detection model, it is first necessary to comprehensively extract and fuse multi-scale features in the image. The overall structure of the multi-scale feature extraction module is as Figure 2 shown.
[0077] Step1: Data collection and preprocessing: Collect insulator image data from transmission lines, covering normal and damaged states under different environmental conditions. Clean, crop, and normalize the image data to eliminate noise and unify the data format.
[0078] Step 2: Construction of multi-scale feature extraction model: The preprocessed insulator image is input into the multi-scale feature extraction module based on the pre-trained ResNet-50 architecture. This module extracts feature maps layer by layer through the residual neural network, generating multi-scale features with different resolutions. Low-level texture features of H / 4×W / 4 with 256 channels are extracted at the Conv2_x (R2) stage; local spatial features of H / 8×W / 8 with 512 channels are extracted at the Conv3_x (R3) stage; intermediate semantic features of H / 16×W / 16 with 1024 channels are extracted at the Conv4_x (R4) stage; and global semantic features of H / 32×W / 32 with 2048 channels are extracted at the Conv5_x (R5) stage. These feature maps cover multi-level information from low to high, providing a basis for feature fusion.
[0079] Step 3: Feature fusion: The extracted multi-scale feature maps are input into the top-down inverse residual neural network, and through the upsampling operation, the high-resolution semantic information is gradually transmitted and fused to the low-resolution feature maps. First, the feature map of R5 is resized through the upsampling operation and horizontally fused with the feature map of R4 to generate a feature map of H / 16×W / 16, and the number of channels is reduced to 256. Subsequently, the fused feature of R4 is processed through the same upsampling and horizontally fused with the feature map of R3 to generate a feature map of H / 8×W / 8 with 256 channels. Finally, the fused feature of R3 is fused with the feature map of R2 to generate a feature map of H / 4×W / 4 with 256 channels. The last layer of feature F4 is the feature map of R5 (Conv5) further optimized through the convolution operation (Conv2d) and batch normalization processing (Groupnorm). Through this feature fusion process, the high-resolution semantic information and the low-resolution spatial information are fully integrated.
[0080] The fused feature map is further processed through the convolution operation to generate multi-scale feature maps F1, F2, F3, and F4. The size of F1 is H / 16×W / 16, which is used to capture global semantic information; the size of F2 is H / 32×W / 32, which is used for fine-grained pattern extraction; the sizes of F3 and F4 are both H / 64×W / 64, which are suitable for detecting large-range defects. The multi-scale feature maps integrate local and global information, providing strong feature support for defect classification and localization.
[0081] After completing feature extraction and fusion, the present invention introduces a multi-head hybrid attention mechanism to enhance the model's attention to detail features and its ability to understand the overall context. This mechanism can effectively process multi-scale defect features in feature maps at different levels by combining local window attention and global attention. The structure of the multi-head hybrid attention mechanism is as Figure 3 shown.
[0082] In each layer of the model, local window attention and global attention are calculated in a fused manner to adapt to the characteristics of feature maps at different levels. In low-level feature maps (such as F1, F2, and F3), the model mainly adopts the local window attention mechanism to concentrate on processing the detailed information in the image, such as small defects on the surface of insulators. By restricting the attention calculation within a fixed-size window (such as k×k), local window attention can accurately capture the correlation relationships of spatially local feature points. In high-level feature maps (such as F4), the model applies the global attention mechanism to capture the semantic context of the entire image to improve the detection ability for large-scale or complex defects. Through this fusion of local and global attention, the model achieves a good balance between detail recognition and overall understanding.
[0083] To fully exploit the potential of the multi-head hybrid attention mechanism, the present invention optimizes the attention weight parameters so that the model can efficiently process multi-scale defect features. In the local attention mechanism, the weight ratios of different windows are adjusted to ensure that the detailed information of small defects is not overlooked; in the global attention mechanism, appropriate weights are assigned to the global features of different layers to ensure the model's ability to locate large-scale defects. At the same time, the optimization of attention weights can adapt to the characteristics of different input images, enabling the model to remain robust in complex scenarios.
[0084] The improved Transformer model in the present invention consists of multiple layers of encoders and decoders, which are mainly used for in-depth analysis and feature fusion of complex image data to achieve the tasks of defect classification and localization. The input to the model is four feature maps with different resolutions from the multi-scale feature extraction module, namely F1 (H / 16×W / 16), F2 (H / 32×W / 32), F3 (H / 64×W / 64), and F4 (H / 128×W / 128). These feature maps are first flattened to convert each feature map into a one-dimensional sequence (tokens). Each feature point is represented as an input token and combined with sine position encoding to introduce spatial position information, enabling the model to capture the spatial relationships between feature points in the image.
[0085] In the encoder stage, the input feature vectors are processed by the multi-head self-attention mechanism. Each feature point interacts with other feature points globally, capturing the long-range dependencies and spatial relationships in different regions of the image, and enhancing the model's recognition ability in complex backgrounds. The feed-forward neural network module in each layer further performs non-linear mapping on these features and is optimized through residual connections and LayerNorm operations to ensure the stable transmission of feature information in the multi-layer structure, while improving the efficiency and stability of training. Through this multi-layer stacking process, the encoder effectively fuses semantic information at different scales, enabling each feature point to contain richer context information and providing accurate input for the decoder.
[0086] In the decoder stage, the features output by the encoder are further enhanced by the multi-head self-attention mechanism to improve the information interaction between features and enhance the correlation of feature points inside the decoder. The decoder can not only handle the internal feature interaction, but also fuse the global context information transmitted by the encoder with the input features of the decoder through the cross-attention mechanism, enhancing the ability to understand image details. The cross-attention mechanism helps the decoder accurately classify and locate defects at each position. Through this process, the decoder can generate the final defect class probability distribution and bounding box coordinates, providing accurate output for object detection and defect recognition.
[0087] In the output stage of the decoder, the present invention designs a prediction unit for converting the hidden feature representation into the target category and bounding box coordinates. Through the implementation and optimization of this unit, the model can achieve high-precision defect detection in multi-object scenarios and improve the detection performance for small defects.
[0088] The prediction unit uses 50 learnable object queries, which are dynamically adjusted according to the characteristics of the dataset during model training. The object queries are used to adapt to different datasets and detection scenarios, and capture the target-related information in the feature map through the cross-attention mechanism between the decoder and the encoder to ensure the accurate parsing of targets in complex scenarios.
[0089] To solve the problem of target assignment in multi-object scenarios, this embodiment introduces the Hungarian bipartite matching algorithm for optimization during model training. The Hungarian algorithm is a classic optimization method used to find the optimal assignment scheme in matching tasks. In object detection tasks, the Hungarian algorithm can accurately correspond each predicted box to the real target by calculating the matching cost between the predicted box and the real box, thereby improving the detection accuracy.
[0090] Specifically, in the object detection framework of the present invention, the Hungarian algorithm is applied to the decoder part. The decoder generates multiple candidate prediction boxes, and these prediction boxes need to be matched with the ground truth bounding boxes. To perform effective matching, first, the intersection over union (IoU) between each prediction box and the ground truth box is calculated to obtain the matching cost for each pair of boxes. When the IoU value between the prediction box and the ground truth box is lower than the set threshold, the matching cost is set to 1, indicating that these two boxes do not match; when the IoU value is greater than or equal to the set threshold, the matching cost is set to 0, indicating that they match. Then, the Hungarian algorithm selects the optimal matching scheme by minimizing the total cost, that is, assigns a ground truth object to each prediction box.
[0091] In this way, the Hungarian algorithm can effectively avoid the situation of object loss and misassignment. Especially in the multi-object scenario, it can accurately match multiple prediction boxes with multiple ground truth boxes, thereby improving the detection accuracy of the model. Especially when dealing with complex backgrounds and small defects, the introduction of the Hungarian algorithm significantly reduces the probability of false detection and missed detection, and improves the overall performance of object detection.
[0092] In summary, the present invention realizes high-precision detection of insulator defects through the design and optimization of multi-scale feature fusion, multi-head hybrid attention mechanism and improved Transformer model, and has broad application prospects.
[0093] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.
[0094] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. An insulator defect detection method based on hybrid attention and multi-scale features, characterized in that Including the following steps: Obtain the insulator image of the transmission line to be detected, and preprocess the insulator image; Use a multi-scale feature extraction module to extract and fuse features from the preprocessed image. The multi-scale feature extraction module includes a bottom-up residual neural network and a top-down inverse residual neural network. The bottom-up residual neural network extracts feature maps of different levels through multiple residual blocks. The feature maps of different levels include low-level texture feature maps, local spatial feature maps, intermediate semantic feature maps, and global semantic feature maps. The top-down inverse residual neural network performs horizontal fusion of high-level feature maps and low-level feature maps through upsampling operations, and sequentially outputs the first fusion feature map, the second fusion feature map, the third fusion feature map, and the fourth fusion feature map; Input the fused feature map into a multi-head hybrid attention mechanism for processing. The multi-head hybrid attention mechanism includes local attention and global attention. Local attention processes the first fusion feature map, the second fusion feature map, and the third fusion feature map, and global attention processes the fourth fusion feature map; Input the feature map processed by the multi-head hybrid attention mechanism into an improved Transformer model for processing, output the probability distribution of the defect categories and the bounding box coordinates of the insulator image, and realize the defect detection of the insulator according to the output probability distribution of the defect categories and the bounding box coordinates.
2. The insulator defect detection method based on hybrid attention and multi-scale features according to claim 1, wherein The bottom-up residual neural network is a multi-scale feature extraction module based on the pre-trained ResNet-50 architecture, including multiple residual blocks. Each residual block gradually extracts features from the input image through convolutional layers, including: Conv2 stage. In this stage, a convolutional kernel is used to perform a convolutional operation on the input image to extract low-level texture features. The convolutional kernel size of the convolutional operation is 3×3, the stride is 1, and a low-level texture feature map with a resolution of H / 4×W / 4 and 256 channels is generated; Conv3 stage. In this stage, local spatial features are extracted, and a local spatial feature map with a resolution of H / 8×W / 8 and 512 channels is generated; Conv4 stage. In this stage, intermediate semantic features are extracted, and an intermediate semantic feature map with a resolution of H / 16×W / 16 and 1024 channels is generated; Conv5 stage. In this stage, global semantic features are extracted, and a global semantic feature map with a resolution of H / 32×W / 32 and 2048 channels is generated.
3. The insulator defect detection method based on hybrid attention and multi-scale features according to claim 2, characterized in that, Each residual block gradually extracts features from the input image through convolutional layers. The feature extraction process of the residual block is described as: F i (X) = R i (F i-1 (X)) = σ(W i *F i-1 (X) + b i ) + F i-1 (X) Among them, F i (X) represents the feature map output by the residual block in the Convi stage in the Convi stage, and F i-1 (X) represents the feature map output by the residual block in the Convi-1 stage in the Convi-1 stage. R i represents the residual block in the Convi stage, σ is the activation function, and W i and b i are the weight and bias of the convolutional layer respectively. * represents the convolution operation. The value range of i is 2, 3, 4, and 5, corresponding to the feature extraction processes in the Conv2, Conv3, Conv4, and Conv5 stages of the bottom-up residual neural network respectively.
4. A method for detecting insulator defects based on hybrid attention and multi-scale features according to claim 1, characterized in that The top-down inverse residual neural network performs horizontal fusion of high-level feature maps and low-level feature maps through upsampling operations, specifically including: Generate the fourth fusion feature map by performing a convolutional operation and batch normalization on the global semantic feature map; Adjust the size of the global semantic feature map through an upsampling operation, and horizontally fuse it with the intermediate semantic feature map to generate a third fusion feature map with a resolution of H / 16×W / 16 and 256 channels; The intermediate semantic feature map is resized through an upsampling operation and horizontally fused with the local spatial feature map to generate a second fused feature map of H / 8×W / 8 with 256 channels. The local spatial feature map is resized through an upsampling operation and horizontally fused with the low-level texture feature map to generate a first fused feature map of H / 4×W / 4 with 256 channels.
5. The insulator defect detection method based on hybrid attention and multi-scale features according to claim 4, wherein The top-down inverse residual neural network horizontally fuses high-level feature maps and low-level feature maps through an upsampling operation, which is described as: G i (F i (X)) = ReLU(W up * U i+1 (F i+1 (X)) + b up + F i (X)) Among them, G i (F i (X)) is the feature map after fusion in the i-th layer, representing the feature output after upsampling, fusion, and activation. F i (X) represents the feature map output by the residual block in the Convi stage in the Convi stage. W up *U i+1 represents the upsampling operation, and b up is the bias term in the upsampling operation. ReLU represents the ReLU activation function.
6. The insulator defect detection method based on hybrid attention and multi-scale features according to claim 1, wherein, The local attention processes the first fused feature map, the second fused feature map, and the third fused feature map, specifically including: Set a local window of a fixed size on each input feature map, and determine the position and range of each window according to the feature point coordinates in the input feature map; Utilize the self-attention mechanism to calculate the correlation between pixels within each window to obtain the attention score of each pixel and obtain the attention score matrix within each window; Weight the pixels within each local region in each feature map by the obtained attention score matrix to obtain the feature map after weighted processing; Merge the feature maps after weighted processing of all local windows to obtain the feature map integrating local attention features, and obtain the first fused feature map, the second fused feature map, and the third fused feature map after local attention processing.
7. A method for detecting insulator defects based on hybrid attention and multi-scale features according to claim 1, characterized in that The global attention processes the fourth fused feature map, specifically including: Map the number of channels of the fourth fused feature map to the corresponding dimension through convolution operation or linear transformation to obtain the query matrix Q, the key matrix K, and the value matrix V; Perform a dot product operation on the query matrix Q and the key matrix K to calculate the similarity between each pair of positions and obtain the attention score matrix; Use the Softmax function to normalize the attention score matrix, and perform weighted summation of the normalized attention weight matrix and the value matrix V to obtain the weighted feature representation of each position, and output the fourth fused feature map after global attention processing.
8. A method for detecting insulator defects based on hybrid attention and multi-scale features according to claim 1, characterized in that, The improved Transformer model includes an encoder, a decoder, and a prediction unit, and the prediction unit includes 50 preset learnable object query vectors.
9. The insulator defect detection method based on hybrid attention and multi-scale features according to claim 1, characterized in that, Inputting the feature map processed by the multi-head hybrid attention mechanism into the improved Transformer model for processing specifically includes: Input the first fused feature map, the second fused feature map, and the third fused feature map processed by the multi-head hybrid attention mechanism and the fourth fused feature map into the encoder, perform non-linear transformation through a feed-forward neural network, and process through residual connection and LayerNorm operation, and output the fused feature vector, and the formula is: E fused = σ(W f ·[E1, E2, E3, E4] + b f ) Among them, E fused is the fused feature vector output by the encoder. E1, E2, E3, and E4 are the feature vectors obtained by performing non-linear transformation on the first fused feature map, the second fused feature map, the third fused feature map, and the fourth fused feature map through a feed-forward neural network and processing them through residual connection and LayerNorm operation. [E1, E2, E3, E4] represents the concatenation operation of the four feature vectors. W f is a learnable fusion parameter matrix, b f is the bias term, and σ is the activation function; Input the fused feature vector E output by the encoder fused into the multi-head self-attention mechanism of the decoder for processing. Input the feature vector after being processed by the multi-head self-attention mechanism into the cross-attention mechanism for processing. The cross-attention mechanism interacts with the key and value vectors of the feature vector after being processed by the multi-head self-attention mechanism through 50 preset learnable object query vectors, calculates a new feature representation through the cross-attention mechanism, inputs the feature vector after being processed by the cross-attention mechanism into the feed-forward neural network for non-linear transformation and activation function processing, and outputs the probability distribution of the defect category and the bounding box coordinates of the insulator image through residual connection and layer normalization operations.
10. A method for detecting insulator defects based on hybrid attention and multi-scale features according to claim 1, characterized in that, The multi-scale feature extraction module and the improved Transformer model are trained based on the Hungarian bipartite matching algorithm by combining the intersection over union between the predicted bounding boxes and the ground truth bounding boxes, specifically including: During the training process, each time a training image is input, the improved Transformer model generates multiple prediction boxes. The Hungarian bipartite matching algorithm is used to calculate the matching cost between each prediction box and all ground truth boxes, select the optimal matching relationship, calculate the loss function based on the optimal matching relationship, and adjust the parameters of the multi-scale feature extraction module and the improved Transformer model through the gradient descent method by calculating the loss function, so as to realize the training of the multi-scale feature extraction module and the improved Transformer model. The formula is as follows: Among them, L is the loss function, N is the number of generated predicted bounding boxes, M is the number of ground-truth bounding boxes in the input training images, cost(i, j) is the matching cost, which is used to measure the difference between the i-th predicted bounding box and the j-th ground-truth bounding box, and x ij is a binary variable indicating whether the i-th predicted bounding box matches the j-th ground-truth bounding box. x ij = 1 indicates that the i-th predicted bounding box matches the j-th ground-truth bounding box, and x ij = 0 indicates that the i-th predicted bounding box does not match the j-th ground-truth bounding box. is the intersection over union between the i-th predicted bounding box b i and the j-th ground-truth bounding box , and θ is a preset parameter.
Citation Information
Cited By
Data enhancement and yield analysis method and system for multi-defect pattern wafer graph
CN120782776A
Data enhancement and yield analysis method and system for multi-defect pattern wafer map
CN120782776B
An open-world oriented background-defect joint modeling defect detection method and system
CN122530218A