Insulator defect detection method and system based on multi-scale receptive field attention residual network
By using a multi-scale receptive field attention residual network in insulator defect detection, combined with a multi-scale receptive field and coordinate attention module, the problem of low detection accuracy is solved, and higher detection accuracy and calculation efficiency are achieved.
Patent Information
- Application Number
- CN202510155672.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-12
AI Technical Summary
The existing insulator defect detection methods face the problem of low detection accuracy caused by complex background and small defect area.
A multi-scale receptive field attention residual network based on the ResNet50 network framework is adopted, and a multi-scale receptive field and coordinate attention module are combined to improve the model's ability to extract target features and enhance key features.
It improves the accuracy and accuracy of insulator defect detection, can more effectively extract the characteristics of the target to be inspected, and has low computing overhead, ensuring the real-time and efficientness of the overall network model.
Smart Images

Figure CN119991642A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power transmission line detection, and in particular to an insulator defect detection method and system based on a multi-scale receptive field attention residual network. Background Art
[0002] The features extracted by the backbone network are crucial to the effectiveness of the target detection algorithm and determine the final performance of the algorithm. The backbone network is responsible for extracting features from the image to be detected, providing a basis for the algorithm to complete the detection task. Deeper network structures usually have stronger feature expression capabilities, but increasing the number of network layers will cause performance degradation. The Residual Network (ResNet) avoids the performance degradation problem by repeatedly stacking the residual structure, while having a powerful feature extraction capability.
[0003] Compared with other inspection tasks, insulator defect detection faces greater challenges. On the one hand, there are many objects unrelated to the target in the complex background where the insulator is located, and these objects may block the target to be inspected. On the other hand, the insulator defect occupies a very small proportion of the image area, making it difficult to detect. These challenges increase the complexity of the insulator defect detection task, which in turn affects the algorithm's detection accuracy of insulator defects. Summary of the invention
[0004] (I) Purpose of the invention
[0005] The purpose of the present invention is to provide an insulator defect detection method and system based on a multi-scale receptive field attention residual network. By combining multiple scale receptive fields and coordinate attention in a multi-scale receptive field attention module, the model's ability to extract target features can be effectively improved, and key features can be adaptively enhanced. The coordinate attention module combines position information for attention weighting, has low computational overhead, and can effectively improve target positioning accuracy. It is used to solve the problem that the existing insulator defect detection method and system have low detection accuracy when facing complex detection tasks.
[0006] (II) Technical solution
[0007] To solve the above problems, the first aspect of the present invention provides an insulator defect detection method based on a multi-scale receptive field attention residual network, comprising:
[0008] Based on the ResNet50 network framework, the multi-scale receptive field attention module and the coordinate attention module, an overall network model based on the multi-scale receptive field attention residual network is constructed;
[0009] Using the overall network model to extract features from an input insulator image to obtain a feature map;
[0010] Defect detection is performed on the characteristic diagram to obtain a detection result of the insulator defect.
[0011] Furthermore, based on the ResNet50 network framework, the multi-scale receptive field attention module and the coordinate attention module, an overall network model based on the multi-scale receptive field attention residual network is constructed, including:
[0012] Use the ResNet50 network framework as the basic network architecture;
[0013] Embed the multi-scale receptive field attention module into the ResNet50 network framework;
[0014] The coordinate attention module is embedded into the multi-scale receptive field attention module.
[0015] Furthermore, the step of extracting features from the input insulator image using the overall network model to obtain a feature map includes:
[0016] The original insulator image is input into the overall network model, and the first stage feature extraction is performed after 7×7 2D convolution and 3×3 maximum pooling to obtain the output feature map of the first stage;
[0017] The output feature map of the first stage is input into the first residual structure consisting of three 1×1, 3×3 and 1×1 2D convolutions, and the second stage feature extraction is performed in combination with the first multi-scale receptive field attention module to obtain the output feature map of the second stage;
[0018] The output feature map of the second stage is input into the second residual structure consisting of four 1×1, 3×3 and 1×1 2D convolutions, and the third stage feature extraction is performed in combination with the second multi-scale receptive field attention module to obtain the output feature map of the third stage;
[0019] Input the output feature map of the third stage into 6 third residual structures composed of 1×1, 3×3 and 1×1 2D convolutions, and perform the third stage feature extraction in combination with the third multi-scale receptive field attention module to obtain the output feature map of the fourth stage;
[0020] The output feature map of the fourth stage is input into a fourth residual structure consisting of three 2D convolutions of 1×1, 3×3 and 1×1, and the fourth stage feature extraction is performed in combination with the fourth multi-scale receptive field attention module to obtain the output feature map of the fifth stage.
[0021] Furthermore, the output feature map of the first stage is input into three residual structures consisting of 1×1, 3×3 and 1×1 2D convolutions, and the second stage feature extraction is performed in combination with the first multi-scale receptive field attention module to obtain the output feature map of the second stage, including:
[0022] Input the output feature map of the first stage into the first residual structure composed of three 1×1, 3×3 and 1×1 2D convolutions to extract features and obtain the feature map output by the first residual structure;
[0023] Input the feature maps output by the first residual structure into four branches of multi-scale receptive field attention modules with different sizes respectively;
[0024] Among them, the receptive field size of the first branch is 3×3, including 1 1×1 ordinary 2D convolution and 1 3×3 dilated convolution with a dilation rate of 1;
[0025] The receptive field size of the second branch is 7×7, including 1 1×1 ordinary 2D convolution and 1 3×3 dilated convolution with a dilation rate of 3;
[0026] The receptive field size of the third branch is 11×11, including 1 1×1 ordinary 2D convolution and 1 3×3 dilated convolution with a dilation rate of 5;
[0027] The fourth branch is the residual branch, whose receptive field is 1×1, including 1 1×1 ordinary 2D convolution;
[0028] The output feature maps of the first branch, the second branch, and the third branch are connected in the channel dimension, and then the number of channels is adjusted through a 1×1 ordinary 2D convolution to obtain the first stage fusion feature map;
[0029] Input the first-stage fusion feature map into the first coordinate attention module for processing to obtain an output feature map of the first coordinate attention module;
[0030] After the feature map output by the fourth branch is fused with the output feature map of the first coordinate attention module, it is nonlinearly processed by the Relu activation function to obtain the output feature map of the second stage.
[0031] Furthermore, the step of inputting the first-stage fusion feature map into the first coordinate attention module for processing to obtain the output feature map of the first coordinate attention module includes:
[0032] Input the first stage fusion feature map into the first coordinate attention module;
[0033] Pooling the first-stage fused feature map in the vertical direction and the horizontal direction respectively to obtain the first-stage fused feature map after pooling;
[0034] Connecting the pooled first-stage fusion feature maps in the spatial dimension to obtain a first-stage connection feature map;
[0035] The first-stage connection feature map is passed through a 1×1 2D convolution to adjust the number of channels to obtain the first-stage nonlinear processing feature map;
[0036] The first-stage compressed feature map is batch normalized and nonlinearly processed using the Relu activation function to obtain the first-stage connected feature map;
[0037] Split the first-stage connected feature map along the spatial dimension;
[0038] After the split first-stage connection feature maps are respectively restored through a 1×1 2D convolution to restore the number of channels, they are processed by the Sigmoid function to obtain two first-stage attention weights that carry spatial position information;
[0039] The first-stage fusion feature map is weighted with the two first-stage attention weights to obtain the output feature map of the first-stage coordinate attention.
[0040] Preferably, the calculation formula of the equivalent convolution kernel of the dilated convolution is:
[0041] k`=k+(k-1)×(r-1)
[0042] Among them, k represents the size of the convolution kernel of the standard convolution, and r represents the dilation factor.
[0043] Preferably, the calculation formula for the sizes of the four branches of the multi-scale receptive field attention module with different sizes is:
[0044]
[0045] Among them, RF(i) is the receptive field size of the i-th layer, k i is the size of the convolution kernel of the i-th layer; S j is the step size of the j-th convolution kernel; RF(0)=1 is the initial value of the receptive field.
[0046] Furthermore, the first-stage fusion feature map is pooled in the vertical direction and the horizontal direction respectively to obtain the pooled first-stage fusion feature map, including:
[0047] The first stage fusion feature map after pooling includes the first stage fusion feature map after vertical pooling and the first stage fusion feature map after horizontal pooling;
[0048]
[0049] Among them, H is the size of the first stage fusion feature map in the vertical direction; W is the size of the first stage fusion feature map in the horizontal direction; It is the first stage fusion feature map after vertical pooling of the nth channel; is the first stage fusion feature map after horizontal pooling of the nth channel; h is the height of the nth channel; w is the width of the nth channel.
[0050] Furthermore, the number of output channels of the 7×7 2D convolution is 64, and the stride of the 7×7 convolution kernel and the 3×3 maximum pooling kernel are both 2;
[0051] In the first residual structure, the number of output channels of the 1×1 2D convolution is 64, the number of output channels of the 3×3 2D convolution is 64, and the number of output channels of the 1×1 2D convolution is 256;
[0052] In the second residual structure, the number of output channels of the 1×1 2D convolution is 128, the number of output channels of the 3×3 2D convolution is 128, and the number of output channels of the 1×1 2D convolution is 512;
[0053] The number of output channels of the 1×1 2D convolution in the third residual structure is 256, the number of output channels of the 3×3 2D convolution is 256, and the number of output channels of the 1×1 2D convolution is 1024;
[0054] In the fourth residual structure, the number of output channels of the 1×1 2D convolution is 512, the number of output channels of the 3×3 2D convolution is 512, and the number of output channels of the 1×1 2D convolution is 2048.
[0055] According to another aspect of the present invention, there is provided an insulator defect detection system based on a multi-scale receptive field attention residual network, comprising:
[0056] Model building module: Based on the ResNet50 network framework, multi-scale receptive field attention module and coordinate attention module, an overall network model based on a multi-scale receptive field attention residual network is constructed;
[0057] Feature extraction module: used to extract features from the input insulator image using the overall network model to obtain a feature map;
[0058] Defect detection module: used to perform defect detection on the characteristic graph to obtain the detection result of insulator defects.
[0059] (III) Beneficial effects
[0060] The above technical solution of the present invention has the following beneficial technical effects:
[0061] The present invention provides an insulator defect detection method based on a multi-scale receptive field attention residual network, comprising: constructing an overall network model based on a multi-scale receptive field attention residual network based on a ResNet50 network framework, a multi-scale receptive field attention module and a coordinate attention module; performing feature extraction on an input insulator image using the overall network model to obtain a feature map; performing defect detection on the feature map to obtain a detection result of an insulator defect.
[0062] The technical solution provided by the embodiment of the present invention gradually extracts feature maps at different levels and combines multi-scale receptive fields and coordinate attention, so that the feature maps finally outputted can better characterize the key targets in the image, thereby improving the precision and accuracy of insulator defect detection. It can cope with the challenges of complex situations in insulator defect detection tasks and can more effectively extract the features of the targets to be inspected.
[0063] It has low computational overhead. In the design of the coordinate attention module, through reasonable weighting of spatial position information, it can effectively improve the target positioning accuracy while ensuring the real-time and efficiency of the overall network model. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.
[0065] Figure 1 It is a flow chart of the insulator defect detection method based on multi-scale receptive field attention residual network of the present invention;
[0066] Figure 2 It is a feature extraction flow chart of the overall network model of the insulator defect detection method based on the multi-scale receptive field attention residual network of the present invention;
[0067] Figure 3 It is a structure and feature extraction flow chart of a multi-scale receptive field attention module of an insulator defect detection method based on a multi-scale receptive field attention residual network of the present invention;
[0068] Figure 4 It is a structure and feature extraction flow chart of the coordinate attention module of the insulator defect detection method based on the multi-scale receptive field attention residual network of the present invention. DETAILED DESCRIPTION
[0069] In order to make the purpose, technical scheme and advantages of the present invention clearer, the present invention is further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings. It should be understood that these descriptions are only exemplary and are not intended to limit the scope of the present invention. In addition, in the following description, the description of well-known structures and technologies is omitted to avoid unnecessary confusion of the concept of the present invention.
[0070] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in one or more embodiments of the present invention should be understood by people with ordinary skills in the field to which the present invention belongs. The words "first", "second" and similar words used in one or more embodiments of the present invention do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect.
[0071] The technical solution of the present invention is described in detail below with reference to the accompanying drawings. An embodiment of the present invention provides an insulator defect detection method based on a multi-scale receptive field attention residual network. Figure 1 is a flow chart of the insulator defect detection method based on the multi-scale receptive field attention residual network. Figure 1 As shown, the method includes:
[0072] S1. Based on the ResNet50 network framework, the multi-scale receptive field attention module and the coordinate attention module, an overall network model based on the multi-scale receptive field attention residual network is constructed, including:
[0073] S11. Use the ResNet50 network framework as the basic network architecture;
[0074] S12, embed the multi-scale receptive field attention module into the ResNet50 network framework;
[0075] S13. Embed the coordinate attention module into the multi-scale receptive field attention module.
[0076] Combination Figure 1 and Figure 2 :
[0077] S2. Use the overall network model to extract features from the input insulator image to obtain a feature map, including:
[0078] S21, input the original insulator image into the overall network model, perform the first stage feature extraction after 7×7 2D convolution and 3×3 maximum pooling, and obtain the first stage output feature map;
[0079] As a preferred implementation mode,
[0080] Among them, the number of output channels of the 7×7 2D convolution is 64, and the stride of the 7×7 convolution kernel and the 3×3 maximum pooling kernel is 2.
[0081] S22, input the output feature map of the first stage into the first residual structure composed of three 1×1, 3×3 and 1×1 2D convolutions, and perform the second stage feature extraction in combination with the first multi-scale receptive field attention module to obtain the output feature map of the second stage, including:
[0082] Combination Figure 1-Figure 3 :
[0083] S221, input the output feature map of the first stage into three first residual structures composed of 1×1, 3×3 and 1×1 2D convolutions to perform feature extraction to obtain a feature map output by the first residual structure;
[0084] S222, inputting the feature map output by the first residual structure into four branches of multi-scale receptive field attention modules with different sizes respectively;
[0085] Among them, the receptive field size of the first branch is 3×3, including 1 1×1 ordinary 2D convolution and 1 3×3 dilated convolution with a dilation rate of 1;
[0086] The receptive field size of the second branch is 7×7, including 1 1×1 ordinary 2D convolution and 1 3×3 dilated convolution with a dilation rate of 3;
[0087] The receptive field size of the third branch is 11×11, including 1 1×1 ordinary 2D convolution and 1 3×3 dilated convolution with a dilation rate of 5;
[0088] The fourth branch is the residual branch, whose receptive field is 1×1, including 1 1×1 ordinary 2D convolution;
[0089] S223, connecting the output feature maps of the first branch, the second branch, and the third branch in the channel dimension, and then adjusting the number of channels through a 1×1 ordinary 2D convolution to obtain a first-stage fusion feature map;
[0090] S224, inputting the first stage fusion feature map into the first coordinate attention module for processing, and obtaining an output feature map of the first coordinate attention module, including:
[0091] Combination Figure 1-Figure 4 :
[0092] S2241, inputting the first stage fusion feature map into the first coordinate attention module;
[0093] S2242, pooling the first-stage fused feature map in the vertical direction and the horizontal direction respectively to obtain the first-stage fused feature map after pooling;
[0094] Preferably, the first stage fusion feature map after pooling includes the first stage fusion feature map after vertical pooling and the first stage fusion feature map after horizontal pooling;
[0095]
[0096] Among them, H is the size of the first stage fusion feature map in the vertical direction; W is the size of the first stage fusion feature map in the horizontal direction; It is the first stage fusion feature map after vertical pooling of the nth channel; is the first stage fusion feature map after horizontal pooling of the nth channel; h is the height of the nth channel; w is the width of the nth channel.
[0097] S2243, connecting the pooled first-stage fusion feature map in the spatial dimension to obtain a first-stage connection feature map;
[0098] S2244, passing a 1×1 2D convolution through the first-stage connection feature map to adjust the number of channels and obtain a first-stage nonlinear processing feature map;
[0099] S2245, performing batch normalization on the first-stage compressed feature map, performing nonlinear processing using a Relu activation function, and obtaining a first-stage connected feature map;
[0100] S2246, splitting the first stage connection feature map along the spatial dimension;
[0101] S2247, after the split first-stage connection feature map is respectively restored to the number of channels through a 1×1 2D convolution, and then processed by the Sigmoid function to obtain two first-stage attention weights carrying spatial position information;
[0102] S2248. Weight the first-stage fusion feature map and the two first-stage attention weights to obtain the output feature map of the first-stage coordinate attention.
[0103] S225. After fusing the feature map output by the fourth branch with the output feature map of the first coordinate attention module, nonlinear processing is performed through the Relu activation function to obtain the output feature map of the second stage.
[0104] As a preferred implementation mode,
[0105] Among them, the number of output channels of 1×1 2D convolution in the first residual structure is 64, the number of output channels of 3×3 2D convolution is 64, and the number of output channels of 1×1 2D convolution is 256.
[0106] S23, inputting the output feature map of the second stage into the second residual structure consisting of four 1×1, 3×3 and 1×1 2D convolutions, and performing the third stage feature extraction in combination with the second multi-scale receptive field attention module to obtain the output feature map of the third stage;
[0107] As a preferred implementation mode,
[0108] Among them, the number of output channels of 1×1 2D convolution in the second residual structure is 128, the number of output channels of 3×3 2D convolution is 128, and the number of output channels of 1×1 2D convolution is 512.
[0109] S24, inputting the output feature map of the third stage into the third residual structure composed of 6 1×1, 3×3 and 1×1 2D convolutions, combining the third multi-scale receptive field attention module to perform the third stage feature extraction, and obtaining the output feature map of the fourth stage;
[0110] As a preferred implementation mode,
[0111] Among them, the number of output channels of 1×1 2D convolution in the third residual structure is 256, the number of output channels of 3×3 2D convolution is 256, and the number of output channels of 1×1 2D convolution is 1024.
[0112] S25, inputting the output feature map of the fourth stage into three fourth residual structures consisting of 1×1, 3×3 and 1×1 2D convolutions, and performing fourth stage feature extraction in combination with the fourth multi-scale receptive field attention module to obtain the output feature map of the fifth stage;
[0113] As a preferred implementation mode,
[0114] Among them, the number of output channels of 1×1 2D convolution in the fourth residual structure is 512, the number of output channels of 3×3 2D convolution is 512, and the number of output channels of 1×1 2D convolution is 2048.
[0115] Among them, the operation steps of S23, S24 and S25 are similar and will not be repeated here.
[0116] Preferably, the calculation formula of the equivalent convolution kernel of the dilated convolution is:
[0117] k`=k+(k-1)×(r-1)
[0118] Among them, k represents the size of the convolution kernel of the standard convolution, and r represents the dilation factor.
[0119] Preferably, the calculation formula for the sizes of the branches of the four multi-scale receptive field attention modules with different sizes is:
[0120]
[0121] Among them, RF(i) is the receptive field size of the i-th layer, k i is the size of the convolution kernel of the i-th layer; S j is the step size of the j-th convolution kernel; RF(0)=1 is the initial value of the receptive field.
[0122] In one embodiment,
[0123] Combination Figure 1-Figure 4 ;
[0124] S222, input the feature map output by the first residual structure with a size of C×H×W into four branches of multi-scale receptive field attention modules with different sizes respectively;
[0125] Among them, in the first branch, the second branch and the third branch, 1×1 ordinary 2D convolution is used to process the feature map, the number of channels is reduced to C / 4, and the size of the output feature map after convolution is C / 4×H×W.
[0126] S223, connecting the output feature maps of the first branch, the second branch, and the third branch in the channel dimension, and then adjusting the number of channels through a 1×1 ordinary 2D convolution to obtain a first-stage fusion feature map with a size of C×H×W;
[0127] Among them, the 3×3 dilated convolution operations in the first branch, the second branch, and the third branch do not change the spatial size of the feature map, which remains C / 4×H×W.
[0128] The fourth branch is the residual branch and does not change the size of the input feature map.
[0129] S224, inputting the first stage fusion feature map of size C×H×W into the first coordinate attention module for processing, and obtaining the output feature map of the first coordinate attention module, including:
[0130] S2241, input the first stage fusion feature map with a size of C×H×W into the first coordinate attention module;
[0131] S2242, pooling the first stage fusion feature map of size C×H×W in the vertical direction and the horizontal direction respectively to obtain the first stage fusion feature map after pooling;
[0132] Combination Figure 4 , where the pooling kernel has a vertical size of H×1 and a horizontal size of 1×W. Through vertical pooling, the feature map size is compressed to C×1×W. Through horizontal pooling, the feature map size is compressed to C×H×1.
[0133] S2243, connecting the pooled first-stage fusion feature map in the spatial dimension to obtain a first-stage connection feature map;
[0134] S2244, passing a 1×1 2D convolution through the first-stage connection feature map to adjust the number of channels and obtain a first-stage nonlinear processing feature map;
[0135] Among them, the 1×1 2D convolution compresses the number of channels of the first-stage connection feature map to C / 8;
[0136] S2245, performing batch normalization on the first-stage compressed feature map, performing nonlinear processing using a Relu activation function, and obtaining a first-stage connected feature map;
[0137] S2246, splitting the first stage connection feature map along the spatial dimension;
[0138] S2247, after the split first-stage connection feature map is respectively restored to the number of channels through a 1×1 2D convolution, and then processed by the Sigmoid function to obtain two first-stage attention weights carrying spatial position information;
[0139] Among them, the 1×1 2D convolution restores the number of channels of the vertical and horizontal feature maps to C respectively.
[0140] In another embodiment of the present invention, there is also provided an insulator defect detection system based on a multi-scale receptive field attention residual network, comprising:
[0141] Model building module: Based on the ResNet50 network framework, multi-scale receptive field attention module and coordinate attention module, an overall network model based on a multi-scale receptive field attention residual network is constructed;
[0142] Feature extraction module: used to extract features from the input insulator image using the overall network model to obtain a feature map;
[0143] Defect detection module: used to perform defect detection on the characteristic graph and obtain the detection results of insulator defects.
[0144] In summary, an embodiment of the present invention involves an insulator defect detection method based on a multi-scale receptive field attention residual network, comprising: constructing an overall network model based on a multi-scale receptive field attention residual network based on a ResNet50 network framework, a multi-scale receptive field attention module and a coordinate attention module; performing feature extraction on an input insulator image using the overall network model to obtain a feature map; performing defect detection on the feature map to obtain a detection result of an insulator defect.
[0145] The technical solution provided by the embodiment of the present invention improves the ResNet50 network architecture and adds a multi-scale receptive field attention module at the end of each stage except the first stage to improve the feature extraction capability of each stage. At the same time, the coordinate attention module is embedded in the multi-scale receptive field attention module to enhance the focus on important features. Through multi-stage feature extraction, the insulator image is processed and features at different levels are output to achieve accurate detection.
[0146] The technical solution provided by the embodiment of the present invention gradually extracts feature maps at different levels and combines multi-scale receptive fields and coordinate attention, so that the feature maps finally outputted can better characterize the key targets in the image, thereby improving the precision and accuracy of insulator defect detection. It can cope with the challenges of complex situations in insulator defect detection tasks and can more effectively extract the features of the targets to be inspected.
[0147] It has low computational overhead. In the design of the coordinate attention module, through reasonable weighting of spatial position information, it can effectively improve the target positioning accuracy while ensuring the real-time and efficiency of the overall network model.
[0148] It should be understood that the above specific embodiments of the present invention are only used to illustrate or explain the principles of the present invention, and do not constitute a limitation of the present invention. Therefore, any modifications, equivalent substitutions, improvements, etc. made without departing from the spirit and scope of the present invention should be included in the protection scope of the present invention. In addition, the appended claims of the present invention are intended to cover all changes and modifications that fall within the scope and boundaries of the appended claims, or the equivalent forms of such scope and boundaries.
Claims
1. An insulator defect detection method based on a multi-scale receptive field attention residual network, characterized in that: The method comprises: Based on the ResNet50 network framework, the multi-scale receptive field attention module and the coordinate attention module, an overall network model based on the multi-scale receptive field attention residual network is constructed; Using the overall network model to extract features from the input insulator image to obtain a feature map; Defect detection is performed on the characteristic diagram to obtain a detection result of the insulator defect.
2. The method according to claim 1, characterized in that The overall network model based on the multi-scale receptive field attention residual network is constructed based on the ResNet50 network framework, the multi-scale receptive field attention module and the coordinate attention module, including: Use the ResNet50 network framework as the basic network architecture; Embed the multi-scale receptive field attention module into the ResNet50 network framework; The coordinate attention module is embedded into the multi-scale receptive field attention module.
3. The method according to claim 2, characterized in that The step of extracting features from the input insulator image using the overall network model to obtain a feature map includes: The original insulator image is input into the overall network model, and the first stage feature extraction is performed after 7×7 2D convolution and 3×3 maximum pooling to obtain the output feature map of the first stage; The output feature map of the first stage is input into the first residual structure consisting of three 1×1, 3×3 and 1×1 2D convolutions, and the second stage feature extraction is performed in combination with the first multi-scale receptive field attention module to obtain the output feature map of the second stage; The output feature map of the second stage is input into the second residual structure consisting of four 1×1, 3×3 and 1×1 2D convolutions, and the third stage feature extraction is performed in combination with the second multi-scale receptive field attention module to obtain the output feature map of the third stage; Input the output feature map of the third stage into 6 third residual structures composed of 1×1, 3×3 and 1×1 2D convolutions, and perform the third stage feature extraction in combination with the third multi-scale receptive field attention module to obtain the output feature map of the fourth stage; The output feature map of the fourth stage is input into a fourth residual structure consisting of three 2D convolutions of 1×1, 3×3 and 1×1, and the fourth stage feature extraction is performed in combination with the fourth multi-scale receptive field attention module to obtain the output feature map of the fifth stage.
4. The method according to claim 3, characterized in that The output feature map of the first stage is input into three residual structures consisting of 1×1, 3×3 and 1×1 2D convolutions, and the second stage feature extraction is performed in combination with the first multi-scale receptive field attention module to obtain the output feature map of the second stage, including: Input the output feature map of the first stage into the first residual structure composed of three 1×1, 3×3 and 1×1 2D convolutions to extract features and obtain the feature map output by the first residual structure; Input the feature maps output by the first residual structure into four branches of multi-scale receptive field attention modules with different sizes respectively; Among them, the receptive field size of the first branch is 3×3, including 1 1×1 ordinary 2D convolution and 1 3×3 dilated convolution with a dilation rate of 1; The receptive field size of the second branch is 7×7, including 1 1×1 ordinary 2D convolution and 1 3×3 dilated convolution with a dilation rate of 3; The receptive field size of the third branch is 11×11, including 1 1×1 ordinary 2D convolution and 1 3×3 dilated convolution with a dilation rate of 5; The fourth branch is the residual branch, whose receptive field is 1×1, including 1 1×1 ordinary 2D convolution; The output feature maps of the first branch, the second branch, and the third branch are connected in the channel dimension, and then the number of channels is adjusted through a 1×1 ordinary 2D convolution to obtain the first stage fusion feature map; Input the first-stage fusion feature map into the first coordinate attention module for processing to obtain an output feature map of the first coordinate attention module; After the feature map output by the fourth branch is fused with the output feature map of the first coordinate attention module, it is nonlinearly processed by the Relu activation function to obtain the output feature map of the second stage.
5. The method according to claim 4, characterized in that The step of inputting the first-stage fusion feature map into the first coordinate attention module for processing to obtain an output feature map of the first coordinate attention module includes: Input the first stage fusion feature map into the first coordinate attention module; Pooling the first-stage fused feature map in the vertical direction and the horizontal direction respectively to obtain the first-stage fused feature map after pooling; Connecting the pooled first-stage fusion feature maps in the spatial dimension to obtain a first-stage connection feature map; The first-stage connection feature map is passed through a 1×1 2D convolution to adjust the number of channels to obtain the first-stage nonlinear processing feature map; The first-stage compressed feature map is batch normalized and nonlinearly processed using the Relu activation function to obtain the first-stage connected feature map; Split the first-stage connected feature map along the spatial dimension; After the split first-stage connection feature maps are respectively restored through a 1×1 2D convolution to restore the number of channels, they are processed by the Sigmoid function to obtain two first-stage attention weights that carry spatial position information; The first-stage fusion feature map is weighted with the two first-stage attention weights to obtain the output feature map of the first-stage coordinate attention.
6. The method according to claim 3, characterized in that: The calculation formula of the equivalent convolution kernel of the dilated convolution is: k`=k+(k-1)×(r-1) Among them, k represents the size of the convolution kernel of the standard convolution, and r represents the dilation factor.
7. The method according to claim 6, characterized in that The calculation formula for the sizes of the four branches of the multi-scale receptive field attention module with different sizes is: Among them, RF(i) is the receptive field size of the i-th layer, k i is the size of the convolution kernel of the i-th layer; S j is the step size of the j-th convolution kernel; RF(0)=1 is the initial value of the receptive field.
8. The method according to claim 5, characterized in that The first stage fusion feature map is pooled in the vertical direction and the horizontal direction respectively to obtain the first stage fusion feature map after pooling, include: The first stage fusion feature map after pooling includes the first stage fusion feature map after vertical pooling and the first stage fusion feature map after horizontal pooling; Among them, H is the size of the first stage fusion feature map in the vertical direction; W is the size of the first stage fusion feature map in the horizontal direction; It is the first stage fusion feature map after vertical pooling of the nth channel; is the first stage fusion feature map after horizontal pooling of the nth channel; h is the height of the nth channel; w is the width of the nth channel.
9. The method according to claim 2, characterized in that: The number of output channels of the 7×7 2D convolution is 64, and the stride of the 7×7 convolution kernel and the 3×3 maximum pooling kernel are both 2; In the first residual structure, the number of output channels of the 1×1 2D convolution is 64, the number of output channels of the 3×3 2D convolution is 64, and the number of output channels of the 1×1 2D convolution is 256; In the second residual structure, the number of output channels of the 1×1 2D convolution is 128, the number of output channels of the 3×3 2D convolution is 128, and the number of output channels of the 1×1 2D convolution is 512; The number of output channels of the 1×1 2D convolution in the third residual structure is 256, the number of output channels of the 3×3 2D convolution is 256, and the number of output channels of the 1×1 2D convolution is 1024; In the fourth residual structure, the number of output channels of the 1×1 2D convolution is 512, the number of output channels of the 3×3 2D convolution is 512, and the number of output channels of the 1×1 2D convolution is 2048.
10. An insulator defect detection system based on a multi-scale receptive field attention residual network, characterized in that: The detection system comprises: Model building module: Based on the ResNet50 network framework, multi-scale receptive field attention module and coordinate attention module, an overall network model based on a multi-scale receptive field attention residual network is constructed; Feature extraction module: used to extract features from the input insulator image using the overall network model to obtain a feature map; Defect detection module: used to perform defect detection on the characteristic graph to obtain the detection result of insulator defects.
Citation Information
Patent Citations
Aerial photography small target detection method based on adaptive receptive field enhancement
CN115719450A
Insulator defect detection method and device, medium and program product
CN115984226A
SAR image ship target identification method based on attention mechanism
CN117218612A
Strip steel defect detection method based on redundant feature reuse and receptive field enhancement
CN117830283A
Radio frequency interference suppression method and device based on DeepLabV3 + model
CN118395274A
Cited By
Gearbox fault diagnosis method based on domain adaptation and deep residual network
CN121365331A
Insulator defect identification method based on multi-scale edge fusion and cross-space attention
CN121526952A