Meter defect detection method based on efficient local attention and large separable nuclear attention
By introducing separable convolution and efficient local attention modules into the YOLOv8 network model, combined with specific loss functions, the problems of low detection efficiency and insufficient data in substation meter defect detection are solved, and more efficient and accurate detection effects are achieved.
Patent Information
- Application Number
- CN202411808766.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-10
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art has problems in the detection of substation meter defects, such as low detection efficiency, easy to miss or missed detection, and due to insufficient defect data and insufficient model training samples, resulting in poor detection results.
The meter defect detection method based on efficient local attention and large separable nuclear attention is adopted. By improving the YOLOv8 network model, a separable convolutional structure and an efficient local attention module are added, and a border regression is detected by combining DFL Loss and CIoU Loss.
It improves the accuracy and efficiency of defect detection of substation meter, achieves more accurate feature fusion and target positioning, reduces the computational complexity and parameter quantity, and is suitable for substation environments with limited computing resources.
Smart Images

Figure CN119991545A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electric power visual target detection, and in particular to a meter defect detection method based on efficient local attention and large separable kernel attention. Background Art
[0002] In recent years, the State Grid has continuously promoted the development of new power systems, and the construction of substations has gradually expanded and upgraded. As the most important infrastructure in the power system, the stable operation of substations has an important impact on the entire power system and people's daily lives. However, affected by the natural environment and human activities, the meters in the substations are prone to defects such as damage, blur and dirt, which bring huge hidden dangers to the safe operation of the substations. In order to detect such defects in time, it is necessary to accurately locate and identify the defective meters. Therefore, the study of meter defect detection is of great significance to ensure the safe operation of the power system.
[0003] Some studies have used computer vision technology to identify defects in substation meters. For example, in the existing solution "Substation meter defect detection based on PCBAM-YOLOv5 network", a parallel hybrid attention module is added to YOLOv5 to enhance the meter defect features to improve detection accuracy; in "Substation meter defect detection algorithm based on improved YOLOv5", a coordinate attention mechanism is introduced to highlight the difference in defect features, and EDIOU is used to make the prediction bounding box regression more accurate and fast. However, there are still two difficulties in using deep learning target detection methods to detect meter defects in substation images:
[0004] 1. The background of substation meters is complex. There are many types of meter equipment, the defect characteristics are complex, and the model learning characteristics are difficult, resulting in low detection efficiency and easy omission or misdetection.
[0005] 2. Insufficient meter defect data. Due to the particularity of the substation scene, it is difficult to collect defect data on site. Most defect images are taken by inspection personnel during work, and the collection process is not unified, resulting in most defect data being unusable, insufficient model training samples, and poor model detection effect. Summary of the invention
[0006] In order to solve the problems existing in the above-mentioned prior art, the purpose of the present invention is to propose a meter defect detection method based on efficient local attention and large separable kernel attention, which provides important technical support for substation intelligent inspection and effectively improves the detection accuracy of substation meter defects.
[0007] To achieve the above object, the present invention provides the following solutions:
[0008] A meter defect detection method based on efficient local attention and large separable kernel attention, comprising:
[0009] Acquire defect images of substation meters to be inspected;
[0010] Inputting the defect image of the substation meter to be detected into the improved YOLOv8 network model to obtain a defect detection result; the improved YOLOv8 network model is obtained by training using a training set;
[0011] Improve the YOLOv8 network model: add a separable convolution structure after the SPPF layer in the backbone network of the YOLOv8 network model, and add an efficient local attention module after the neck network is skipped, that is, after the upsampling and the shallow features of the same size are spliced;
[0012] The separable convolution structure is used for high-level feature extraction;
[0013] The efficient local attention module is used to efficiently fuse features of different sizes;
[0014] The training set includes: substation meter defect images and defect labels.
[0015] Optionally, obtaining the training set includes:
[0016] The collected original meter defect images are cleaned according to the front orientation of the target object, obvious defect features, and the defect area occupies a preset area of the entire image, so as to obtain a picture set of the original substation meter defect images.
[0017] Optionally, obtaining the separable convolution structure includes:
[0018] The 2D weight convolution kernels of the depth convolution and depth extension convolution in the LKA block are divided into two cascaded 1D separable weight convolution kernels to obtain depth separable convolution and depth separable extension convolution, and the basic structure of the LKA block is maintained to obtain the separable convolution structure.
[0019] Optionally, the separable convolution structure includes:
[0020] The depthwise separable convolution is used to divide the number of channels of the input feature map into multiple groups if the size of the input feature map is the same as the size of the output feature map, use a convolution kernel to convolve each group, and splice multiple groups to obtain the channels of the output feature map to capture local spatial information;
[0021] The depthwise separable dilated convolution is used to introduce a dilation rate into the convolution layer, thereby defining the spacing between values when the convolution kernel in the convolution layer processes data, that is, the number of intervals between each point of the convolution kernel, and adding holes to expand the receptive field to capture global spatial information.
[0022] Optionally, the formula for expanding the receptive field is:
[0023] F i+1 =(2 i+2 -1)×(2 i+2 -1)
[0024] Among them, F i+1 is to expand the receptive field of the convolution, and i+1 is the expansion rate.
[0025] Optionally, the formula for high-level feature extraction by adding a separable convolution structure is:
[0026]
[0027] A C =W 1×1 *Z C
[0028]
[0029] Among them, * is convolution, is the Hadamard product, d is the expansion rate, represents the output of the deep convolution with kernel size (2d-1)×1 and 1×(2d-1), which captures the local spatial information and compensates for the grid effect of the subsequent deep expansion convolution. C The kernel size is and The depth of the dilated convolution, A C To obtain the attention map by performing kernel convolution on the output of the depth convolution, Note that Figure A C With the input feature map F C The output after Hadamard product, is a convolution kernel of size (2d-1)×1, For size The convolution kernel is , H is the height of the feature map, and W is the width of the feature map.
[0030] Optionally, the efficient local attention module includes:
[0031] A positioning information embedding unit is used to perform horizontal and vertical stripe pooling on the input feature map, calculate the average value using the element values in the pooling kernel, take the average value as the pooling output value, perform convolution on the output feature map, and expand it horizontally and vertically respectively, sum the pixels corresponding to the same position of the expanded feature map pixel by pixel to obtain a spliced feature map, process the spliced feature map by convolution and sigmoid, and then multiply it with the corresponding pixels of the input feature map to obtain a positioning information embedding result;
[0032] The positioning attention generation unit is used to enhance the positioning information based on a one-dimensional convolution sequence, process the enhanced position information through group normalization, and obtain position attention in the horizontal and vertical directions.
[0033] Optionally, the formula for obtaining the positioning information embedding result is:
[0034]
[0035] in, are the positioning information embedding results in the horizontal and vertical directions respectively, H is the height of the input feature map, x c (h,i) is the output channel in the horizontal direction, c, height h, strip pooling is performed on the i-th feature, W is the width of the input feature map, x c (j,w) is the output channel c along the horizontal direction, the width is w, and the i-th feature is strip pooled.
[0036] Optionally, the formula for obtaining the position attention in the horizontal and vertical directions is:
[0037]
[0038] Y = x c ×y h ×y w
[0039] Among them, y h and w are used to obtain the position attention in the horizontal and vertical directions respectively, σ is a nonlinear activation function, and F h and F w Respectively represented as one-dimensional convolution, x c is a convolution with c channels, and Y is the final output of ELA.
[0040] Optionally, obtaining the improved YOLOv8 network model by training with a training set includes:
[0041] Input the training set, and optimize the model parameters in combination with the loss function to obtain the improved YOLOv8 network model; wherein the loss function includes: binary cross entropy, DFL loss function and CIoU loss function;
[0042] The binary cross entropy is used to calculate the classification loss of the target, determine the target category, and output the confidence level;
[0043] The DFL loss function and the CIoU loss function are used to calculate the regression loss of the target; the DFL loss function optimizes the probabilities of the left and right positions closest to the label in the form of cross entropy, thereby focusing on the distribution of the target position and the adjacent area.
[0044] The beneficial effects of the present invention are:
[0045] The present invention designs a spatial pyramid structure based on large kernel separable convolution, which efficiently processes high-level features in meter defect images, reduces computational complexity and parameter quantity, and reduces storage requirements, so that the model can run more efficiently in a substation environment with limited computing resources, thereby improving computational efficiency and detection effect; by adding an efficient local attention mechanism to the spliced features in the YOLOv8 neck structure, the neck network is enhanced to capture more comprehensive information by combining different features, and feature fusion is efficiently completed, so that the model can accurately locate meter defects and improve detection accuracy; DFL Loss and CIoU Loss are used together for detection border regression, so that the model can focus on the target position and adjacent areas more quickly; compared with the existing substation meter defect detection method, the model trained by the present invention effectively improves the detection accuracy of meter defects, and can realize the positioning and detection of meter defects in substation images. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0047] Figure 1 A flow chart of a meter defect detection method based on efficient local attention and large separable kernel attention according to an embodiment of the present invention;
[0048] Figure 2 This is a schematic diagram of the improved YOLOv8 network structure of an embodiment of the present invention;
[0049] Figure 3 A schematic diagram of data processing and annotation according to an embodiment of the present invention;
[0050] Figure 4 A schematic diagram of a spatial pyramid structure based on large kernel separable convolution according to an embodiment of the present invention;
[0051] Figure 5 A schematic diagram of the structure of an efficient attention ELA module according to an embodiment of the present invention;
[0052] Figure 6 This is a diagram showing the effect of substation meter defect detection according to an embodiment of the present invention. DETAILED DESCRIPTION
[0053] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0054] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0055] like Figure 1-2 As shown, this embodiment discloses a meter defect detection method based on efficient local attention and large separable kernel attention, including: obtaining a defect image of a substation meter to be detected; inputting the defect image of the substation meter to be detected into an improved YOLOv8 network model to obtain a defect detection result; the improved YOLOv8 network model is obtained by training using a training set; improving the YOLOv8 network model: adding a separable convolution structure after the SPPF layer in the backbone network of the YOLOv8 network model, and adding an efficient local attention module after the neck network is skipped, that is, after upsampling and splicing shallow features of the same size; the separable convolution structure is used for high-level feature extraction; the efficient local attention module is used for efficient fusion of features of different sizes; wherein the training set includes: substation meter defect images and defect labels.
[0056] Specifically:
[0057] S1: Construct a substation meter defect dataset. Based on the collected substation defect images, screen and clean them to obtain labeled images of substation meter defects covering different devices, different angles, and different types. Substation meter defects include blurred meters and damaged meters.
[0058] S2: Design a spatial pyramid structure based on large kernel separable convolution to efficiently complete high-level feature extraction in the backbone network, solve the problem of large model parameters and long training time caused by the introduction of optimization methods into the model, and balance the accuracy and speed of the model;
[0059] S3: Achieve efficient feature fusion, add an efficient local attention mechanism to the neck network of the model, achieve efficient fusion of features of different sizes in the neck network, ensure rich feature information input to the detection structure, and improve the model's positioning accuracy for meter defects;
[0060] S4: Training the model, using DFL Loss and CIoU Loss in YOLOv8 for learning and training. DFL Loss can effectively deal with the problem of category imbalance and ensure that the model learns about small categories or difficult-to-detect categories.
[0061] Furthermore, obtaining the training set includes: cleaning the collected original substation meter defect images according to the standards of the front orientation of the target object, obvious defect features, and the defect area occupying a preset area of the entire image, so as to obtain a picture set of the original substation meter defect images.
[0062] Specifically, step S1 includes:
[0063] S11: Cleaning the acquired meter defect images according to the criteria that the target object is facing in the front direction, the defect features are obvious, and the defect area occupies 1 / 3 of the entire image, and selecting the images with clear images and a large number of substation meter defects to obtain a picture set;
[0064] S12: Use LabelImg to label the image set. The labeling category is metering defects. When multiple target objects appear in an image at the same time, pay attention to the gap between the labeling boxes. If the gap size of the labeling boxes is less than 1 / 2 of the target size, it is labeled as one box to avoid overlapping labeling boxes. Figure 3 .
[0065] Furthermore, obtaining the separable convolution structure includes: dividing the 2D weight convolution kernel of the depth convolution and the depth extension convolution in the LKA block into two cascaded 1D separable weight convolution kernels to obtain the depth separable convolution and the depth separable extension convolution, and maintaining the basic structure of the LKA block to obtain the separable convolution structure.
[0066] Furthermore, the separable convolution structure includes: depthwise separable convolution, which is used to divide the number of channels of the input feature map into multiple groups if the size of the input feature map is the same as the size of the output feature map, and use the convolution kernel to convolve each group, and splice multiple groups to obtain the channels of the output feature map to capture local spatial information; depthwise separable expanded convolution, which is used to introduce an expansion rate into the convolution layer, thereby defining the spacing between values when the convolution kernel in the convolution layer processes data, that is, the number of intervals between each point of the convolution kernel, and adding holes to expand the receptive field to capture global spatial information.
[0067] Specifically, step S2 includes:
[0068] YOLOv8 with high detection accuracy and fast detection speed is selected as the benchmark model. The model backbone network CSPDarknet53 consists of Conv, C2f, SPPF and bottleneck modules. A large kernel separation convolution (LSKC) structure is added after the SPPF layer to enhance the model backbone network's ability to extract and process high-level features.
[0069] The large kernel separable convolution LSKC structure consists of two parts: depth convolution and depth expansion convolution. The depth convolution captures local spatial information, and the depth expansion convolution is responsible for capturing the global spatial information of the depth convolution output. Both convolution modules are small kernel size convolution modules decomposed from large kernel size convolution modules. The equivalent improved LKA structure is cascaded, and finally the Hadamard product operation is performed to obtain the output with the same feature size as the input, such as Figure 4 shown.
[0070] Large Kernel Separable Convolution (LSKC) consists of depthwise convolution and depthwise scalable convolution. The two convolution modules are spatially separable. The convolution module with a large kernel size is decomposed into a convolution module with a small kernel size, that is, the 2D convolution kernel is split into two cascaded 1D convolution kernels. Kernel decomposition helps alleviate the problem of quadratic increase in computational cost caused by using a large kernel size alone.
[0071] LSKC uses the basic LSA block structure, given an input feature map F∈R C×H×W , where C is the number of input channels, H and W represent the height and width of the feature map respectively. The conventional method to design LKA is to use a large convolution kernel in a 2D depth convolution. The calculation process of LKA is shown in the following formula:
[0072]
[0073] A C =W 1×1 *Z C (2)
[0074]
[0075] Where * is convolution, is the Hadamard product, Z C It is done by adding a convolution kernel of size k×k The output of the depthwise convolution is obtained by convolving the input feature map F. It should be noted that each channel C in F is convolved with the corresponding channel in the convolution kernel W. The k in formula (1) also represents the maximum receptive field of the convolution kernel W. Then, through the convolution kernel W1×1 To obtain the attention map A C The output of LKA is the attention map A C And the input feature map F C Hadamard calculation results.
[0076] The equivalent and improved configuration of LKA is obtained by dividing the 2D weight convolution kernel of the depth convolution and depth dilation convolution into two cascaded 1D separable weight convolution kernels, that is, obtaining depth separable convolution and depth separable dilation convolution, and maintaining the LKA structure to obtain LSKC. The calculation process of LSKC is shown in the formula:
[0077]
[0078]
[0079] A C =W 1×1 *Z C (6)
[0080]
[0081] Where * is convolution, is the Hadamard product, d is the expansion rate, represents the output of the deep convolution with kernel size (2d-1)×1 and 1×(2d-1), which captures the local spatial information and compensates for the grid effect of the subsequent deep expansion convolution. C The kernel size is and The depth of the dilated convolution, A C To obtain the attention map by performing kernel convolution on the output of the depth convolution, Note that Figure A C With the input feature map F C The output after Hadamard product, is a convolution kernel of size (2d-1)×1, For size The convolution kernel is , H is the height of the feature map, and W is the width of the feature map.
[0082] Among them, for depth convolution, if the size of the input feature map is equal to the size of the output feature map, and the number of channels C of the input feature map is divided into C groups, that is, each group has only 1 channel, and then a 1×1×K×K convolution kernel is used to convolve each group, and finally the results of the C groups are concatenated, the channels of the output feature map are still C×H×W. This is the calculation process of depth-separable convolution.
[0083] Deep dilated convolution, also known as hole convolution, introduces a parameter called dilation rate to the convolution layer, which defines the spacing between values when the convolution kernel processes data, representing the number of intervals between points in the convolution kernel. Deep dilated convolution uses the addition of holes to expand the receptive field. For example, the original 3x3 convolution kernel can have a 5x5 (when the dilation rate is 2) or larger receptive field with the same number of parameters and computation, thus eliminating the need for downsampling. The calculation formula for the receptive field of the extended convolution is:
[0084] F i+1 =(2 i+2 -1)×(2 i+2 -1) (8)
[0085] Where F i+1 is to expand the receptive field of the convolution, and i+1 is the expansion rate.
[0086] Furthermore, the efficient local attention module includes: a positioning information embedding unit, which is used to pool the input feature map horizontally and vertically, and then use the element values in the pooling kernel to find the average value, and use the average value as the pooling output value, and convolve the output feature map, and expand it horizontally and vertically respectively, and sum the pixels corresponding to the same position of the expanded feature map pixel by pixel to obtain a spliced feature map, and multiply the spliced feature map with the corresponding pixels of the input feature map after convolution and sigmoid processing to obtain the positioning information embedding result; a positioning attention generation unit, which is used to enhance the positioning information based on the one-dimensional convolution sequence, and process the enhanced position information through group normalization to obtain position attention in the horizontal and vertical directions.
[0087] Specifically, step S3 includes:
[0088] The baseline model YOLOv8 uses PANet in its neck network. To better integrate semantic information at different levels, the Efficient Local Attention (ELA) module is added after the deep features are sampled and concatenated with the shallow features of the same size.
[0089] ELA includes two steps: localization information embedding and localization attention generation. Localization information embedding captures long-range spatial dependencies by utilizing strip pooling and average pooling each channel in two spatial ranges; localization attention generation uses group normalization and one-dimensional convolution sequence to enhance localization information, such as Figure 5 shown.
[0090] ELA mainly includes two steps: localization information embedding and localization attention generation. Localization information embedding captures long-range spatial dependencies by utilizing strip pooling operations and performs average pooling on each channel in two spatial ranges: along the horizontal direction (H, 1) and along the vertical direction (1, W). The output in the horizontal direction represents a channel of c and a height of h, and the output in the vertical direction represents a channel of c and a width of w, which are expressed as:
[0091]
[0092]
[0093] in, are the positioning information embedding results in the horizontal and vertical directions respectively, H is the height of the input feature map, x c (h,i) is the output channel in the horizontal direction, c, height h, strip pooling is performed on the i-th feature, W is the width of the input feature map, x c (j,w) is the output channel c along the horizontal direction, the width is w, and the i-th feature is strip pooled.
[0094] From the results of the localization information embedding, we can see that the localization information is a sequential signal in the channel. Therefore, using a one-dimensional convolution sequence signal can effectively enhance the interactive ability of localization information embedding. At the same time, since the computational complexity of one-dimensional convolution is small, it will not bring additional burden to the model, so that the entire ELA can quickly and accurately locate the area of interest. Localization attention generation uses GN to process the enhanced location information, and the localization attention is represented in the horizontal and vertical directions as follows:
[0095]
[0096]
[0097] Y = x c ×y h ×y w (13)
[0098] Where y h and w are used to obtain the position attention in the horizontal and vertical directions respectively, σ is a nonlinear activation function, and F h and F w Respectively represented as one-dimensional convolution, x c is a convolution with c channels, and Y is the final output of ELA.
[0099] The calculation process of strip pooling includes: input a feature map C×H×W, and convert the input feature map into H×1 and 1×W after horizontal and vertical stripe pooling; use the averaging method to average the element values in the pooling kernel, and use the value as the pooling output value; then expand the output feature map left and right and up and down respectively through convolution, and the two feature maps have the same size after expansion. The pixel-by-pixel summation of the corresponding position of the expanded feature map is performed to obtain the H×W feature map, and finally, the final output result is obtained by multiplying the corresponding pixels of the original input map through 1×1 convolution and sigmoid processing.
[0100] Group normalization (GN) is between layer normalization and instance normalization. GN first divides the channels into many groups, changes the dimension of the feature vector from [N, C, H, W] to [N, G, C / / G, H, W], and then normalizes each group to a dimension of [C / / G, H, W]. The normalization calculation formula is:
[0101]
[0102]
[0103] In the formula, μ i is the mean, σ i is the calculation area of the standard deviation, x i is the feature calculated for one layer, i is the index, ε is a small constant, S i is the set of pixels for which the mean and standard deviation are calculated, and m is the size of the set.
[0104] The definition formula of the group normalization set is:
[0105]
[0106] In the formula, S i is a set of definitions, k N (or i N ) represents the sub-index of k (or i) along the N dimension, C / G represents the number of channels in each group, and floor represents the rounding down operation. It means that each group of channels is assumed to be stored sequentially along the C dimension, and indexes i and k are in the same group of channels.
[0107] Furthermore, using the training set to train an improved YOLOv8 network model includes: inputting the training set, and optimizing the model parameters in combination with the loss function to obtain the improved YOLOv8 network model; wherein the loss functions include: binary cross entropy, DFL loss function and CIoU loss function; binary cross entropy is used to calculate the classification loss of the target, judge the target category, and output the confidence; the DFL loss function and the CIoU loss function are used to calculate the regression loss of the target; the DFL loss function optimizes the probabilities of the left and right positions closest to the label in the form of cross entropy, thereby focusing on the distribution of the target position and the adjacent area.
[0108] Specifically, step S4 includes:
[0109] S41: The data set is divided into a training set and a validation set in a ratio of 8:2, where the validation set is also the test set. The data enhancement method is used to expand and transform the images, effectively increasing the sample images and avoiding the performance degradation of the model due to insufficient data.
[0110] S42: In the training loss, binary cross entropy (BCE) is used to calculate the classification loss of the target. For each category, it is judged whether it is of this type and the confidence is output. DFL Loss and CIoU Loss are used to calculate the regression loss of the target. DFL Loss optimizes the probability of the two positions closest to the label, one on the left and one on the right, in the form of cross entropy, so that the network can focus on the distribution of the target position and the adjacent area more quickly.
[0111] After YOLOv8 introduced the Center-based methods of Anchor-Free, the model changed from outputting the predicted box size offset to outputting the distance between the left, top, right, and bottom borders of the predicted target box and the target center point. In order to cooperate with Anchor-Free and improve generalization, DFL Loss was added. The calculation process is as follows:
[0112] DFL(S i ,S i+1 )=-((y i+1 -y)log(S i )+(yy i )log(S i+1 )) (17)
[0113] In the formula, S i is the predicted value output by the network, S i+1 is the network's nearest predicted value; y is the actual value of the label, y i is the label integral value, y i+1 is the adjacent label integral value, DFL(S i ,S i+1) is the loss between the predicted value and the leading predicted value. The specific process of converting the label to DFL form is: convert the label value from the output prediction box size offset x, y, w, h to the distance l, t, r, b from the left, top, right, and bottom borders of the output prediction target box to the target center point; calculate the four values of l, t, r, and b of the label and convert them into the integral form required by DFL.
[0114] S43: Based on the trained optimal model, detect and identify the defects of substation meters and evaluate the model performance. The evaluation indicators include precision, recall, and average precision (mAP). The calculation process is as follows:
[0115]
[0116]
[0117]
[0118]
[0119] Where TP represents the number of true meter defects detected, FP represents the number of falsely detected flagged defects, and FN represents the number of undetected meter defects.
[0120] The advantages of the model of the present invention can be illustrated by comparing the YOLOv8 basic model, the spatial pyramid structure based on large kernel separable convolution used by YOLOv8, the efficient local attention mechanism ELA used by YOLOv8, and the model of the present invention (using two optimization methods at the same time). Table 1 is a comparison between the model of the present invention and the above model settings. It can be seen that the model of the present invention performs well in terms of mAP score, precision and recall rate indicators.
[0121] Table 1
[0122]
[0123] The detection effect of the method of the present invention is as follows Figure 6As shown. Based on the YOLOv8 model, the present invention first designs and uses a spatial pyramid structure based on large kernel separable convolution to efficiently process high-level features, reduce computational complexity and parameter quantity, and improve computational efficiency. Then, by adding an efficient local attention mechanism to the neck structure of the YOLOv8 model, the model feature fusion capability and the positioning accuracy of the model are improved; the experimental results on the constructed substation meter defect dataset show that the model can significantly improve the accuracy, recall rate and average accuracy of meter defect detection, indicating that the detection method based on efficient local attention and large kernel separable convolution ensures that the model accurately locates the area of interest, and can efficiently process high-level features, and has a good detection effect on substation meter defects.
[0124] The embodiments described above are only descriptions of the preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the design spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by ordinary technicians in this field should all fall within the protection scope determined by the claims of the present invention.
Claims
1. A meter defect detection method based on efficient local attention and large separable kernel attention, characterized in that: include: Acquire defect images of substation meters to be inspected; Inputting the defect image of the substation meter to be detected into the improved YOLOv8 network model to obtain a defect detection result; the improved YOLOv8 network model is obtained by training using a training set; Improve the YOLOv8 network model: add a separable convolution structure after the SPPF layer in the backbone network of the YOLOv8 network model, and add an efficient local attention module after the neck network is skipped, that is, after the upsampling and the shallow features of the same size are spliced; The separable convolution structure is used for high-level feature extraction; The efficient local attention module is used to efficiently fuse features of different sizes; The training set includes: substation meter defect images and defect labels.
2. The meter defect detection method based on efficient local attention and large separable kernel attention according to claim 1 is characterized in that: Acquiring the training set includes: The collected original meter defect images are cleaned according to the front orientation of the target object, obvious defect features, and the defect area occupies a preset area of the entire image, so as to obtain a picture set of the original substation meter defect images.
3. The meter defect detection method based on efficient local attention and large separable kernel attention according to claim 1 is characterized in that: Obtaining the separable convolution structure includes: The 2D weight convolution kernels of the depth convolution and depth extension convolution in the LKA block are divided into two cascaded 1D separable weight convolution kernels to obtain depth separable convolution and depth separable extension convolution, and the basic structure of the LKA block is maintained to obtain the separable convolution structure.
4. The meter defect detection method based on efficient local attention and large separable kernel attention according to claim 3 is characterized in that: The separable convolution structure includes: The depthwise separable convolution is used to divide the number of channels of the input feature map into multiple groups if the size of the input feature map is the same as the size of the output feature map, use a convolution kernel to convolve each group, and splice multiple groups to obtain the channels of the output feature map to capture local spatial information; The depthwise separable dilated convolution is used to introduce a dilation rate into the convolution layer, thereby defining the spacing between values when the convolution kernel in the convolution layer processes data, that is, the number of intervals between each point of the convolution kernel, and adding holes to expand the receptive field to capture global spatial information.
5. The meter defect detection method based on efficient local attention and large separable kernel attention according to claim 4 is characterized in that: The formula for expanding the receptive field is: F i+1 =(2 i+2 -1)×(2 i+2 -1) Among them, F i+1 is to expand the receptive field of the convolution, and i+1 is the expansion rate.
6. The meter defect detection method based on efficient local attention and large separable kernel attention according to claim 1, characterized in that: The formula for high-level feature extraction by adding a separable convolution structure is: AND C =In 1×1 *WITH C Among them, * is convolution, is the Hadamard product, d is the expansion rate, represents the output of the deep convolution with kernel size (2d-1)×1 and 1×(2d-1), which captures the local spatial information and compensates for the grid effect of the subsequent deep expansion convolution. C The kernel size is and The depth of the dilated convolution, A C To obtain the attention map by performing kernel convolution on the output of the depth convolution, Note that Figure A C With the input feature map F C The output after Hadamard product, is a convolution kernel of size (2d-1)×1, For size The convolution kernel is , H is the height of the feature map, and W is the width of the feature map.
7. The meter defect detection method based on efficient local attention and large separable kernel attention according to claim 1 is characterized in that: The efficient local attention module includes: A positioning information embedding unit is used to perform horizontal and vertical stripe pooling on the input feature map, calculate the average value using the element values in the pooling kernel, take the average value as the pooling output value, perform convolution on the output feature map, and expand it horizontally and vertically respectively, sum the pixels corresponding to the same position of the expanded feature map pixel by pixel to obtain a spliced feature map, process the spliced feature map by convolution and sigmoid, and then multiply it with the corresponding pixels of the input feature map to obtain a positioning information embedding result; The positioning attention generation unit is used to enhance the positioning information based on a one-dimensional convolution sequence, process the enhanced position information through group normalization, and obtain position attention in the horizontal and vertical directions.
8. The meter defect detection method based on efficient local attention and large separable kernel attention according to claim 7 is characterized in that: The formula for obtaining the embedding result of the positioning information is: in, are the positioning information embedding results in the horizontal and vertical directions respectively, H is the height of the input feature map, x c (h,i) is the output channel along the horizontal direction, c, height h, strip pooling is performed on the i-th feature, W is the width of the input feature map, x c (j,w) is the output channel c along the horizontal direction, the width is w, and the i-th feature is strip pooled.
9. The meter defect detection method based on efficient local attention and large separable kernel attention according to claim 7, characterized in that: The formula for obtaining position attention in the horizontal and vertical directions is: Y=x c ×y h ×y w Among them, y h and w are used to obtain the position attention in the horizontal and vertical directions respectively, σ is a nonlinear activation function, and F h and F w Represented as one-dimensional convolution, x c is a convolution with c channels, and Y is the final output of ELA.
10. The meter defect detection method based on efficient local attention and large separable kernel attention according to claim 1, characterized in that: The improved YOLOv8 network model obtained by training with the training set includes: Input the training set, and optimize the model parameters in combination with the loss function to obtain the improved YOLOv8 network model; wherein the loss function includes: binary cross entropy, DFL loss function and CIoU loss function; The binary cross entropy is used to calculate the classification loss of the target, determine the target category, and output the confidence level; The DFL loss function and the CIoU loss function are used to calculate the regression loss of the target; the DFL loss function optimizes the probabilities of the left and right positions closest to the label in the form of cross entropy, thereby focusing on the distribution of the target position and the adjacent area.
Citation Information
Cited By
Convolution kernel decomposition and cascade quantization optimization method based on Z transformation, processing method and related device
CN122491354A