Substation equipment defect detection method based on improved Yolov8s network

By improving the YOLOv8s network, the cross-branch feature fusion module CBFF and Inner-IoU loss function are introduced, which solves the problem of insufficient accuracy in substation equipment defect detection in complex environments, and achieves higher detection accuracy and lower leakage detection rate.

CN120147262APending Publication Date: 2025-06-13CHINA THREE GORGES UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510219295.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In complex environments, the accuracy of defect detection of substation equipment is insufficient and problems of error detection and missed detection are prone to occur.

Method used

Improve the YOLOv8s network, introduce the cross-branch feature fusion module CBFF, adopt multi-branch feature extraction and fusion strategies, enhance the feature extraction capabilities of the network, and optimize the model using the Inner-IoU loss function.

Benefits of technology

The accuracy and accuracy of defect detection of substation equipment is improved, the leakage detection rate is reduced, and the parameter quantity of the model is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147262A_ABST
    Figure CN120147262A_ABST
Patent Text Reader

Abstract

The invention relates to a transformer substation equipment defect detection method based on an improved Yolov8s network, and the method comprises the following steps: collecting an image of a transformer substation equipment defect, and constructing a data set D1; an improved YOLOv8s network is constructed; the data set is input into the improved YOLOv8s network, and feature maps of different scales are generated; inputting a training set in the data set D1 into the improved YOLOv8s network for training, and performing optimization according to an Inner-IoU loss function; and inputting a verification set in the data set D1 into the improved YOLOv8s network, and evaluating the network by using an evaluation index. The substation equipment defect detection method based on the improved Yolov8s network adapts to various complex geographical conditions and environments, a cross-branch feature fusion module CBFF is introduced, and through a multi-branch feature extraction and fusion strategy, the feature extraction capability of the network is enhanced, and the substation equipment defect detection precision and accuracy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of object detection, and specifically relates to a substation equipment defect detection method based on an improved Yolov8s network. Background Art

[0002] A substation is a key hub facility in the power system, undertaking the tasks of efficient power transmission and reasonable power distribution. Its operating status is directly related to the reliability and safety of the entire power system. The substation is long-term exposed to a changeable and complex outdoor environment, resulting in the equipment facing multiple challenges such as wind and rain erosion, temperature difference changes, and electromagnetic interference. Therefore, regular inspection and maintenance of substation equipment are an indispensable part of ensuring the stable operation of the substation and the overall safety of the power system.

[0003] Traditional inspection work mainly relies on inspectors to conduct inspections. However, manual inspection highly depends on the experience and judgment of inspectors, is easily affected by environmental and personal factors, and in the face of complex geographical conditions or bad weather, some areas may not be covered in time. Moreover, the inspection work is often carried out in a high-voltage environment and may face adverse factors such as electromagnetic interference and bad weather, increasing the work risk of inspectors. Summary of the Invention

[0004] The technical problem of the present invention is: to propose an improved substation equipment defect detection and identification method to solve the problems of insufficient detection accuracy, misdetection, and missed detection in complex environments.

[0005] To solve the above problems, the present invention constructs an improved YOLOv8s network, introduces a cross-branch feature fusion module CBFF in the backbone network, adopts a multi-branch feature extraction and fusion strategy, enhances the feature extraction ability of the network, improves the detection accuracy and convergence speed of the model, and weights the recognition accuracy of the predicted target position.

[0006] The technical solution of the present invention is a substation equipment defect detection method based on an improved Yolov8s network, including the following steps: S1: Collect images of substation equipment defects and construct a data set D1; S2: Construct an improved YOLOv8s network; S3: Input the data set into the improved YOLOv8s network to generate feature maps of different scales; S4: Input the training set in the data set D1 into the improved YOLOv8s network for training and optimize it according to the Inner-IoU loss function; S5: Input the validation set in the data set D1 into the improved YOLOv8s network and evaluate the network with evaluation indicators.

[0007] Further, in step S1, images of substation equipment defects are selected, including annotating the images; the annotation includes bounding box annotation of the defect areas and assigning corresponding category labels to each defect; the dataset D1 also includes preprocessing the dataset, including unifying the image size, normalizing, and data augmentation, which are used to improve the robustness and generalization ability of the model.

[0008] Further, in step S2, the YOLOv8s network is improved, including a backbone network, a neck sub-network, and a decoupled detection head; the backbone network extracts multi-level features and generates feature maps of different scales; the neck sub-network performs feature fusion and enhancement on the feature maps of different scales to improve the expression ability of features at different scales and provide richer feature representations for subsequent detection; the decoupled detection head outputs detection results including the defect target category, position coordinates, and confidence.

[0009] Preferably, the backbone network includes a convolutional module and a cross-branch feature fusion module CBFF connected in series in sequence; the cross-branch feature fusion module CBFF includes an initial convolutional layer, a Split layer, RFC_neck layers in multiple branches, a Concat layer, and a convolutional layer.

[0010] Preferably, the cross-branch feature fusion module CBFF includes the following steps: 1) The input feature map is processed through a 1×1 convolutional layer for dimension conversion to increase the number of channels. 2) The dimension-converted feature map is channel-segmented through the Split layer to generate different branches. 3) The different branches are processed through the RFC_neck layer to obtain optimized branch outputs. 4) The different optimized branch outputs are concatenated through a Concat operation and then processed through a 1×1 convolutional layer to restore the number of channels and obtain the final output result.

[0011] Further, the RFC_neck layer is composed of a conventional convolutional layer and an RFCBAMConv layer connected in series. The conventional convolutional layer is used to extract basic features and perform preliminary feature fusion. The RFCBAMConv layer combines the channel attention mechanism and the spatial attention mechanism to enhance the feature extraction ability of the model from both spatial and channel dimensions; the RFC_neck layer is used to strengthen the network's ability to learn features.

[0012] Furthermore, in the RFCBAMConv layer, the input feature map compresses the spatial dimension through global average pooling in the left branch and generates a vector representation of channel information through a fully connected layer. The input feature map passes through the right branch, where convolution is used to extract receptive field spatial features, and average pooling and max pooling operations are performed to capture spatial distribution information and generate a vector representation of spatial information. Then, the vector representations of the two parts of information obtained are used to weight the transformed feature map. The outputs of the left and right branches integrate information through ordinary convolution and are connected to the input feature map with a residual connection to obtain the output feature map.

[0013] Furthermore, the neck sub-network includes a Spatial Pyramid Pooling Module (SPPF), a Convolution Module (Conv), an Upsample Module, and a C2f Module. C2f is used to decode the prediction head of the detection layer and output the category, position coordinates, and confidence of the defective target.

[0014] Preferably, in step S4, the Inner-IoU loss function introduces a ratio factor to handle small objects and morphological changes by focusing on the overlapping region inside the target box.

[0015] Preferably, the calculation formula of the Inner-IoU loss function is: ; ; ; ; ; ; ; In the formula, respectively represent the center point coordinates of the ground truth box, respectively represent the center point coordinates of the predicted box, , respectively represent the width and height of the ground truth box, , respectively represent the width and height of the predicted box, represents the scale factor, represents the lower left corner coordinate of the ground truth box, represents the lower right corner coordinate of the ground truth box, represents the upper left corner coordinate of the ground truth box, represents the upper right corner coordinate of the ground truth box, represents the lower left corner coordinate of the predicted box, represents the lower right corner coordinate of the predicted box, represents the upper right corner coordinate of the predicted box, represents the upper right corner coordinate of the predicted box, Indicates the area of the intersection region between the predicted bounding box and the ground truth bounding box. Indicates the area of the union region between the predicted bounding box and the ground truth bounding box.

[0016] Compared with the prior art, the beneficial effects of the present invention include: 1) The present invention proposes a method for detecting substation equipment defects based on an improved Yolov8s network, which is adaptable to various complex geographical conditions and environments. And a cross-branch feature fusion module CBFF is introduced. Through a multi-branch feature extraction and fusion strategy, the feature extraction ability of the network is enhanced, and the accuracy and precision of substation equipment defect detection are improved.

[0017] 2) The present invention proposes a method for detecting substation equipment defects based on an improved Yolov8s network. The cross-branch feature fusion module CBFF integrates the RFC_neck sub-module with a double convolutional structure, which improves the improved Yolov8s network's ability to extract and process feature information in images.

[0018] 3) The present invention proposes a method for detecting substation equipment defects based on an improved Yolov8s network. The Inner-IoU function is used as the loss function to improve the accuracy of the network in processing the scale and shape changes of targets, thereby improving the accuracy of substation equipment defect detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The present invention will be further described below in conjunction with the drawings and embodiments.

[0020] Figure 1 is the network structure diagram of the improved Yolov8s network of the present invention; Figure 2 is the network structure diagram of the cross-branch feature fusion module CBFF of the improved Yolov8s network of the present invention; Figure 3 is the structure diagram of the RFC_neck layer of the improved Yolov8s network of the present invention; Figure 4 is the structure diagram of the RFCBAMConv layer of the improved Yolov8s network of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0021] As Figure 1 shown, a method for detecting substation equipment defects based on an improved Yolov8s network includes the following steps: S1: Collect images of substation equipment defects and construct a data set D1; In step S1, images of substation equipment defects are selected, including annotating the images; the annotation includes bounding box annotation of the defect areas and assigning corresponding class labels to each defect; the dataset D1 also includes preprocessing the dataset, including unifying the image size, normalizing, and data augmentation, which are used to improve the robustness and generalization ability of the model.

[0022] S2: Construct an improved YOLOv8s network; The improved YOLOv8s network includes a backbone network, a neck sub-network, and a decoupled detection head; the backbone network extracts multi-level features and generates feature maps of different scales; the neck sub-network performs feature fusion and enhancement on the feature maps of different scales, improves the expression ability of features at different scales, and provides a richer feature representation for subsequent detection; the decoupled detection head outputs detection results including the defect target category, position coordinates, and confidence.

[0023] The backbone network includes a convolutional module and a cross-branch feature fusion module CBFF connected in series in sequence; the cross-branch feature fusion module CBFF includes an initial convolutional layer, a Split layer, RFC_neck layers in multiple branches, a Concat layer, and a convolutional layer.

[0024] As Figure 2 shown, the cross-branch feature fusion module CBFF includes the following steps: 1) The input feature map is processed through a 1×1 convolutional layer for dimension conversion to increase the channel dimension; 2) The dimension-converted feature map is channel-segmented through the Split layer to generate different branches; 3) The different branches are processed through the RFC_neck layer to obtain optimized branch outputs; 4) After the different optimized branch outputs are concatenated through the Concat operation, they are processed through a 1×1 convolutional layer to restore the channel dimension and obtain the final output result.

[0025] As Figure 3 shown, the RFC_neck layer is composed of a conventional convolutional layer and an RFCBAMConv layer connected in series. The conventional convolutional layer is used to extract basic features and perform preliminary feature fusion. The RFCBAMConv layer combines the channel attention mechanism and the spatial attention mechanism to enhance the feature extraction ability of the model from both the spatial and channel dimensions; the RFC_neck layer is used to strengthen the network's ability to learn features.

[0026] As Figure 4As shown in the figure, in the RFCBAMConv layer, the input feature map compresses the spatial dimension through global average pooling in the left branch, and passes through a fully connected layer to generate a vector representation of the channel information; the input feature map passes through the right branch, where convolution is used to extract the receptive field spatial features, and average pooling and max pooling operations are performed to capture the spatial distribution information and generate a vector representation of the spatial information. Then, the vector representations of the two parts of information obtained are used to weight the transformed feature map; the outputs of the left and right branches integrate information through ordinary convolution and are connected to the input feature map through a residual connection to obtain the output feature map.

[0027] The neck sub-network includes a spatial pyramid pooling module SPPF, a convolution module Conv, an Upsample module, and a C2f module; the C2f is used for decoding the prediction head of the detection layer and outputs the category, position coordinates, and confidence of the defect target.

[0028] S3: Input the data set into the improved YOLOv8s network to generate feature maps of different scales; S4: Input the training set in the data set D1 into the improved YOLOv8s network for training and optimize it according to the Inner-IoU loss function; In step S4, the Inner-IoU loss function introduces a ratio factor to handle small objects and morphological changes by focusing on the overlapping area inside the target box.

[0029] The calculation formula of the Inner-IoU loss function is: ; ; ; ; ; ; ; In the formula, respectively represent the center point coordinates of the ground truth box, respectively represent the center point coordinates of the predicted box, , respectively represent the width and height of the ground truth box, , respectively represent the width and height of the predicted box, represents the scale factor, represents the lower left corner coordinate of the ground truth box, represents the lower right corner coordinate of the ground truth box, represents the upper left corner coordinate of the ground truth box, represents the upper right corner coordinate of the ground truth box, Represents the lower left corner coordinates of the prediction box, Represents the lower right corner coordinates of the prediction box, Represents the upper right corner coordinates of the prediction box, Represents the upper right corner coordinates of the prediction box, Represents the area of the intersection region between the prediction box and the ground truth box, Represents the area of the union region between the prediction box and the ground truth box.

[0030] S5: Input the validation set in the dataset D1 into the improved YOLOv8s network and evaluate the network using evaluation metrics.

[0031] Use the trained model to evaluate on the validation set, evaluate the detection performance of the model, and use precision Precision, recall Recall, mean average precision mAP, and the number of parameters Params as the evaluation metrics of the model. Pression represents the proportion of actual positive samples among the samples predicted as positive by the model, reflecting the accuracy of the detection results. Recall represents the proportion of samples correctly predicted as positive by the model among the actual positive samples, reflecting the integrity of the model and its coverage ability for the target. mAP is an indicator that comprehensively considers precision and recall. A high mAP indicates good overall performance of the model in the detection task. Params represents the number of all trainable parameters in the model, reflecting the complexity and storage requirements of the model; The calculation formula of Pression is as follows: ; In the formula, TP is the number of defective targets successfully identified by the model, and FP is the number of normal devices misidentified as defective.

[0032] The calculation formula of Recall is as follows: ; In the formula, FN represents the number of substation devices with actual defects that are misidentified as normal by the model, that is, the missed defective targets.

[0033] The calculation formula of mAP is as follows: ; ; In the formula, AP represents the area calculated by integration under the curve with recall Recall as the abscissa and precision Precision as the ordinate, and its value ranges between 0 and 1.

[0034] The present invention adopts the Ubuntu 22.04 operating system, is equipped with an Intel Core i5-12400F 2.5GHz processor, 16GB of running memory, and an NVIDIA GeForce RTX 4060 Ti 16GB graphics card. The algorithm implementation is based on the Pytorch 2.3.1 deep learning framework, uses the Python 3.9 language to write programs, and accelerates the training and inference of the network model through CUDA 12.4. To ensure the reproducibility of the experiment, the training of the network model in the experiment is completed under unified settings. The size of the input image during training is 640×640. If the performance of the model does not show effective improvement within 50 consecutive training cycles, an early stopping mechanism is triggered to terminate the training.

[0035] In the present invention, different values are set for the ratio parameter in the Inner-IoU loss function for experiments, and the experimental results are shown in Table 1.

[0036] Table 1

[0037] To verify the superiority of the method of the present invention, this article conducts a comparison with the current mainstream detection methods. The comparison methods include FasterR-CNN, RetinaNet, DDOD, RTMDet, and the YOLO series. The experimental comparison results are shown in Table 2.

[0038] Table 2

[0039] As shown in Table 2, compared with other mainstream algorithms, the method proposed in the present invention has achieved the best results in performance indicators such as accuracy, recall rate, and mean average precision. For the original YOLOv8s algorithm, the improved algorithm has increased the accuracy P, recall rate R, and mean average precision mAP by 1.4%, 3.2%, and 3.7% respectively, and the number of parameters has decreased by 0.6M year-on-year. Compared with the two-stage algorithms of FasterR-CNN and RetinaNet, the improved algorithm has increased both P and mAP by more than 20% and R by more than 10%. Compared with the Anchor-Free-based DDOD and RTMDet algorithms, there are varying degrees of improvements in P, R, and mAP. Compared with the YOLO series networks YOLOv5s, YOLOv7-tiny, YOLOv7, YOLOv9, and YOLOv10s, the improved algorithm has also achieved higher detection accuracy and lower missed detection rate.

[0040] In summary, compared with other currently advanced detection models, the method proposed in the present invention effectively reduces the missed detection rate while improving the detection accuracy. Its parameter quantity has decreased, and the overall model still maintains a reasonable computational complexity. These results fully prove that the method proposed in the present invention can meet the actual needs of substation equipment defect detection.

Claims

1. A substation equipment defect detection method based on an improved Yolov8s network, characterized in that: The following steps are involved: S1: Collect images of substation equipment defects and construct data set D1; S2: Build and improve the YOLOv8s network; S3: Input the data set into the improved YOLOv8s network to generate feature maps of different scales; S4: Input the training set in the dataset D1 into the improved YOLOv8s network for training, and optimize it according to the Inner-IoU loss function; S5: Input the validation set in dataset D1 into the improved YOLOv8s network and evaluate the network using evaluation indicators.

2. According to claim 1, a substation equipment defect detection method based on an improved Yolov8s network is characterized in that: In step S1, the image of the substation equipment defect is selected, including labeling the image; the labeling includes labeling the defect area with a bounding box and assigning a corresponding category label to each defect; the data set D1 also includes preprocessing the data set, including unifying the image size, normalization and data enhancement, so as to improve the robustness and generalization ability of the model.

3. According to claim 1, a substation equipment defect detection method based on an improved Yolov8s network is characterized in that: In step S2, the improved YOLOv8s network includes a backbone network, a neck subnetwork and a decoupling detection head; the backbone network extracts multi-level features to generate feature maps of different scales; the neck subnetwork fuses and enhances the feature maps of different scales to improve the expression ability of features of different scales and provide richer feature representation for subsequent detection; the decoupling detection head outputs detection results including defect target category, location coordinates and confidence.

4. According to claim 3, a method for detecting defects in substation equipment based on an improved Yolov8s network is characterized in that: The backbone network includes a convolution module and a cross-branch feature fusion module CBFF connected in series in sequence; the cross-branch feature fusion module CBFF includes an initial convolution layer, a Split layer, an RFC_neck layer in multiple branches, a Concat layer and a convolution layer.

5. According to claim 4, a method for detecting defects in substation equipment based on an improved Yolov8s network is characterized in that: The cross-branch feature fusion module CBFF comprises the following steps: 1) The input feature map is processed by a 1×1 convolution layer to perform dimension conversion and increase the channel dimension; 2) The feature map after dimension conversion is passed through the Split layer to perform channel segmentation and generate different branches; 3) Process different branches through the RFC_neck layer to obtain optimized branch output; 4) The different optimized branch outputs are concatenated through the Concat operation and processed through a 1×1 convolution layer to restore the channel dimension and obtain the final output result.

6. According to claim 5, a method for detecting defects in substation equipment based on an improved Yolov8s network is characterized in that: The RFC_neck layer includes a conventional convolution layer and an RFCBAMConv layer connected in series. The conventional convolution layer is used to extract basic features and perform preliminary feature fusion. The RFCBAMConv layer combines the channel attention mechanism and the spatial attention mechanism to enhance the feature extraction capability of the model from both spatial and channel dimensions. The RFC_neck layer is used to enhance the network's ability to learn features.

7. According to claim 6, a method for detecting defects in substation equipment based on an improved Yolov8s network is characterized in that: In the RFCBAMConv layer, the input feature map is subjected to global average pooling in the left branch to compress the spatial dimension, and passes through a fully connected layer to generate a vector representation of the channel information; The input feature map passes through the right branch, and the convolution extracts the spatial features of the receptive field. After average pooling and maximum pooling operations, the spatial distribution information is captured and a vector representation of the spatial information is generated. The transformed feature map is then weighted using the vector representation of the two parts of information obtained. The outputs of the left and right branches integrate information through ordinary convolution and perform residual connection with the input feature map to obtain the output feature map.

8. According to claim 3, a method for detecting defects in substation equipment based on an improved Yolov8s network is characterized in that: The neck subnetwork includes a spatial pyramid pooling module SPPF, a convolution module Conv, an Upsample module and a C2f module; the C2f is used for prediction head decoding of the detection layer to output the category, position coordinates and confidence of the defect target.

9. According to claim 1, a method for detecting defects in substation equipment based on an improved Yolov8s network, characterized in that: In step S4, the Inner-IoU loss function introduces a ratio factor to handle small objects and morphological changes by focusing on the overlapping area inside the target box.

10. According to claim 9, a method for detecting defects in substation equipment based on an improved Yolov8s network is characterized in that: The Inner-IoU loss function is calculated as: ; ; ; ; ; ; ; In the formula, Respectively represent the center point coordinates of the real frame, Respectively represent the center point coordinates of the prediction box, , Represent the width and height of the real frame respectively, , Represent the width and height of the prediction box respectively, represents the scale factor, Represents the coordinates of the lower left corner of the real box, Represents the coordinates of the lower right corner of the real box, Indicates the coordinates of the upper left corner of the real box, Represents the coordinates of the upper right corner of the real box, Represents the coordinate of the lower left corner of the prediction box, Represents the coordinates of the lower right corner of the prediction box, Indicates the coordinates of the upper right corner of the prediction box, Indicates the coordinates of the upper right corner of the prediction box, Represents the intersection area of ​​the predicted box and the real box, Represents the union area of ​​the predicted box and the true box.

Citation Information

Cited By

  • Substation equipment defect detection method and device based on deep learning

    CN121095143A