Power transformation equipment defect detection method based on improved YOLOv8n model
By introducing the IRB-SE attention mechanism, BiFPN feature pyramid network and small object detection layer in the YOLOv8n model, the problems of small target feature loss and computing resource limitation in external defect detection of substation equipment are solved, and high-precision and real-time defect detection are achieved.
Patent Information
- Application Number
- CN202411949043.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art has problems such as small target feature loss, neglecting inter-channel correlation and redundant information interference, and computing resource limitations in the detection of external defects of substation equipment, resulting in low detection accuracy and difficulty in real-time detection.
By introducing the IRB-SE attention mechanism into the C2f module of the YOLOv8n model, the model's attention to key local features is enhanced; the bidirectional feature pyramid network (BiFPN) is introduced to optimize the feature fusion layer; and a small object detection layer is added to the detection layer to reduce the loss of small object features.
The model's detection accuracy of substation equipment defects is improved, network parameters are reduced, the accuracy of power meter reading detection is improved, and the requirements of real-time detection are realized.
Smart Images

Figure CN120047714A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of power equipment detection, and in particular to a substation equipment defect detection method based on an improved YOLOv8n model. Background Art
[0002] In the power system, the safety and reliability of substation equipment are crucial for the stable operation of the power system. However, substation equipment is exposed to the natural environment for a long time, and various external defects are likely to occur, such as metal corrosion, equipment oil leakage, hanging objects (foreign objects, etc.), discoloration of the breather silica gel, and breather oil leakage. If these defects cannot be discovered and processed in time, it is very likely to cause equipment failures and even major power accidents. Therefore, it is of great significance to detect external defects of substation equipment in a timely and accurate manner.
[0003] The traditional method for detecting substation equipment defects is manual inspection, which requires a large amount of manpower and material resources, and has low detection efficiency and difficult to guarantee accuracy. To address these issues, researchers have used traditional image recognition methods, such as methods based on histogram of oriented gradients features, infrared images, and acoustic imaging technology, to detect equipment defects. However, these methods are costly, and the imaging effect is easily affected by environmental factors, making it impossible to accurately locate the defect position.
[0004] With the rapid development of computer vision and image processing technologies, automated defect detection methods have begun to be applied. Object detection methods based on deep learning emerge in an endless stream. These methods can be mainly divided into two categories in terms of structure: one is the "two-stage" network represented by Faster R-CNN (Region Convolutional Neural Network) and SPPNet; the other is the "one-stage" network represented by SSD (Single Shot MultiBox Detector) and YOLO (You Only Look Once) series. The most fundamental difference between the two is that the "two-stage" network needs to first generate a series of potential target candidate regions, and then perform in-depth feature extraction and recognition analysis on these regions, while the "one-stage" network directly predicts the classification and location of the target through a one-time neural network operation. Chen et al. improved the YOLOv4 backbone network to address the positioning difficulty problem in power equipment object detection, and introduced the focal loss function to solve the problem of low detection accuracy caused by the imbalance between positive and negative samples. Finally, the improved YOLOv4 algorithm was implemented for surface defect detection of power equipment. Zhang Mingquan et al. based on Faster R-CNN improved the defect detection accuracy by enhancing the input image data, adding an SPP structure to the network, improving the feature fusion method, and improving the classification and bounding box regression loss functions.
[0005] Although the algorithms of these studies perform well in the detection of external defects of power transformation equipment, they still face problems such as complex defects, low detection accuracy for small target defects, and difficulties in mobile deployment in practical applications.
[0006] For the defect detection of power transformation equipment, due to the large variety and wide span of power transformation equipment defects, small-scale defects account for a small proportion in the image, which is very likely to cause problems such as missed detection and misdetection. Using the YOLOv8n model for detection, although the detection effect is better than other object detection models, there are still many shortcomings in the original YOLOv8n model for the detection of various types of defects and small targets. The following are several main shortcomings summarized from the perspective of technical logic:
[0007] (1) Loss of small target features: Since the YOLOv8n model performs continuous downsampling operations, it is easy to lose the feature information of small targets during the feature extraction process. This may lead to the inability to effectively capture the high-detail and small-size features of power transformation equipment, thus affecting the detection accuracy of small targets.
[0008] (2) Ignoring the correlation between channels and interference from redundant information: The YOLOv8n model treats different feature channels equally, ignoring the correlation between channels and resulting in interference from redundant information, thereby reducing the model performance. In the detection of power transformation equipment, different channels may contain different useful information. Ignoring these differences and the complex background or irrelevant features may lead to a decline in model performance.
[0009] (3) Computational resource limitations: Although YOLOv8n is a lightweight model, it is still limited by computational resources in practical applications. In the detection of power transformation equipment defects, if the computational resources are limited, the performance of the model may not be fully utilized, affecting the detection speed and accuracy. Summary of the Invention
[0010] The purpose of the present invention is to overcome the deficiencies of the prior art and propose a method for detecting defects in power transformation equipment based on an improved YOLOv8n model. By adding an IRB-SE attention mechanism to the C2f module, the model's attention to key local features is increased, and the network's ability to extract features at various scales is improved, thereby enhancing the detection accuracy of the model for power transformation equipment defects. In addition, a bidirectional feature pyramid network (BiFPN) is introduced to optimize the feature fusion layer to promote the effective flow and fusion of information between features at different scales; finally, based on the three detection layers in the original model, a small target detection layer is introduced to reduce the loss of small target features and improve the model's detection ability for small targets. This model reduces the network parameters and also improves the accuracy of electricity meter reading detection, meeting the requirements of real-time detection.
[0011] The technical problem solved by the present invention is achieved by adopting the following technical solutions:
[0012] A substation equipment defect detection method based on an improved YOLOv8n model, comprising the following steps:
[0013] Step 1, by improving the YOLOv8n model, an improved YOLOv8n model is obtained;
[0014] Step 2, using the improved YOLOv8n model to detect substation equipment;
[0015] Step 3, obtaining the defect information of the substation equipment according to the calculation result of the improved YOLOv8n model.
[0016] Moreover, the said Step 1 includes the following steps:
[0017] Step 1.1, combining SENetv2 and IRB to construct an IRB-SE attention mechanism, and adding it to the YOLOv8n model;
[0018] Step 1.2, introducing a bidirectional feature pyramid network BiFPN into the YOLOv8n model;
[0019] Step 1.3, adding a small target detection layer to the detection layer of the YOLOv8n model.
[0020] Moreover, the specific implementation method of the said Step 1.1 is: fusing SENetv2 and IRB to obtain an IRB-SE attention mechanism, which is used to capture the advantages of feature information at different levels and scales, and at the same time enabling the network to effectively learn the complex relationships between channels, enhancing the network's selective transmission ability for key features; and introducing the IRB-SE attention mechanism into the C2f module of YOLOv8 to improve it, obtaining an IRB-SEC2f module.
[0021] Moreover, the specific implementation method of obtaining the IRB-SEC2f module is: adding an IRB-SE attention mechanism at the end of the Bottleneck module in the C2f module.
[0022] Moreover, the specific implementation method of the said Step 1.2 is: fusing the outputs of the second and third IRB-SEC2f modules from top to bottom in the backbone network and the outputs of the first and second IRB-SEC2f modules from bottom to top in the neck network in the fourth and fifth Concat modules in the neck network.
[0023] Moreover, the specific implementation method of step 1.3 is as follows: Based on the three detection layers of the YOLOv8n model, a small object detection layer is introduced, which is specifically used for features with high detail and small size characteristics; it is used to capture and retain the features of small objects, so as to improve the ability and effect of the model in small object detection, reduce false detections and missed detections caused by feature loss, and in order to reduce the computational amount of the model and improve the detection speed, the last large object detection layer is deleted.
[0024] Moreover, the specific implementation method of step 2 is as follows:
[0025] Step 2.1: First, convert the dataset into the YOLO format and divide it into a training set and a validation set at a ratio of 8:2 for use during training and validation.
[0026] Step 2.2: Introduce the improved method into the model to generate the corresponding network structure.
[0027] Step 2.3: Build a suitable training environment and find the optimal parameter combination through multiple experiments.
[0028] Step 2.4: Start training the model. The model will calculate the prediction results through forward propagation based on the input image data and annotation information, compare them with the true labels, and then use the backpropagation algorithm to adjust the parameters of the model to minimize the loss function and make the prediction results of the model closer to the real situation.
[0029] Step 2.5: During the training process, use the validation set to regularly evaluate the trained model.
[0030] Step 2.6: Use the trained model to detect the specified picture to obtain the detection results.
[0031] The advantages and positive effects of the present invention are:
[0032] By combining SENetv2 and IRB, the present invention obtains a new attention mechanism, the IRB-SE attention mechanism, which adaptively adjusts the weights of different feature channels, emphasizes important features and suppresses unnecessary information, and can capture feature information at different levels and scales, thereby improving the performance of the model. In addition, in order to promote the effective flow and fusion of information between features of different scales, a bidirectional feature pyramid network (BiFPN) is introduced to optimize the feature fusion layer. Finally, in order to reduce the loss of small object features and improve the small object detection ability, a small object detection layer is introduced on the basis of the three detection layers in the original model. The present invention effectively improves the detection accuracy on the premise of reducing the number of parameters of the model and can meet the real-time requirements, having important practical significance and application value. Description of the Drawings
[0033] Figure 1 It is the structure diagram of the YOLOv8n model;
[0034] Figure 2 It is the network structure diagram of the improved model of the present invention;
[0035] Figure 3 It is the schematic diagram of the structural comparison of ResNeXt, ResNeXt and SaENet;
[0036] Figure 4 It is the network structure diagram of SaENet;
[0037] Figure 5 It is the structural comparison diagram of the traditional residual module and the inverted residual module;
[0038] Figure 6 It is the schematic diagram of the IRB-SE attention mechanism of the present invention;
[0039] Figure 7 It is the schematic diagram of the IRB-SEC2f module of the present invention;
[0040] Figure 8 It is the structural comparison diagram of FPN, PANet, NAS-FPN and BiFPN of the present invention;
[0041] Figure 9 It is the schematic diagram of the dataset example of the present invention. Detailed implementation manners
[0042] The present invention will be further described in detail below with reference to the accompanying drawings.
[0043] A substation equipment defect detection method based on an improved YOLOv8n model includes the following steps:
[0044] Step 1, obtain an improved YOLOv8n model by improving the YOLOv8n model.
[0045] The original YOLOv8n model is as Figure 1 shown. Use the improved IRB-SEC2f module to replace the C2f of the original model. By adaptively adjusting the weights of different feature channels, it highlights the key local features while suppressing the interference of complex environments, thereby improving the detection accuracy of the model; introduce BiFPN to optimize the feature fusion layer, improve the model's ability to extract features of different scales of defects, and enable the model to learn more effective feature representations; at the same time, on the basis of the three detection layers in the original model, introduce a small target detection layer to improve the model's small target detection ability.
[0046] The model structure of the improved model is as Figure 2As shown, after the input image, the image is first preprocessed, adjusted to the corresponding size, and converted into a tensor. Then it enters the Backbone of the model, whose main function is to perform multi-level feature extraction on the input image, gradually from low-level features to high-level features. After that, the output is transmitted to the Neck, which fuses feature information of different scales, enabling the model to detect targets of different sizes. Finally, the Head is responsible for completing the object classification and bounding box regression in the detection task.
[0047] Step 1.1: Combine SENetv2 and IRB to construct the IRB-SE attention mechanism and add it to the YOLOv8n model.
[0048] Since the YOLOv8n model treats different feature channels equally, ignoring the correlation between channels, resulting in redundant information interference, thus reducing the model performance. Therefore, the present invention improves on the basis of the YOLOv8n model, introducing the IRB-SE attention mechanism into the original C2f module, enabling the network to effectively learn the complex relationships between channels, enhancing the network's selective transmission ability for key features, and being able to capture feature information of different levels and scales, improving the model's feature representation ability and adaptability to complex data and task scenarios.
[0049] SENetv2 is an improved SENet network, in which an upgraded SE module, namely the Squeeze Aggregation Excitation (SaE) module, is introduced. Compared with the SE module, the SaE module uses multi-branch dense layers to enhance the network's feature representation ability, thus improving the model's performance. The comparison between the Aggregated Residual (ResNeXt) module, the Squeeze and Excitation (ResNeXt) module, and the proposed Squeeze Aggregation Excitation module is as Figure 3 shown. This figure shows that both the SE module and the SaE module selectively transmit key features. However, the SaE module optimizes this stage by increasing the cardinality between layers, enabling different levels and types of key features to be transmitted and utilized more effectively.
[0050] After the input in the SaENet network undergoes standard convolution operations, first, global average pooling operations are used to squeeze the features, then channel weights are obtained through multi-branch fully connected layers and activation functions, and finally, the convolutional features are scaled and restored to their original form. Then the scaled output is connected to the input in the residual module. The working mechanism is as Figure 4 . In this way, the network can more effectively learn the complex relationships between channels, enhancing the network's selective transmission ability for key features.
[0051] The Inverted Residual Block is a key structure in the MobileNetV2 network. It improves the traditional residual block to enhance the network's efficiency and performance. It first uses a 1×1 convolutional layer to expand the channel dimension, enabling the network to obtain richer feature information in subsequent operations. Then, depthwise separable convolution is performed, which can effectively reduce the computational amount while extracting spatial features within different channels. Finally, another 1×1 convolutional layer is used to compress the number of channels back to an appropriate value. Meanwhile, the residual connection in the inverted residual block helps alleviate the vanishing gradient problem, ensuring the smooth propagation of information in the network and facilitating network training. Comparison of the structures of the traditional residual block and the inverted residual block Figure 5 . Jiangning Zhang et al. rethought the Inverted Residual Block (IRB) and the MHSA / FFN module in Transformer, derived the modernized iRMB, and constructed a lightweight attention model EMO based on iRMB. Through numerous experiments, the superiority of the method was demonstrated.
[0052] According to the iRMB, SENetv2 and IRB are fused to obtain the IRB-SE attention mechanism. That is, it has the advantage of the IRB that can capture feature information at different levels and scales, and can also enable the network to effectively learn the complex relationships between channels, enhancing the network's selective transmission ability for key features. The specific structure of the IRB-SE attention mechanism is as Figure 6 . Among them, 1x1 convolution is used for compressing and expanding the number of channels to optimize the computational efficiency. Depthwise separable convolution (DW-Conv) is used to capture spatial features, while the SENetv2 attention mechanism is used to capture the global dependencies between features. Finally, a residual connection is used to add the input and the compressed output to alleviate the vanishing gradient problem and ensure the smooth propagation of information in the network.
[0053] The IRB-SE is introduced into the C2f module of YOLOv8 and improved to obtain the IRB-SEC2f module. The structure is as Figure 7 , specifically, an IRB-SE attention mechanism is added at the end of the Bottleneck module in the C2f module.
[0054] Step 1.2: Introduce the Bidirectional Feature Pyramid Network (BiFPN) into the YOLOv8n model.
[0055] To enable more effective flow and fusion of information between features of different scales in the YOLOv8n model, integrate feature information from different levels, and accelerate the extraction speed of feature information at different scales. Therefore, the present invention introduces the Bidirectional Feature Pyramid Network (BiFPN) into the original model.
[0056] The Bidirectional Feature Pyramid Network is an efficient multi-scale feature fusion network that optimizes on the basis of the traditional Feature Pyramid Network (FPN). The comparison of the network structures of FPN, PANet, NAS-FPN and BiFPN is as follows Figure 8 , as can be seen from the figure, FPN introduces a top-down path to fuse multi-scale features, enrich the semantics of low-level features, and integrate global and local information; the PANet network used in YOLOv8n adds an additional bottom-up path on the basis of FPN, and reversely transmits the detailed information of the low layer to the high layer, forming a complement to the top-down semantic information; NAS-FPN uses Neural Architecture Search (NAS) to find an irregular feature network topology, which can automatically explore and discover a feature network topology more suitable for specific tasks and datasets, can adapt to features of different scales, and at the same time repeatedly applies the same efficient blocks, reducing the number of model parameters, reducing the computational complexity, and at the same time improving the training and inference speed of the model, but it takes a long time during the search.
[0057] BiFPN balances accuracy and efficiency through efficient bidirectional cross-scale connections and repeated block structures. Bidirectional feature fusion in BiFPN refers to a mechanism that allows information in the feature network layer to flow and fuse in both top-down and bottom-up directions. It enhances the network's ability to fuse features, enabling the network to more effectively utilize information of different scales, thereby improving the performance of object detection. The weighted fusion mechanism adds an additional weight to each input and allows the network to learn the importance of each input feature. Introducing BiFPN in YOLOv8n enables efficient multi-scale feature fusion and transmission, effectively improving the performance and representational ability of the object detection model.
[0058] Step 1.3: Add a small object detection layer to the detection layer of the YOLOv8n model.
[0059] During the continuous downsampling operation of YOLOv8n, the feature information of small objects is very likely to be lost during the transmission process, which will have an adverse impact on the detection performance. Based on this problem, on the basis of the three detection layers of the original model, a small object detection layer is introduced, which is specifically used for features with high details and small sizes. This can better capture and retain the features of small objects, improve the model's ability and effect in small object detection, and reduce problems such as false detection and missed detection caused by feature loss. At the same time, in order to reduce the computational load of the model and improve the detection speed, the last large object detection layer is deleted.
[0060] Step 2: Use the improved YOLOv8n model to detect substation equipment.
[0061] Step 2.1: First, convert the dataset into the YOLO format and divide it into a training set and a validation set at a ratio of 8:2 for use during training and validation.
[0062] Step 2.2: Introduce the improved method into the model to generate the corresponding network structure.
[0063] Step 2.3: Set up a suitable training environment, such as the learning rate, batch size, number of training epochs, etc. The optimal parameter combination can be found through multiple experiments.
[0064] Step 2.4: Start training the model. The model will calculate the prediction results through forward propagation based on the input image data and annotation information, compare them with the true labels, and then use the backpropagation algorithm to adjust the model's parameters to minimize the loss function and make the model's prediction results closer to the real situation.
[0065] Step 2.5: During the training process, use the validation set to regularly evaluate the trained model and monitor metrics such as the model's accuracy, recall rate, F1 value, etc.
[0066] Step 2.6: Use the trained model to detect the specified image and obtain the detection results. After inputting the image, first preprocess the image, adjust it to the corresponding size, and convert it into a tensor; then enter the Backbone of the model, whose main function is to perform multi-level feature extraction on the input image, gradually from low-level features to high-level features; then the output will be transmitted to the Neck, which fuses feature information at different scales so that the model can detect targets of different sizes; finally, the Head is responsible for completing the object classification and bounding box regression in the detection task.
[0067] Step 3: Obtain the defect information of the substation equipment based on the calculation results of the improved YOLOv8n model.
[0068] Example 1
[0069] According to the above substation equipment defect detection method based on the improved YOLOv8n model, the effect of the present invention has been verified through constructing an experimental environment and dataset for calculation.
[0070] In the experimental environment of the embodiment of the present invention, the operating system is Linux Ubuntu 20.04, the CPU is the 13th Gen Intel(R) Core(TM) i9-13900KF, the GPU is the NVIDIA GeForce RTX 4090 with 24GB video memory. CUDA version 12.0 is used to accelerate model training. The deep learning framework is Pytorch 2.1.1, and the Python version is 3.8.18. During training, the hyperparameters are set as follows: the batch size is 8, the number of epochs is 300, the input image size is 640×640, the initial learning rate is 0.01, pre-trained weights are not used, and the SGD optimizer is used.
[0071] In the embodiment of the present invention, 1560 pictures of a substation device with external defects are selected for annotation as the experimental data set. Examples are as follows Figure 9 . The image resolution of each picture is 1024×768. The defects include equipment metal rust, normal breather, discolored breather silica gel, oil leakage from the breather, hanging objects, and oil leakage from the equipment. The original data is divided into a training set and a validation set in a ratio of 8:2, and the proportion of each category is kept consistent during the division process.
[0072] To objectively evaluate the performance of the network model, the present invention uses Precision (P), Recall (R), mean average precision (mAP), the number of model parameters (Params), the amount of computation (GFLOPs), Frames Per Second (FPS), and the model size (Size) as evaluation metrics. The calculation formulas for each evaluation metric are as follows:
[0073] (1)
[0074] (2)
[0075] (3)
[0076] (4)
[0077] (5)
[0078] Among them, TP, FP, and FN represent the numbers of true positives, false positives, and false negatives correctly detected, respectively; AP represents the area enclosed by the P-R curve and the coordinate axes; N represents the number of categories of targets to be detected; Framenum represents the total number of detected pictures; and ElapsedTime represents the total time required for detection.
[0079] Example 2
[0080] According to the above-mentioned method for detecting substation equipment defects based on the improved YOLOv8n model, defect detection experiments have been carried out on the same dataset and experimental environment to verify the improvement effect of the IRB-SE attention mechanism, P2 small target detection layer, and BiFPN bidirectional feature pyramid network introduced in the present invention on the performance of the original YOLOv8n model.
[0081] Based on the original model, first, the IRB-SE attention mechanism designed in the present invention is added to form model1; then, after introducing the P2 small target detection layer, model2 is formed; finally, the BiFPN bidirectional feature pyramid network is introduced to form model3. Comparative verification is carried out through the designed ablation experiment, and the results are shown in Table 1.
[0082] Table 1 Ablation Experiment
[0083] As can be seen from Table 1, after using the IRB-SE attention mechanism designed in the present invention to improve the original C2f module on the original model, the detection accuracy of the model is improved by 1.5%, and the computational cost, number of parameters, and model size all decrease. It has the advantage of lightweight while improving performance. At this time, after introducing the P2 small target detection layer on the original model, the detection accuracy of the model is improved by 0.4% again, and the number of parameters and size of the model are reduced by 14% and 8.9% respectively, but the computational cost increases by 8.7 GFLOPs. After further introducing BiFPN, although mAP50 decreases by 0.1%, the computational cost, number of parameters, and model size of the model are reduced by 32.4%, 16.7%, and 13.7% respectively, achieving a lightweight effect on the model while ensuring the model accuracy. Comparing the improved model with the original model, mAP50 is improved by 1.8%, the number of parameters and model size are reduced by 33.3% and 26.7% respectively, but the computational cost is increased by 30%.
[0084] In order to illustrate that the improved IRB-SE attention mechanism can improve the performance of the model more effectively than the original SENetv2, the present invention designs a comparative experiment between the two. The results are shown in Table 2. As can be seen from the table, compared with SENetv2, the IRB-SE attention mechanism improves mAP50 by 0.5%, and improves the amount of computation, the amount of parameters, and the model size, which decrease by 13.6%, 6.7%, and 6.7%, respectively. This shows that the improved IRB-SE attention mechanism can improve the performance of the model more effectively than the original SENetv2.
[0085] Table 2 Comparative experiments of different attention mechanisms
[0086]
[0087] In order to verify the advancedness of the improved model, the improved model was compared with a variety of currently popular detection algorithms under the condition that the experimental environment and parameters remain unchanged. The experimental results are shown in Table 3. It can be observed that the precision, recall and mAP50 of the improved model proposed in the present invention reached 77.7%, 64.2% and 65.6% respectively, and the mAP50 is significantly better than other popular detection algorithms. At the same time, under the premise of maintaining a low amount of calculation, parameter amount and model size, the model can still maintain a high real-time performance.
[0088] Table 3 Comparative experiments of different algorithms
[0089]
[0090] It should be emphasized that the embodiments described in the present invention are illustrative rather than restrictive. Therefore, the present invention includes but is not limited to the embodiments described in the specific implementation manner. Any other implementation manners derived by those skilled in the art based on the technical solution of the present invention also fall within the scope of protection of the present invention.
Claims
1. A defect detection method for substation equipment based on an improved YOLOv8n model, characterized in that: The following steps are involved: Step 1, by improving the YOLOv8n model, an improved YOLOv8n model is obtained; Step 2: Use the improved YOLOv8n model to detect the substation equipment; Step 3: Obtain defect information of substation equipment based on the calculation results of the improved YOLOv8n model.
2. According to claim 1, a method for detecting defects in substation equipment based on an improved YOLOv8n model is characterized in that: The step 1 comprises the following steps: Step 1.1, combine SENetv2 and IRB, build the IRB-SE attention mechanism, and add it to the YOLOv8n model; Step 1.2, introduce the bidirectional feature pyramid network BiFPN into the YOLOv8n model; Step 1.3: Add a small target detection layer to the YOLOv8n model detection layer.
3. According to claim 2, a method for detecting defects in substation equipment based on an improved YOLOv8n model is characterized in that: The specific implementation method of step 1.1 is as follows: SENetv2 is integrated with IRB to obtain the IRB-SE attention mechanism, which is used to capture the advantages of feature information at different levels and scales, while allowing the network to effectively learn the complex relationship between channels and enhance the network's selective transmission capability for key features; and the IRB-SE attention mechanism is introduced into the C2f module of YOLOv8, and improved to obtain the IRB-SEC2f module.
4. According to claim 3, a method for detecting defects in substation equipment based on an improved YOLOv8n model is characterized in that: The specific implementation method of obtaining the IRB-SEC2f module is: adding an IRB-SE attention mechanism at the end of the Bottleneck module in the C2f module.
5. According to claim 2, a method for detecting defects in substation equipment based on an improved YOLOv8n model is characterized in that: The specific implementation method of step 1.2 is: feature fusion of the outputs of the second and third IRB-SEC2f modules from top to bottom in the backbone network and the outputs of the first and second IRB-SEC2f modules from bottom to top in the neck network in the fourth and fifth Concat modules in the neck network.
6. According to claim 2, a method for detecting defects in substation equipment based on an improved YOLOv8n model is characterized in that: The specific implementation method of step 1.3 is as follows: on the basis of the three detection layers of the YOLOv8n model, a small target detection layer is introduced, which is specifically used for features with high details and small size characteristics; it is used to capture and retain the features of small targets to improve the model's ability and effect in small target detection, reduce false detections and missed detections caused by feature loss, and in order to reduce the amount of calculation of the model and improve the detection speed, the last large target detection layer is deleted.
7. According to claim 1, a method for detecting defects in substation equipment based on an improved YOLOv8n model is characterized in that: The specific implementation method of step 2 is: Step 2.1: First, convert the dataset into YOLO format and divide it into training set and validation set in a ratio of 8:2 for use and verification during training; Step 2.2, introduce the improved method into the model to generate the corresponding network structure; Step 2.3: Build a suitable training environment and find the optimal parameter combination through multiple experiments; Step 2.4: Start training the model. The model will calculate the prediction results through forward propagation based on the input image data and annotation information, and compare them with the actual labels. Then, the model parameters will be adjusted using the back propagation algorithm to minimize the loss function and make the model's prediction results closer to the actual situation. Step 2.5: During the training process, use the validation set to regularly evaluate the trained model. Step 2.6: Use the trained model to detect the specified image and obtain the detection result.
Citation Information
Cited By
Underwater target identification method based on improved YOLOv8 algorithm
CN120356084A