A coal mine power equipment defect detection method based on improved YOLOv5s

By improving the YOLOv5s detection algorithm, the multi-branch coordinate attention module, feature fusion network module and fast spatial pyramid pooling average pooling module are added, which solves the problems of low accuracy of defect detection and difficulty in positioning of coal mine power equipment, and achieves more efficient and accurate defect detection.

CN116029982BActive Publication Date: 2025-05-23XUZHOU NORMAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211540752.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-02
Publication Date
2025-05-23
Estimated Expiration
2042-12-02

AI Technical Summary

Technical Problem

The detection accuracy of coal mine power equipment defects is low and the positioning is difficult, which affects production safety and efficiency.

Method used

Improve the YOLOv5s detection algorithm, add multi-branch coordinate attention module, feature fusion network module and fast spatial pyramid pooling average pooling module, enhance object detection accuracy and feature expression capabilities, and improve defect positioning accuracy.

Benefits of technology

It significantly improves the accuracy and positioning accuracy of coal mine power equipment defect detection, improves detection efficiency and accuracy, and meets the needs of coal mine production safety and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116029982B_ABST
    Figure CN116029982B_ABST
Patent Text Reader

Abstract

The present invention provides a method for detecting defects of coal mine power equipment based on improved YOLOv5s. It belongs to the technical field of target detection of computer vision. The coal mine power equipment image is input into the improved YOLOv5s network model for fault identification; the improved YOLOv5s network model uses a multi-branch coordinate attention module to enhance the ability of identifying coal mine power equipment faults in the coal mine power equipment image, and the coal mine power equipment faults include four fault conditions: oil leakage, dial damage, shell damage, and respirator silicone discoloration; the feature fusion network module is used to cross-layer connect the non-adjacent feature information between the backbone network and the neck network to enhance the feature expression and fusion capabilities; the fast spatial pyramid pooling average pooling module is embedded between the path fusion networks of the neck network to enhance the ability of the network to transmit shallow positioning information to the deep layer; and the result of the coal mine power equipment image detection is output. The steps are simple, the use is convenient, and the detection efficiency is high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision target detection technology, and in particular to a coal mine power equipment defect detection method based on improved YOLOv5s. Background Art

[0002] Coal is an important strategic energy source in my country and has promoted the overall economic development of my country. However, due to the harsh environment and complex working conditions of coal mines, some power equipment has been in a harsh environment for a long time. Affected by factors such as high temperature, severe cold, and air pollution, power equipment is prone to aging or damage. In order to ensure the production safety and efficiency of coal mines, it is necessary to conduct regular inspections on the operating status of coal mine power equipment. This type of inspection has important engineering significance and practical value.

[0003] Traditional coal mine power equipment inspection usually adopts manual inspection, but due to the characteristics of high labor intensity, scattered inspection quality and high difficulty, the inspection efficiency is low. With the rapid development of artificial intelligence, my country's coal mine intelligent detection technology continues to improve. This type of technology mainly uses intelligent robots to replace manual labor to achieve automatic inspection. At present, the defect detection of coal mine power equipment based on automatic inspection robots mainly involves methods such as image processing, machine learning and deep learning. Compared with the former two, the detection method based on deep learning has the advantages of high accuracy and strong real-time performance, and also has good advantages in model deployment in engineering practice.

[0004] Object detection algorithms based on deep learning are mainly divided into one-stage detection algorithms and two-stage detection algorithms. One-stage detection algorithms mainly include methods such as SSD and YOLO. This type of method uses the idea of ​​regression to predict all categories and corresponding confidence and bounding box information, and performs end-to-end object recognition, so it has certain advantages in detection speed. Two-stage detection algorithms mainly include methods such as R-CNN, FastR-CNN and FasterR-CNN. This type of method can obtain slightly higher detection accuracy by sliding the filter window on the image to extract the region of interest; but compared with the one-stage detection algorithm, the detection time is longer and it is not suitable for places with high real-time requirements. Among the existing one-stage detection algorithms, the YOLO series has shown better performance and can provide better detection results, so this detection algorithm has been widely used in engineering practice. Summary of the invention

[0005] Aiming at the problem of low accuracy of defects in coal mine power equipment, a coal mine power equipment defect detection method based on improved YOLOv5s is provided. The YOLOv5s detection algorithm is used, which includes a multi-branch coordinate attention module to improve the target detection accuracy; secondly, it includes a feature fusion network module to enable non-adjacent feature information to communicate effectively, thereby enhancing the model feature expression and fusion capabilities; finally, in order to solve the problem of difficulty in locating defects in coal mine power equipment, a fast spatial pyramid pooling and average pooling module is included to make the model pay more attention to the location of equipment defects.

[0006] To solve the above problems, the present invention provides a method for detecting defects in coal mine power equipment based on improved YOLOv5s, and the steps are as follows:

[0007] 1) Perform digitization processing on coal mine power equipment images and preprocess coal mine power equipment images;

[0008] 2) Input the preprocessed coal mine power equipment images into the improved YOLOv5s network model for fault identification;

[0009] 3) Using the multi-branch coordinate attention module, the improved YOLOv5s model is used to enhance the ability to identify coal mine power equipment faults in coal mine power equipment images. Coal mine power equipment faults include: oil leakage, dial damage, shell damage, and respirator silicone discoloration;

[0010] 4) The feature fusion network module is used to connect non-adjacent feature information between the backbone network and the neck network across layers to enhance the feature expression and fusion capabilities of the improved YOLOv5s model; the fast spatial pyramid pooling average pooling module is embedded between the path fusion networks of the neck network to enhance the ability of the network to transfer shallow positioning information to deep layers;

[0011] 5) The improved YOLOv5s network model outputs the probabilities of five states, namely, oil leakage, damaged dial, damaged shell, discoloration of respirator silicone, and no fault, based on the results of coal mine power equipment image detection, and outputs the result with the highest probability; thereby determining whether the coal mine power equipment has a fault, and if so, which of the following faults is it: oil leakage, damaged dial, damaged shell, or discoloration of respirator silicone;

[0012] The improved YOLOv5s network model is based on the conventional YOLOv5s network model, and a multi-branch coordinate attention module MCA is added to the trunk of the YOLOv5s network model, and sequentially connected CBS modules, CBS modules, C3 modules, CBS modules, C3 modules, CBS modules, C3 modules, CBS modules, C3 modules, MCA coordinate attention modules, and fast spatial pyramid pooling modules SPPF modules are constructed as the trunk of the improved YOLOv5s network model; the CBS module includes: a convolutional layer, a batch normalization layer, and a SiLU activation function;

[0013] In the YOLOv5s network model, the third C3 module and the fourth C3 module are directly transmitted to the PAN layer of the neck by crossing the FPN layer of the neck, which ensures the effective communication of feature information between non-adjacent layers, and further reuses the key feature information to achieve information fusion between multiple feature layers and construct a feature fusion module:

[0014] In the neck of the YOLOv5s network model, add the fast spatial pyramid pooling average pooling module SPPFA after the first C3 and the second C3 to construct the neck part;

[0015] The backbone part of the YOLOv5s network model, the neck part of the YOLOv5s network model, and the head part of the YOLOv5s network model are connected in sequence. The output end of the backbone part of the YOLOv5s network model is connected to the input end of the neck part of the YOLOv5s network model. The output end of the neck part of the YOLOv5s network model is connected to the input end of the output part of the YOLOv5s network model. The feature fusion module connects the non-adjacent feature information between the backbone network and the neck network across layers to enhance the feature expression and fusion capabilities of the YOLOv5s model.

[0016] Collect defective sample images of coal mine power equipment and normal images of coal mine power equipment, manually mark whether the coal mine power equipment in the image is faulty and the type of fault, and then train the improved YOLOv5s network model with the manually labeled images until the training requirements are met to obtain an improved YOLOv5s network model with good catenary.

[0017] Furthermore, the MCA module includes sequentially connected Y channel average pooling, X channel average pooling, global average pooling, convolution 1, convolution 2, convolution 3, non-linear activation function 1, non-linear activation function 2, and non-linear activation function 3;

[0018] The three input ends of the MCA module are Y channel average pooling, X channel average pooling, and global average pooling; the output end of the Y channel average pooling is connected to the input end of convolution 1, and the output end of convolution 1 is connected to the input end of nonlinear activation function 1; the output end of the X channel average pooling is connected to the input end of convolution 2, and the output end of convolution 2 is connected to the input end of nonlinear activation function 2; the output end of the global average pooling is connected to the input end of convolution 3, and the output end of convolution 3 is connected to the input end of nonlinear activation function 3; the normalized weight g finally generated by nonlinear activation function 1, nonlinear activation function 2, and nonlinear activation function 3 h , g w , g hw Multiply them together as the output of the MCA coordinate attention module.

[0019] Further, the feature fusion module includes 20×20 feature layer 1, 20×20 feature layer 2, 20×20 feature layer 3, 20×20 feature layer 4, 40×40 feature layer 1, 40×40 feature layer 2, 40×40 feature layer 3, 80×80 feature layer 1, 80×80 feature layer 2, and 80×80 feature layer 3;

[0020] The 80×80 feature layer 1 is used as the input end, the output end of the 80×80 feature layer 1 is connected to the input ends of the 40×40 feature layer 1 and the 80×80 feature layer 2; the output end of the 40×40 feature layer 1 is connected to the input ends of the 20×20 feature layer 1 and the 40×40 feature layer 2; the output end of the 20×20 feature layer 1 is connected to the input end of the 20×20 feature layer 2; the output end of the 20×20 feature layer 2 is connected to the input end of the 20×20 feature layer 3; the output end of the 20×20 feature layer 3 is connected to the input ends of the 40×40 feature layer 2 and the 20×20 feature layer 4; the output end of the 40×40 feature layer 2 is connected to the 8 The input ends of the 0×80 feature layer 2 and the 40×40 feature layer 3 are connected; the output end of the 80×80 feature layer 2 is connected to the input end of the 80×80 feature layer 3; the output end of the 80×80 feature layer 3 is connected to the input end of the 40×40 feature layer 3; the output end of the 40×40 feature layer 3 is connected to the input end of the 20×20 feature layer 4; the output ends of the 40×40 feature layer 1 and the 20×20 feature layer 1 are connected to the input ends of the 20×20 feature layer 4 and the 40×40 feature layer 3; the output ends of the 20×20 feature layer 4, the 40×40 feature layer 3, and the 80×80 feature layer 3 serve as the output ends of the feature fusion module.

[0021] Further, the SPPFA module includes sequentially connected ConvBNSiLU1, ConvBNSiLU2, average pooling 1, average pooling 2, average pooling 3, and Concat modules;

[0022] ConvBNSiLU1 is used as the input, and the output of ConvBNSiLU1 is connected to the input of average pooling 1 and the input of the Concat module; the output of average pooling 1 is connected to the input of average pooling 2, and the output of average pooling 2 is connected to the input of average pooling 3; the outputs of average pooling 1, average pooling 2, and average pooling 3 are connected to the input of the Concat module; the output of the Concat module is connected to the input of ConvBNSiLU2; the output of ConvBNSiLU2 serves as the output of the SPPFA module.

[0023] Furthermore, the improved YOLOv5s model after training was subjected to performance evaluation tests. The evaluation indicators included precision P, recall R, average precision AP, mean average precision mAP, and weighted average F of precision P and recall R. 1 The specific formula is as follows:

[0024]

[0025]

[0026]

[0027]

[0028]

[0029] Where P represents the precision, that is, the ratio of the number of correctly detected coal mine power equipment defects to the number of detected coal mine power equipment defects; R represents the recall rate, that is, the ratio of the number of correctly detected coal mine power equipment defects to the number of coal mine power equipment defects in the data set; T P Indicates the number of correctly detected coal mine power equipment defects, F P represents the number of falsely detected coal mine power equipment defects, F N It represents the number of falsely detected defects of non-coal mine power equipment; n is the total number of categories, and i represents the serial number of the current category.

[0030] Furthermore, the manually annotated coal mine power equipment defect images and fault-free images are combined into an image dataset and divided into a training set and a validation set in a ratio of 8:2.

[0031] Furthermore, all images in the training set and validation set are unified into a preset size.

[0032] Furthermore, Mosaic data enhancement is performed on the training set images.

[0033] Beneficial effects: The present invention proposes a method for detecting defects in coal mine power equipment in view of the characteristics of underground coal mines. The improved YOLOv5s model is used in the detection method to improve the accuracy of defect detection in coal mine power equipment; 2. The improved YOLOv5s model enhances the ability of the YOLOv5s model to obtain equipment oil leakage, dial damage, shell damage, and respirator silicone discoloration area information by utilizing a multi-branch coordinate attention module; 3. The improved YOLOv5s model enhances the feature expression and fusion ability of the YOLOv5s model by utilizing a feature fusion network module to cross-layer connect non-adjacent feature information between the backbone network and the neck network; 3. The fast spatial pyramid pooling average pooling module is embedded between the path fusion networks of the neck network of the YOLOv5s model to improve the ability of the network to transmit shallow positioning information to deep layers. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 A specific network diagram of the coal mine power equipment defect detection method of improving YOLOv5s according to an embodiment of the present invention;

[0035] Figure 2 It is an MCA module diagram of the coal mine power equipment defect detection method of improving YOLOv5s in an embodiment of the present invention;

[0036] Figure 3 It is a feature fusion module diagram of the coal mine power equipment defect detection method of improving YOLOv5s in an embodiment of the present invention;

[0037] Figure 4 This is a SPPFA module diagram of the coal mine power equipment defect detection method based on improved YOLOv5s according to an embodiment of the present invention. DETAILED DESCRIPTION

[0038] In order to better understand the technical content of the present invention, specific embodiments are given and described as follows in conjunction with the accompanying drawings.

[0039] Various aspects of the invention are described herein with reference to the accompanying drawings, in which many illustrative embodiments are shown. The embodiments of the invention are not limited to those described in the accompanying drawings. The invention is implemented by any of the various concepts and embodiments described above, as well as the concepts and embodiments described in detail below, because the concepts and embodiments disclosed in the invention are not limited to any implementation. In addition, some aspects disclosed in the invention may be used alone or in any appropriate combination with other aspects disclosed in the invention.

[0040] The present invention provides a coal mine power equipment defect detection method based on improved YOLOv5s, based on a coal mine power equipment defect data set containing corresponding preset labels. The coal mine power equipment defect data contains 4 types of pictures, a total of 3250 pictures, namely equipment oil leakage, dial damage, meter housing damage, and respirator silicone discoloration. The labels of the targets in the pictures are marked using Labelimg software. The data set generates a training set and a validation set in a ratio of 8:2.

[0041] After that, all images in the training set and validation set are unified into a size of 640×640. MCA attention module, feature fusion module and SPPFA module are added to the baseline YOLOv5s network model. Based on the preset number of coal mine power equipment defect sample images in the coal mine power equipment defect dataset, each coal mine power equipment defect sample image is used as input and each target equipment defect sample image is used as output to train the improved YOLOv5s coal mine power equipment defect detection model, and obtain the improved YOLOv5s coal mine power equipment defect detection model.

[0042] like Figure 1 As shown in the figure, in the improved YOLOv5s model, a multi-branch coordinate attention module is proposed to improve the target detection accuracy; secondly, through the feature fusion network module, non-adjacent feature information is effectively communicated, which enhances the model's feature expression and fusion capabilities; finally, in view of the difficulty in locating defects in coal mine power equipment, a fast spatial pyramid pooling average pooling module is proposed to make the model pay more attention to the location of equipment defects. Through steps one to four, the coal mine power equipment defect detection model of the improved YOLOv5s is constructed;

[0043] Step 1: Based on the YOLOv5s network model, add a multi-branched coordinate attention module (MCA) to the backbone of the YOLOv5s network model, and sequentially connect the CBS module, CBS module, C3 module, CBS module, C3 module, CBS module, C3 module, CBS module, C3 module, CBS module, C3 module, MCA coordinate attention module, and SPPF module to construct the backbone of the improved YOLOv5s network model;

[0044] Step 2: Based on the YOLOv5s network model, in the feature fusion module of the YOLOv5s network model, the third and fourth C3 modules are directly transferred to the PAN layer of the neck through the FPN layer across the neck;

[0045] Step 3: Based on the YOLOv5s network model, add a Spatial Pyramid Pooling Fast Avgpooling (SPPFA) module to the neck of the YOLOv5s network model to construct the improved neck part of the YOLOv5s network model;

[0046] Step 4: Sequentially connect the backbone part of the improved YOLOv5s network model, the neck part of the improved YOLOv5s network model, and the output part of the YOLOv5s network model. The output end of the backbone part of the improved YOLOv5s network model is connected to the input end of the neck of the YOLOv5s network model, and the output end of the neck part of the improved YOLOv5s network model is connected to the input end of the output part of the YOLOv5s network model; to form an improved YOLOv5s coal mine power equipment defect detection model with the coal mine power equipment defect sample image as the input and the true value label corresponding to the coal mine power equipment defect sample image as the output;

[0047] The MCA coordinate attention module is as Figure 2 shown, consisting of three branches in total. The average pooling of the Y channel, the average pooling of the X channel, and the global average pooling are used as the three input ends of the MCA module; the output end of the average pooling of the Y channel is connected to the input end of Convolution 1, and the output end of Convolution 1 is connected to the input end of Nonlinear Activation Function 1; the output end of the average pooling of the X channel is connected to the input end of Convolution 2, and the output end of Convolution 2 is connected to the input end of Nonlinear Activation Function 2; the output end of the global average pooling is connected to the input end of Convolution 3, and the output end of Convolution 3 is connected to the input end of Nonlinear Activation Function 3; the normalized weights g h 、g w 、g hw generated by Nonlinear Activation Function 1, Nonlinear Activation Function 2, and Nonlinear Activation Function 3 are multiplied as the output end of the MCA coordinate attention module.

[0048] In the MCA coordinate attention, the channels of Branch 1 and Branch 2 are decomposed into two parallel one-dimensional feature codes, and then the input feature image is averaged pooled along the width and height directions respectively, and the normalized weights g h ,g w are generated through convolution and nonlinear activation functions. In Branch 3, the global information is combined through global average pooling to capture wider and higher image features in the receptive field, and the normalized weight g hwThe coordinate attention module can be viewed as a computational unit that enhances the expressiveness of learned features in mobile networks. It can take any intermediate feature tensor as input and output a transformed tensor with the same size as the input. The attention module can effectively make the model focus on more interesting areas, thereby improving the model's defect detection capabilities.

[0049] Feature fusion module such as Figure 3 As shown. The feature fusion module ensures the effective exchange of feature information between non-adjacent layers, and further reuses the key feature information to achieve information fusion between multiple feature layers. The feature fusion module includes 20×20 feature layer 1, 20×20 feature layer 2, 20×20 feature layer 3, 20×20 feature layer 4, 40×40 feature layer 1, 40×40 feature layer 2, 40×40 feature layer 3, 80×80 feature layer 1, 80×80 feature layer 2, 80×80 feature layer 3;

[0050] The 80×80 feature layer 1 is used as the input end, the output end of the 80×80 feature layer 1 is connected to the input ends of the 40×40 feature layer 1 and the 80×80 feature layer 2; the output end of the 40×40 feature layer 1 is connected to the input ends of the 20×20 feature layer 1 and the 40×40 feature layer 2; the output end of the 20×20 feature layer 1 is connected to the input end of the 20×20 feature layer 2; the output end of the 20×20 feature layer 2 is connected to the input end of the 20×20 feature layer 3; the output end of the 20×20 feature layer 3 is connected to the input ends of the 40×40 feature layer 2 and the 20×20 feature layer 4; the output end of the 40×40 feature layer 2 is connected to the 8 The input ends of the 0×80 feature layer 2 and the 40×40 feature layer 3 are connected; the output end of the 80×80 feature layer 2 is connected to the input end of the 80×80 feature layer 3; the output end of the 80×80 feature layer 3 is connected to the input end of the 40×40 feature layer 3; the output end of the 40×40 feature layer 3 is connected to the input end of the 20×20 feature layer 4; the output ends of the 40×40 feature layer 1 and the 20×20 feature layer 1 are connected to the input ends of the 20×20 feature layer 4 and the 40×40 feature layer 3; the output ends of the 20×20 feature layer 4, the 40×40 feature layer 3, and the 80×80 feature layer 3 serve as the output ends of the feature fusion module.

[0051] SPPFA module such as Figure 4 As shown, the module includes CBS1, CBS2, average pooling 1, average pooling 2, average pooling 3, and Concat modules;

[0052] CBS1 is used as the input end, and the output end of CBS1 is connected to the input end of average pooling 1 and the input end of the Concat module; the output end of average pooling 1 is connected to the input end of average pooling 2, and the output end of average pooling 2 is connected to the input end of average pooling 3; the output ends of average pooling 1, average pooling 2, and average pooling 3 are connected to the input end of the Concat module; the output end of the Concat module is connected to the input end of CBS2; the output end of CBS2 serves as the output end of the SPPFA module.

[0053] The average pooling layer is used in deep networks, which can often retain the characteristics of the overall data, is more sensitive to background information, and can effectively obtain global context information. The PAN layer uses downsampling to transfer shallow positioning information to the deep layer by fusing bottom-up feature information, thereby enhancing multi-scale positioning capabilities. Therefore, placing the SPPFA module between PAN layers can effectively focus on the location of defects in coal mine power equipment.

[0054] The training set is sent to the improved YOLOv5s coal mine power equipment defect detection model for training. The experimental environment of the present invention uses Windows 11 operating system, 12th Gen Intel Core i5-12400 2.50 GHz processor, NVIDIA GeForce RTX 3060 graphics card. The deep learning framework uses Pytorch 1.8 version and CUDA 11.1 version. The training parameters have an initial learning rate of 0.01, a momentum value of 0.937, a bitchsize of 32, and a maximum number of iterations of 200.

[0055] The validation set is input into the trained improved YOLOv5s coal mine power equipment defect detection model for performance analysis and experimental results. The evaluation indicators include precision, recall, mean average precision (mAP), F 1 The specific formula is as follows:

[0056]

[0057]

[0058]

[0059]

[0060]

[0061] Where P represents the detection accuracy, that is, the ratio of the number of correctly detected coal mine power equipment defects to the number of detected coal mine power equipment defects. R represents the recall rate, that is, the ratio of the number of correctly detected coal mine power equipment defects to the number of coal mine power equipment defects in the data set. P Indicates the number of correctly detected coal mine power equipment defects, F P represents the number of falsely detected coal mine power equipment defects, F N Indicates the number of falsely detected non-coal mine power equipment defects. n is the total number of categories, and i is the sequence number of the current category.

[0062] Aiming at the problem of low defect detection of coal mine power equipment, experiments have proved that the detection accuracy can be greatly improved. The data test comparison is shown in Table 1, which more fully proves the feasibility of the algorithm of the present invention. The innovation of the present invention is reflected in the proposal of a multi-branch coordinate attention module, a feature fusion network module, and a fast spatial pyramid pooling average pooling module, which are used in the YOLOv5s detection algorithm. Compared with the original baseline algorithm, the defect detection performance of coal mine power equipment is improved.

[0063] In order to verify the influence of each module on the network performance, the present invention designed an ablation experiment, as shown in Table 1. Model 1 is a single MCA coordinate attention module; Model 2 is a single feature fusion network module; Model 3 is a single SPPFA module; Models 4 to 6 are modules that are improved and combined in pairs from Models 1 to 3; Model 7 is an overall module after the three improved modules are combined. It can be seen from Table 1 that the indicators of Models 1, 2, 3, 4, 5, and 6 have been improved to varying degrees compared with the original YOLOv5s model. Model 7 is better than the original model in terms of P, R, mAP@0.5, mAP@0.95, and F 1 The scores were increased by 1.8%, 2.2%, 3.1%, 2.3% and 3% respectively. All indicators have been greatly improved, which fully proves the effectiveness of the improvements made by the present invention.

[0064] Table 1

[0065]

[0066] The above is only a preferred embodiment of the present invention, but it is not intended to limit the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention, and these improvements and modifications should also be regarded as the protection scope of the present invention.

Claims

1. A coal mine power equipment defect detection method based on improved YOLOv5s, Features Here are the steps: 1) Perform digitization processing on coal mine power equipment images and pre-process coal mine power equipment images; 2) Input the preprocessed coal mine power equipment image into the improved YOLOv5s network model output. According to the results of coal mine power equipment image detection, the probabilities of five states, namely, oil leakage, dial damage, shell damage, respirator silicone discoloration, and no fault, are specifically generated, and the result with the highest probability is output; thereby judging whether the coal mine power equipment has a fault, and if a fault occurs, which of the faults is oil leakage, dial damage, shell damage, and respirator silicone discoloration; The improved YOLOv5s network model is based on the conventional YOLOv5s network model, and a multi-branch coordinate attention module MCA is added before the SPPF module of the trunk of the YOLOv5s network model. The MCA module includes sequentially connected Y channel average pooling, X channel average pooling, global average pooling, convolution 1, convolution 2, convolution 3, nonlinear activation function 1, nonlinear activation function 2, and nonlinear activation function 3; The three input ends of the MCA module are Y channel average pooling, X channel average pooling, and global average pooling; the output end of the Y channel average pooling is connected to the input end of convolution 1, and the output end of convolution 1 is connected to the input end of nonlinear activation function 1; the output end of the X channel average pooling is connected to the input end of convolution 2, and the output end of convolution 2 is connected to the input end of nonlinear activation function 2; the output end of the global average pooling is connected to the input end of convolution 3, and the output end of convolution 3 is connected to the input end of nonlinear activation function 3; the normalized weight g finally generated by nonlinear activation function 1, nonlinear activation function 2, and nonlinear activation function 3 h , g w , g hw Multiply as the output of the MCA coordinate attention module; In the YOLOv5s network model, the outputs of the 3rd C3 module and the 4th C3 module of the backbone network are directly transmitted to the PAN layer of the neck by crossing the FPN layer of the neck, realizing the information fusion between multiple feature layers and constructing a feature fusion module. In the PAN layer of the neck of the YOLOv5s network model, a fast spatial pyramid pooling average pooling module SPPFA is added after the first C3 and the second C3 to construct the neck part; the SPPFA module is obtained by replacing the maximum pooling in the SPPF module with average pooling. Collect defective sample images of coal mine power equipment and normal images of coal mine power equipment, manually mark whether the coal mine power equipment in the image is faulty and the type of fault, and then train the improved YOLOv5s network model with the manually labeled images until the training requirements are met to obtain a trained improved YOLOv5s network model.

2. According to claim 1, a method for detecting defects in coal mine power equipment based on improved YOLOv5s, It is characterized in that The feature fusion module includes 20×20 feature layer 1, 20×20 feature layer 2, 20×20 feature layer 3, 20×20 feature layer 4, 40×40 feature layer 1, 40×40 feature layer 2, 40×40 feature layer 3, 80×80 feature layer 1, 80×80 feature layer 2, and 80×80 feature layer 3; Taking the 80×80 feature layer 1 as the input end, the output end of the 80×80 feature layer 1 is connected to the input ends of the 40×40 feature layer 1 and the 80×80 feature layer 2; the output end of the 40×40 feature layer 1 is connected to the input ends of the 20×20 feature layer 1 and the 40×40 feature layer 2; the output end of the 20×20 feature layer 1 is connected to the input end of the 20×20 feature layer 2; the output end of the 20×20 feature layer 2 is connected to the input end of the 20×20 feature layer 3; the output end of the 20×20 feature layer 3 is connected to the input ends of the 40×40 feature layer 2 and the 20×20 feature layer 4; the output end of the 40×40 feature layer 2 is connected to the input ends of the 80×80 feature layer 2 and the 40×40 feature layer 3; the output end of the 80×80 feature layer 2 is connected to the input end of the 80×80 feature layer 3; the output end of the 80×80 feature layer 3 is connected to the input end of the 40×40 feature layer 3; the output end of the 40×40 feature layer 3 is connected to the input end of the 20×20 feature layer 4; the output ends of the 40×40 feature layer 1 and the 20×20 feature layer 1 are connected to the input ends of the 20×20 feature layer 4 and the 40×40 feature layer 3; the output ends of the 20×20 feature layer 4, the 40×40 feature layer 3, and the 80×80 feature layer 3 are used as the output end of the feature fusion module.

3. A method for detecting defects of coal mine power equipment based on improved YOLOv5s according to claim 1, characterized in that, the SPPFA module includes ConvBNSiLU1, ConvBNSiLU2, average pooling 1, average pooling 2, average pooling 3, and Concat module connected in sequence; taking ConvBNSiLU1 as the input end, the output end of ConvBNSiLU1 is connected to the input ends of average pooling 1 and the Concat module; the output end of average pooling 1 is connected to the input end of average pooling 2, and the output end of average pooling 2 is connected to the input end of average pooling 3; the output ends of average pooling 1, average pooling 2, and average pooling 3 are connected to the input end of the Concat module; the output end of the Concat module is connected to the input end of ConvBNSiLU2; the output end of ConvBNSiLU2 is used as the output end of the SPPFA module.

4. A method for detecting defects of coal mine power equipment based on improved YOLOv5s according to claim 1, characterized in that, The improved YOLOv5s model after training is tested for performance evaluation. The evaluation indicators include precision P, recall R, average precision AP, average precision mAP, and weighted average F of precision P and recall R. 1 The specific formula is as follows: Where P represents the precision, that is, the ratio of the number of correctly detected coal mine power equipment defects to the number of detected coal mine power equipment defects; R represents the recall rate, that is, the ratio of the number of correctly detected coal mine power equipment defects to the number of coal mine power equipment defects in the data set; T P Indicates the number of correctly detected coal mine power equipment defects, F P represents the number of falsely detected coal mine power equipment defects, F N It represents the number of falsely detected defects of non-coal mine power equipment; n is the total number of categories, and i represents the serial number of the current category.

5. A method for detecting defects of coal mine power equipment based on improved YOLOv5s according to claim 1, characterized in that, The defect images of coal mine power equipment after manual annotation and the fault-free images are combined to form an image data set, which is divided into a training set and a validation set according to a ratio of 8:

2.

6. A method for detecting defects of coal mine power equipment based on improved YOLOv5s according to claim 5, characterized in that, All the images in the training set and the validation set are unified to a preset size.

7. A method for detecting defects of coal mine power equipment based on improved YOLOv5s according to claim 5, characterized in that, Perform Mosaic data enhancement on the training set images.