Sealing plug surface defect detection deep learning method based on YOLOv5
By adding attention mechanism SE module and CBAM module to the YOLOv5 network algorithm, the YOLOv5 model is optimized to improve detection accuracy and robustness, and the problem of poor detection effect of sealing and plugging surface defects in the prior art is solved, achieving more efficient and accurate detection performance.
Patent Information
- Application Number
- CN202510131166.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-06-10
AI Technical Summary
The existing Yolov5-based sealing surface defect detection method is not significantly improved when detecting fine and small sealing surface defects, making it difficult to meet the needs of modern manufacturing for high-precision and high-speed detection.
The attention mechanism SE module is added before the feature fusion of the last layer of the backbone network of the YOLOv5 network algorithm, and the attention mechanism CBAM module is added after the two upsamplings and fusions of the neck network to optimize the YOLOv5 model to improve detection accuracy and robustness.
It significantly improves the detection accuracy, robustness and efficiency of the YOLOv5 model, can better handle small defects, complex backgrounds and detection tasks of different scales, and improves the performance of sealing and plugging surface defect detection.
Smart Images

Figure CN120124697A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning defect detection, and in particular to a deep learning method for sealing plug surface defect detection based on YOLOv5. Background Art
[0002] As a key component in automotive wiring harnesses and connectors, the surface quality of precision small sealing plugs has an important impact on sealing performance and long-term reliability. In the actual production process, due to the limitations of manufacturing equipment precision, processing technology and environmental conditions, defects such as gas accumulation, glue deficiency, hole blockage, inner hole defects, and flash are prone to occur on the surface of the sealing plug. These surface defects may not only weaken its sealing performance, but also cause leakage or equipment failure under high pressure or high temperature conditions, posing a serious threat to the safety and stability of the product. Therefore, how to efficiently and accurately detect surface defects of sealing plugs, improve detection quality and optimize production efficiency has become an important issue that needs to be solved in the field of precision manufacturing.
[0003] Traditional manual inspection methods are time-consuming, inefficient, and rely on subjective judgment, often accompanied by a high risk of misjudgment and missed detection, and are difficult to meet the needs of modern manufacturing for high-precision and high-speed inspection. Therefore, the Yolo intelligent network algorithm is introduced to achieve efficient and intelligent inspection process, and has become a more popular inspection method in recent years, and has achieved remarkable results in various computer vision tasks. For example, the Chinese invention patent with announcement number CN118071673A discloses a tile surface defect detection method based on Yolov5, which improves the performance of the Yolov5 detection model by adding four attention modules, namely SE, CA, ECA and CBAM. However, it is only added to the backbone network, and is affected by the addition position and the algorithm of the added module. For small and delicate parts such as sealing plugs, the improvement in the inspection effect is not very obvious. Summary of the invention
[0004] In order to overcome the shortcomings of the background technology and solve the existing technical problems, the present invention discloses a deep learning method for sealing plug surface defect detection based on YOLOv5, which aims to achieve fast, accurate and high-speed detection of surface defects of precision small sealing plugs.
[0005] To achieve the above object, the present invention adopts the following technical solution:
[0006] A deep learning method for detecting surface defects of a sealing plug based on YOLOv5, comprising the following steps: S1. Optimize the YOLOv5 network algorithm; S1.1. Add an attention mechanism SE module before the last layer feature fusion of the backbone network of the YOLOv5 network algorithm; the attention mechanism SE module is used for adaptive feature extraction and correction of feature weights to improve the model's attention to defect targets; S1.2. Add an attention mechanism CBAM module after the two upsampling fusions of the neck network of the YOLOv5 network algorithm; the attention mechanism CBAM module combines the channel attention mechanism and the spatial attention mechanism to extract more comprehensive feature information in the defect area and expand the global receptive field of the feature extraction network.
[0007] S2. Collect different defect categories of the sealing plug and make a surface defect dataset of the sealing plug; S3. Train the optimized YOLOv5 network algorithm through the surface defect dataset of the sealing plug to obtain a YOLOv5 object detection network for detecting surface defects of the sealing plug.
[0008] Furthermore, in S1.1, the attention mechanism SE module includes a compression part and an excitation part. The compression part is used to compress the spatial information in the input feature map, compressing H×W×C into 1×1×C, that is, using global average pooling in the spatial dimension to obtain a 1×1×C feature map; the excitation part is used to combine the learned attention information with the input feature and obtain a feature map U with channel attention.
[0009] Furthermore, the compression operation output of the compression part is obtained through global average pooling, calculating the global spatial average value for each channel c to generate a channel description vector z∈R C : where z c represents the global description of the c-th channel; the excitation part passes the compressed channel description vector z through a two-layer fully connected network to generate a channel weight s∈R C : s = σ(W2·δ(W1·z)); where W 1 ∈R C / r×C and W 2 ∈R C×C / r are the weight matrices of the two-layer fully connected network; δ(·) is the ReLU activation function; σ(·) is the Sigmoid activation function; r is the channel scaling coefficient, and r = 16 can be taken to reduce the amount of calculation; the final output of the SE module is expressed as: Fc′ = s c ·F c , where Fc′ is the weighted c-th channel and s c is the corresponding channel weight.
[0010] Further, in S1.2, the attention mechanism CBAM module includes a channel attention module CAM and a spatial attention module SAM. The CAM module is used to weight the importance of each channel of the input feature map, compress the feature map in the spatial dimension through global average pooling and global max pooling, and obtain a global feature of 1×1×C. The SAM module is used to assign different weights to each spatial position in the feature map. By performing global average pooling and max pooling on the channel dimension, an H×W feature map is generated to express the importance of spatial positions. The input of CBAM is a feature map F∈R C×H×W , where C is the number of channels, and H and W are the height and width of the feature map respectively.
[0011] Further, the CAM module performs global average pooling and global max pooling on the input feature map F in the spatial dimension to generate two vectors describing the global features:
[0012]
[0013] favg and fmax respectively pass through two fully connected networks (MLP) with shared weights, are added and activated to generate the channel attention weight M c : M c =σ(MLP(AvgPool(F)) + MLP(MaxPool(F))); where σ is the Sigmoid activation function, MLP represents a multi-layer perceptron for learning the non-linear relationship between channels, F is the input feature map, and M c is the channel weight; the output of CAM is multiplied by the input feature map with the weight M c to obtain the channel-weighted feature map F′, and the input feature map is weighted according to the channel attention weight: F′ = M c ·F(4);
[0014] The SAM module performs global average pooling and global max pooling on the input feature map F′ in the channel dimension to generate two feature maps f′ avg and f′ max :
[0015]
[0016] Concatenate the f′ avg and f′ max feature maps, and pass through a 7×7 convolutional layer to generate the spatial attention weight M s : M s =σ(Conv7×7([AvgPool(F′),MaxPool(F′)])); where Conv7×7 represents a convolutional kernel of size 7×7, · represents the feature map concatenation operation, F′ is the channel-weighted feature map, and Ms is the spatial weight; the output of SAM is multiplied by the weight M s to multiply with the input feature map F′, obtaining the final feature map F″ with channel and spatial attention; the final output of the CBAM module can be expressed as: F″ = M s ·F′(7), where · represents the element-wise multiplication operation, M s and M c are the spatial weight and channel weight respectively, and F″ is the finally enhanced feature map.
[0017] Furthermore, in the S1 optimized YOLOv5 network algorithm, S1.3 is also included. All C3 modules in the backbone network and neck network of the YOLOv5 network algorithm are replaced with CF2 modules.
[0018] Furthermore, in the S1 optimized YOLOv5 network algorithm, S1.4 is also included. The Feature Pyramid Network BiFPN module is introduced into the neck network of the YOLOv5 network algorithm, and the original Path Aggregation Network PANet structure of the YOLOv5 network algorithm is further removed to enhance the information extraction ability of the network and further enhance the detection performance of the network for small target defects.
[0019] Furthermore, in the S1 optimized YOLOv5 network algorithm, S1.5 is also included. The P2 small target detection layer is added to the head network of the YOLOv5 network algorithm to enhance the detection accuracy of the algorithm for small targets.
[0020] Furthermore, in the S1 optimized YOLOv5 network algorithm, S1.6 is also included. The EIOU loss function is introduced into the YOLOv5 network algorithm to replace the original CIOU loss function, optimizing the sample imbalance problem in the bounding box regression task.
[0021] Furthermore, in S2, the sealed plug surface defect dataset includes collecting different defect categories of the sealed plug through a vision device to make the sealed plug dataset, then performing labeling processing on the sealed plug dataset through LabelImg, generating the completed annotation.txt file at the same time, and randomly dividing the sealed plug dataset and the corresponding.txt file into a training set, a validation set, and a test set according to the ratio of 8:1:1.
[0022] Due to the adoption of the above-mentioned technical solution, the present invention has the following beneficial effects:
[0023] The deep learning method for surface defect detection of sealed plugs based on YOLOv5 disclosed in the present invention adds optimization measures of SE and CBAM modules at specific positions in the backbone network and the neck network for the surface defect detection of small and precise components, which can significantly improve the detection accuracy, robustness and efficiency of the YOLOv5 model; the SE module enhances the attention to defect features by optimizing channel attention; the CBAM module improves the comprehensive perception ability of defect regions and the effect of global feature fusion by combining channel and spatial attention mechanisms. These optimizations help the model better handle the detection tasks of small and precise defects, complex backgrounds and different scales, thereby effectively improving the performance of surface defect detection of small and precise components; compared with the existing Yolov5 detection models with various attention modules added, the method of the present invention achieves an excellent balance in detection accuracy and recall rate, and the detection performance is significantly improved, and it performs better in the constructed sealed plug dataset. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 It is a flowchart of an implementation process of the present invention;
[0025] Figure 2 It is another flowchart of an implementation process of the present invention;
[0026] Figure 3 It is a schematic structural diagram of the attention mechanism SE module in the present invention;
[0027] Figure 4 It is a schematic structural diagram of the attention mechanism CBAM module in the present invention;
[0028] Figure 5 It is the backbone network structure with the attention mechanism SE module added in the YOLOv5 network algorithm;
[0029] Figure 6 It is the neck network structure with the attention mechanism CBAM module added in the YOLOv5 network algorithm;
[0030] Figure 7 It is a schematic comparison structure diagram of the CF2 and the original C3 modules in the present invention;
[0031] Figure 8 It is the network structure with the CF2 module replacing the C3 module in the YOLOv5 network algorithm;
[0032] Figure 9 It is a schematic comparison structure diagram of the feature pyramid network BiFPN module with the FPN module and the PANet module;
[0033] Figure 10 It is the network structure with the feature pyramid network BiFPN module added in the YOLOv5 network algorithm;
[0034] Figure 11 The network structure for adding a P2 small target detection layer to the YOLOv5 network algorithm;
[0035] Figure 12 The structural schematic diagram of the EIOU loss function;
[0036] Figure 13 The network structure for replacing the CIOU loss function with the EIOU loss function in the YOLOv5 network algorithm;
[0037] Figure 14 The network structure of the unimproved YOLOv5 network algorithm;
[0038] Figure 15 For Figure 2 The network structure of the YOLOv5 network algorithm improved according to the
[0039] Figure 16 The three-dimensional structural schematic diagram of the visual detection automation device;
[0040] Figure 17 The partial top-down structural schematic diagram of the visual detection automation device. Specific implementation manners
[0041] Next, the technical solutions of the present invention will be described in conjunction with the accompanying drawings in the embodiments of the present invention.
[0042] Combined with the attached Figure 1-15 The deep learning method for detecting surface defects of a sealed plug based on YOLOv5 includes the following steps:
[0043] Step 1, optimize the YOLOv5 network algorithm.
[0044] Step 1.1, add an attention mechanism SE module before the last layer of feature fusion in the backbone network of the YOLOv5 network algorithm; the attention mechanism SE module is used for adaptive feature extraction and correction of feature weights to improve the model's attention to defect targets;
[0045] As shown in the attached Figure 3 and 5 The attention mechanism SE module includes a compression part and an excitation part. The compression part is used to compress the spatial information in the input feature map, compressing H×W×C into 1×1×C, that is, using global average pooling in the spatial dimension to obtain a 1×1×C feature map; the excitation part is used to combine the learned attention information with the input feature and obtain a feature map U with channel attention;
[0046] The output of the compression operation in the compression part is obtained through global average pooling, calculating the global spatial average value for each channel c to generate a channel description vector z∈RC :
[0047]
[0048] Among them, z c represents the global description of the c-th channel;
[0049] The excitation part passes the compressed channel description vector z through a two-layer fully connected network to generate the channel weight s ∈ R C : s = σ(W2 · δ(W1 · z)); where W 1 ∈ R C / r×C and W 2 ∈ R C×C / r are the weight matrices of the two-layer fully connected network; δ(·) is the ReLU activation function; σ(·) is the Sigmoid activation function; r is the channel scaling coefficient, and r = 16 can be taken to reduce the computational amount;
[0050] The final output of the SE module is expressed as: Fc′ = s c · F c , where Fc′ is the weighted c-th channel, and s c is the corresponding channel weight.
[0051] The neck network of the YOLOv5 network algorithm adopts the FPN + PAN structure, which is used to fuse the multi-scale features extracted by the feature backbone network; the backbone network output provides three groups of feature maps P3, P4, and P5, corresponding to the small-scale feature map, medium-scale feature map, and large-scale feature map respectively; FPN is the upsampling fusion from bottom to top. P5 is adjusted to the same size as P4 through the first upsampling fusion operation, and then adjusted to the same size as P3 through the second upsampling fusion operation; PAN is the downsampling fusion from top to bottom. P3 is adjusted to the same resolution as P4 through the first downsampling fusion operation, and P4 is adjusted to the same resolution as P5 through the first downsampling fusion operation; and the CSP structure connects the FPN and PAN of the neck network.
[0052] Step 1.2, add the attention mechanism CBAM module after the two upsampling fusions of the neck network of the YOLOv5 network algorithm; the attention mechanism CBAM module combines the channel attention mechanism and the spatial attention mechanism to extract more comprehensive feature information in the defect area and increase the global receptive field of the feature extraction network;
[0053] As shown in the appendix Figure 4 and 6As shown, the attention mechanism CBAM module includes a channel attention module CAM and a spatial attention module SAM. The CAM module is used to weight the importance of each channel of the input feature map. By performing global average pooling and global max pooling on the feature map in the spatial dimension, a global feature of 1×1×C is obtained. The SAM module is used to assign different weights to each spatial position in the feature map. By performing global average pooling and max pooling on the channel dimension, a feature map of H×W is generated to express the importance of spatial positions. The input of CBAM is a feature map F∈R C×H×W , where C is the number of channels, and H and W are the height and width of the feature map respectively;
[0054] The CAM module performs global average pooling and global max pooling on the input feature map F in the spatial dimension to generate two vectors describing global features:
[0055]
[0056] favg and fmax respectively pass through a two-layer fully connected network (MLP) with shared weights, are added and activated to generate the channel attention weight M c :
[0057] M c =σ(MLP(AvgPool(F)) + MLP(MaxPool(F))) (3)
[0058] where σ is the Sigmoid activation function, MLP represents a multi-layer perceptron for learning the non-linear relationship between channels, F is the input feature map, and M c is the channel weight; the output of CAM is multiplied by the input feature map with the weight M c to obtain the channel-weighted feature map F′, and the input feature map is weighted according to the channel attention weight: F′ = M c ·F(4);
[0059] The SAM module performs global average pooling and global max pooling on the input feature map F′ in the channel dimension to generate two feature maps f′ avg and f′ max :
[0060]
[0061] The f′ avg and f′ max feature maps are concatenated and passed through a 7×7 convolutional layer to generate the spatial attention weight M s :
[0062] M s= σ(Conv7×7([AvgPool(F′), MaxPool(F′)])) (6)
[0063] where Conv7×7 represents a convolutional kernel of size 7×7, · represents the feature map concatenation operation, F′ is the feature map after channel weighting, and M s is the spatial weight; the output of SAM is multiplied by the weight M s and the input feature map F′ to obtain the final feature map F″ with channel and spatial attention;
[0064] The final output of the CBAM module can be expressed as: F″ = M s ·F′ (7), where · represents the element-wise multiplication operation, M s and M c are the spatial weight and channel weight respectively, and F″ is the finally enhanced feature map;
[0065] The present invention is tested and compared with various improved methods on the constructed sealed plug dataset. Method 1: Refer to the four attention mechanisms SE, CA, ECA, and CBAM proposed in the invention patent "A Method for Detecting Tile Surface Defects Based on Yolov5" and add them to the backbone network; Method 2: Add the CBAM attention mechanism to the neck network; Method 3: Add the SE attention mechanism to the backbone network; Method 4: Add the SimAM attention mechanism to the backbone network; Method 5: Add the CA attention mechanism to the backbone network; The method proposed by the present invention is to add the SE attention mechanism to the backbone network and the CBAM attention mechanism to the neck network; The comparison results are shown in Table 1:
[0066] Table 1 Comparison of the results of the method of the present invention with other methods
[0067] Model Precision(%) Recall(%) mAP@0.5(%) mAP@0.5:0.95(%) FPS YOLOv5s 86.3 83.7 88.4 53.1 95 Method 1 90.2 86.7 90.8 56.5 72 Method 2 89.1 85.8 91 55.8 78 Method 3 88.7 85.2 90.6 55.2 76 Method 4 87.5 84.1 89.7 54.1 77 Method 5 88.3 85.5 90.3 55.4 75 The method of the present invention 91.3 87.5 93.1 57.1 73 .
[0068] As can be seen from the above table, the method of the present invention has the highest Precision, Recall, mAP@0.5, and mAP@0.5:0.95 compared with other methods, indicating that the method of the present invention achieves an excellent balance in detection accuracy and recall rate, significantly improves the detection performance, and performs better in the constructed sealed plug dataset.
[0069] According to requirements, in step 1 of optimizing the YOLOv5 network algorithm, step 1.3 is further included, and all C3 modules in the backbone network and neck network of the YOLOv5 network algorithm are replaced with CF2 modules;
[0070] As shown in the appendix Figure 7 and 8As shown, the original C3 module in the YOLOv5 network mainly draws on the idea of CSPNet for extraction and splitting, and at the same time combines the idea of the residual structure to design the C3Block. The CSP main branch gradient module is the BottleNeck module, and the number of stacked ones is controlled by the parameter n;
[0071] The C2F module consists of four parts: the main branch Main Path, the branch split Split, the Bottleneck residual block, and the fusion operation Concat&Fusion; the main branch is the preliminary feature representation of the input feature X through a convolutional layer; the branch split divides the input feature into two parts: the first part passes directly, and the second part enters the Bottleneck residual block; the Bottleneck residual block includes several depthwise separable convolutional layers, ReLU activation functions, and skip connections Shortcut; the fusion operation is to perform a Concat splicing operation on the features of the direct path and the residual path of the main branch and complete the channel fusion process through a convolutional layer;
[0072] Feature representation of the main branch: F main = Conv(X)(8); where Conv represents the convolution operation;
[0073] Feature representation of the residual path:
[0074]
[0075] Among them, Bottleneck i represents the i-th Bottleneck block;
[0076] Feature fusion representation: F fused = Conv(Concat(F main , F residual ))(10); where Concat represents the splicing operation;
[0077] The final output of the C2F module can be expressed as: Y = F fused (11).
[0078] According to needs, in step 1 of optimizing the YOLOv5 network algorithm, there is also step 1.4, which introduces the Feature Pyramid Network BiFPN module into the neck network of the YOLOv5 network algorithm, and further removes the original Path Aggregation Network PANet structure of the YOLOv5 network algorithm, enhancing the information extraction ability of the network and further enhancing the detection performance of the network for small target defects;
[0079] As shown in the appendix Figure 9 and 10As shown, the Feature Pyramid Network (FPN) is a pioneering work that proposed a top-down approach to combine multi-scale features; the Path Aggregation Network (PANet) added a bottom-up path aggregation network on top of FPN, and the unimproved YOLOv5 network architecture uses it as the neck; the Bidirectional Feature Pyramid Network (BIFPN) optimizes multi-scale feature fusion in a more intuitive and principled way.
[0080] Introduce the BiFPN module of the Feature Pyramid Network into the neck network of the YOLOv5 network algorithm and remove the original PANet structure of the YOLOv5 network algorithm to obtain the optimized YOLOv5 network algorithm. Specifically, after the backbone network extracts the input image, it must be processed by the neck network and then output to the detection part. In order to better fuse the feature information, the BiFPN module of the Feature Pyramid Network combines the deep and shallow feature fusions in both top-down and bottom-up directions and repeats the same layer multiple times to achieve higher-level feature fusion. When fusing features with different resolutions, fast normalization is used in BiFPN to weight the fused features:
[0081]
[0082] where, F i is the i-th input feature map; w i is the corresponding weight; ∈ is a small value for numerical stability;
[0083] The present invention uses the BiFPN feature extraction network to replace the original PAN structure of YOLOv5s, which can not only assign different weights to objects of different scales, but also enhance the information extraction ability of the network, enabling the low-level unknown information to be combined with the high-level semantic information, thereby improving the target detection performance of the network and being able to adaptively adjust the difference degree between different features, reducing feature redundancy and improving computational efficiency, and enhancing the accuracy of small targets.
[0084] According to needs, step 1.5 is also included in the optimization of the YOLOv5 network algorithm in step 1, adding a P2 small target detection layer to the head network of the YOLOv5 network algorithm to enhance the detection accuracy of the algorithm for small targets;
[0085] Such as attached Figure 11As shown, in the unimproved YOLOv5 network structure, there are only 3 detection layers, P5, P4, and P3. The output sizes of the detection heads are 20×20×255, 40×40×255, and 80×80×255 respectively, corresponding to the initialized anchor boxes, which are used to detect larger, medium-sized, and smaller targets respectively; for the dataset of the present invention, the P3 detection layer is not ideal for detecting smaller defects. To improve its detection ability, a P2 detection layer is added to enhance the algorithm's detection ability for small targets, with an output size of 160×160×255, and a set of initial anchor boxes is added, which is obtained by the K-means clustering algorithm to obtain prior boxes with better clustering effects;
[0086] The newly added P2 detection layer is upsampled on the original P3 detection layer and then fused with the low-level features to form the shallow output features. The upsampling and feature fusion formula is:
[0087] P 2 = Conv(Cat(Upsample(P3), Feature Backbone_P2 )) (13)
[0088] Among them, P3: the feature output of the previous layer; FeatureBackbone_P2: the shallow features in the backbone network; Upsample: upsample the P3 feature map to match its resolution with P2; Cat: channel dimension concatenation operation; Conv: obtain the final P2 features through convolution operation after fusion;
[0089] The detection layer calculation formula of the newly added P2 detection layer is:
[0090] Detection P2 = {Anchor P2 , Cls P2 , Box P2} (14).
[0091] According to needs, step 1.6 is also included in optimizing the YOLOv5 network algorithm in step 1. The EIOU loss function is introduced in the YOLOv5 network algorithm to replace the original CIOU loss function, and the sample imbalance problem in the bounding box regression task is optimized;
[0092] In the unimproved YOLOv5 network structure, the CIoU loss function is used for bounding box regression. During the calculation of the loss value, the aspect ratio of the predicted box and the ground truth box is considered, effectively solving the problem of providing the moving direction for the bounding box in the case of non-overlap. However, since the CIoU loss function calculates all loss variables as a whole, it may lead to phenomena such as slow convergence speed and instability, and it also fails to consider the imbalance problem of easy and difficult samples. For the small target sample dataset in this case, directly using the CIoU loss function results in poor detection effects. The EIOU loss function of the present invention further improves and optimizes the CIOU. By adding a direct optimization term for the width and height differences, the loss function becomes more sensitive to the adjustment of the box boundary. At the same time, during the training process, it will actively optimize the matching degree between the width and height of the predicted box and the ground truth box, improving the detection accuracy for small targets and optimizing the sample imbalance problem in the bounding box regression task. The EIOU loss formula is expressed as:
[0093]
[0094] Among them, IoU: represents the overlapping area between the predicted box and the ground truth box; Center point distance, representing the position error of the box; + Box width difference, normalized to the maximum width of the box; Box height difference, normalized to the maximum height of the box.
[0095] Step 2, collect different defect categories of the seal plug and make a surface defect dataset of the seal plug; specifically, the surface defect dataset of the seal plug includes collecting different defect categories of the seal plug through a vision device and making a seal plug dataset, then using LabelImg to label the seal plug dataset, and at the same time generating a labeled.txt file, and randomly dividing the seal plug dataset and the corresponding.txt file into a training set, a validation set and a test set according to the ratio of 8:1:1.
[0096] Step 3, train the optimized YOLOv5 network algorithm through the surface defect dataset of the seal plug to obtain a YOLOv5 object detection network for surface defect detection of the seal plug.
[0097] Implementing the deep learning method for surface defect detection of the seal plug based on YOLOv5 of the present invention and applying it in a vision detection automation device, such as Figure 16 and 17As shown in the figure, the visual inspection automation device includes an inclined vibration feeding tray 1, a linear vibration feeder 2, a gear clamping turntable 3, a circular turntable 4, an infrared digital laser sensor 5, a sealing plug upper surface detection module 6, a diffuse reflection infrared sensor 7, a sealing plug lower surface detection module 8, and a solenoid valve rejection module 9. The sealing plug upper surface detection module consists of a DAheng USB2.0 industrial line array camera MER-125-30UM, a ring dot matrix light source, and a whiteboard. The industrial line array camera takes pictures of the upper surface of the sealing plug from top to bottom. The sealing plug lower surface detection module consists of a DAheng USB2.0 industrial line array camera MER-125-30UM, a ring dot matrix light source, and a whiteboard. The industrial line array camera takes pictures of the lower surface of the sealing plug from bottom to top. The industrial control screen 10 is used to display the operating status of the visual inspection automation device in real time, count the number of defective products of various types, and the change of the qualified rate of the sealing plug. The alarm lamp 11 is used for equipment failure and maintenance. The electrical control cabinet 12 controls the normal operation of the equipment. The placement station of the sealing plug is designed as a clockwise rotation station, and the upper and lower surfaces of the sealing plug can be collected with complete images in cooperation with the line array camera.
[0098] Taking a production line in a certain sealing plug production workshop as an application background, an example is described according to the above technical solution. 1000 products are randomly produced and compared with 5 mainstream object detection algorithms such as Faster R-CNN, YOLOv3, YOLOv4, YOLOv5, and YOLOv7 and the improved algorithm of the present invention. In this experiment, the Windows 11 operating system is used for training, the CPU processor model is Intel(R) Core(TM) i7-10870H CPU@2.20GHz, the GPU model is NVIDIA GeForce GTX 1660Ti, the programming language is Python 3.8.18, the operating environment is built based on the Pytorch 1.8.2 framework, and the CUDA 11.7 and cuDNN acceleration toolboxes are used. The parameter settings are as follows: the input image size is 640×640, the initial parameter learning rate is 0.01, the momentum is 0.937, the weight decay coefficient is 0.0005, the batch sample size is 16, and the epoch is 300.
[0099] The present invention uses precision (P), recall (R), average precision (AP), mean average precision (mAP), frames per second (FPS), and model size as performance evaluation indicators. Among them, P represents the ratio of correct results among all predictions of positive samples; R represents the ratio of correctly predicted positive samples among all positive samples; the PR curve is a curve with Recall as the abscissa and Precision as the ordinate. The larger the area in the lower left of the PR graph, the better the model's performance on this dataset, and the enclosed area is the average precision AP, which represents the average of the precision rates for each category; mAP represents the average of the APs for each category. The larger the mAP value, the better the performance of the model. Considering the actual situation of industrial production, frames per second (FPS) and model size are used as indicators to measure the real-time performance of the model. The specific calculation formulas for each indicator are as follows:
[0100]
[0101]
[0102] TP represents that both the true value and the predicted value are positive samples; FP represents that when the true value is a negative sample, the predicted value is a positive sample; FN represents that when the true value is a positive sample, the predicted value is a negative sample; N represents the number of defect types detected;
[0103]
[0104] Among them, P reTime is the preprocessing time of the image, in ms; I nferTime is the network inference time, in ms; N msTime is the non-maximum suppression time, in ms; n is the number of detected category types;
[0105] The optimized YOLOv5 algorithm of the present invention is compared with 5 mainstream object detection algorithms such as Faster R-CNN, YOLOv3, YOLOv4, YOLOv5, and YOLOv7, using six parameters, namely Precision, Recall, mAP@0.5, mAP@0.5:0.95, Model size, and FPS, as evaluation indicators. The performance of different networks in detecting surface defects of small and precise seals on the dataset is shown in Table 2:
[0106] Table 2 Comparison of the results of the algorithm of the present invention with other algorithms
[0107] Model Precision(%) Recall(%) mAP@0.5(%) mAP@0.5:0.95(%) <![CDATA[Mode lsize(MB)]]> FPS Faster R-CNN 84.5 81.7 85.3 49.2 540 12 YOLOv3 87.1 83.2 88.4 50.3 236 45 YOLOV4 88.7 85.4 90.2 53.7 244 65 YOLOV5 89.3 85.9 91.1 54.6 141 85 YOLOV7 90.1 86.3 92.4 56 143 83 Improved YOLOv5s 915 878 936 578 150 73
[0108] The optimized YOLOv5s of the present invention achieves a Precision of 91.5% in the constructed sealed plug dataset, with both Recall and mAP values improved. Its performance is superior to other models, with an FPS of 73, slightly lower than that of YOLOv7, but with the best comprehensive performance.
[0109] The parts not detailed in the present invention are prior arts. For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any perspective, the above embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be encompassed by the present invention, and any reference signs in the claims should not be regarded as limiting the content of the claims involved.
Claims
1. A deep learning method for detecting surface defects of sealing plugs based on YOLOv5, characterized by: Follow these steps: S1. Optimize YOLOv5 network algorithm; S1.
1. Add the attention mechanism SE module before the last layer of feature fusion of the backbone network of the YOLOv5 network algorithm; The attention mechanism SE module is used for adaptive feature extraction and correction of feature weights to improve the model's attention to defect targets; S1.2, after the two upsampling fusions of the neck network of the YOLOv5 network algorithm, the attention mechanism CBAM module is added; The attention mechanism CBAM module combines the channel attention mechanism with the spatial attention mechanism to extract more comprehensive feature information in the defect area and increase the global perception field of the feature extraction network. S2, collecting different defect categories of sealing plugs and preparing a sealing plug surface defect data set; S3. The YOLOv5 network algorithm optimized by the sealing plug surface defect data set is trained to obtain the YOLOv5 target detection network for sealing plug surface defect detection.
2. The deep learning method for detecting surface defects of sealing plugs based on YOLOv5 according to claim 1 is characterized in that: In S1.1, the attention mechanism SE module includes a compression part and an excitation part. The compression part is used to compress the spatial information in the input feature map and compress H×W×C into 1×1×C, that is, global average pooling is used in the spatial dimension to obtain a 1×1×C feature map; the excitation part is used to combine the learned attention information with the input features and obtain a feature map U with channel attention.
3. The deep learning method for detecting surface defects of sealing plugs based on YOLOv5 according to claim 2 is characterized in that: The compression operation output of the compression part is global average pooling, which calculates the global spatial average value for each channel c to generate a channel description vector z∈R C : Among them, z c represents the global description of the cth channel; The excitation part passes the compressed channel description vector z through a two-layer fully connected network to generate the channel weight s∈R C :s=σ(W2·δ(W1·z)); where W1∈R C / r×C and W2∈R C×C / r is the weight matrix of the two-layer fully connected network; δ(·) is the ReLU activation function; σ(·)\ is the Sigmoid activation function; r is the channel scaling factor, which can be taken as r=16 to reduce the amount of calculation; The final output of the SE module is expressed as: Fc′=s c ·F c , where Fc′ is the weighted c-th channel, s c is the corresponding channel weight.
4. The deep learning method for detecting surface defects of sealing plugs based on YOLOv5 according to claim 1 is characterized in that: In S1.2, the attention mechanism CBAM module includes a channel attention module CAM and a spatial attention module SAM. The CAM module is used to weight the importance of each channel of the input feature map, compress the feature map in the spatial dimension through global average pooling and global maximum pooling, and obtain a 1×1×C global feature; The SAM module is used to assign different weights to each spatial position in the feature map, and generates a H×W feature map to express the importance of the spatial position through global average pooling and maximum pooling in the channel dimension; the input of CBAM is a feature map F∈R C×H×W , where C is the number of channels, H and W are the height and width of the feature map, respectively.
5. The deep learning method for detecting surface defects of sealing plugs based on YOLOv5 according to claim 4 is characterized in that: The CAM module performs global average pooling and global maximum pooling on the input feature map F in the spatial dimension to generate two vectors describing the global features: favg and fmax are added and activated through two layers of shared weighted fully connected networks (MLP) to generate the channel attention weight M c : M c =σ(MLP(AvgPool(F))+MLP(MaxPool(F))) (3) Among them, σ is the Sigmoid activation function, MLP represents the multi-layer perceptron, which is used to learn the nonlinear relationship between channels, F is the input feature map, M c is the channel weight; the output of CAM is passed through the weight M c Multiply it with the input feature map to get the channel-weighted feature map F′. The input feature map is weighted according to the channel attention weight: F′=M c ·F(4); The SAM module performs global average pooling and global maximum pooling on the input feature map F′ in the channel dimension to generate two feature maps f′ avg and f′ max : f′ avg and f′ max The feature maps are concatenated and passed through a 7×7 convolutional layer to generate the spatial attention weight M. s : M s =σ(Conv7×7([AvgPool(F′),MaxPool(F′)])) (6) Among them, Conv7×7 represents the convolution kernel of size 7×7, · represents the feature map concatenation operation, F′ is the feature map after channel weighting, M s is the spatial weight; the output of SAM is passed through the weight M s Multiply it with the input feature map F′ to get the final feature map F″ with channel and spatial attention; The final output of the CBAM module can be expressed as: F″=M s ·F′(7), where · represents the element-by-element multiplication operation, M s and M c are spatial weight and channel weight respectively, and F″ is the final enhanced feature map.
6. The deep learning method for detecting surface defects of sealing plugs based on YOLOv5 according to claim 1 is characterized in that: The S1 optimized YOLOv5 network algorithm also includes S1.3, which replaces all C3 modules with CF2 modules in the backbone network and neck network of the YOLOv5 network algorithm.
7. The deep learning method for detecting surface defects of sealing plugs based on YOLOv5 according to claim 1 is characterized in that: The S1 optimization of the YOLOv5 network algorithm also includes S1.4, which introduces a feature pyramid network BiFPN module into the neck network of the YOLOv5 network algorithm, and further removes the original path aggregation network PANet structure of the YOLOv5 network algorithm to enhance the network's information extraction capability and further enhance the network's detection performance for small target defects.
8. The deep learning method for detecting surface defects of sealing plugs based on YOLOv5 according to claim 1 is characterized in that: The S1 optimized YOLOv5 network algorithm also includes S1.5, which adds a P2 small target detection layer to the head network of the YOLOv5 network algorithm to enhance the algorithm's detection accuracy for small targets.
9. The deep learning method for detecting surface defects of sealing plugs based on YOLOv5 according to claim 1 is characterized in that: The S1 optimization of the YOLOv5 network algorithm also includes S1.6, which introduces the EIOU loss function to replace the original CIOU loss function in the YOLOv5 network algorithm to optimize the sample imbalance problem in the bounding box regression task.
10. The deep learning method for detecting surface defects of sealing plugs based on YOLOv5 according to claim 1 is characterized in that: In S2, the sealing plug surface defect dataset includes collecting different defect categories of the sealing plug through a visual device and making a sealing plug dataset, then labeling the sealing plug dataset through LabelImg, and generating a labeled .txt file at the same time, and randomly dividing the sealing plug dataset and the corresponding .txt file into a training set, a validation set, and a test set in a ratio of 8:1:1.
Citation Information
Patent Citations
Ceramic tile surface defect detection method based on Yolov5
CN118071673A