A Defect Target Detection Method, System, Device and Medium Based on Improved YOLOv5
By improving the YOLOv5 model, adding a micro-object detection layer and a lightweight attention mechanism, and combining a multi-scale feature extraction module, an HT-YOLO model is formed, which solves the problems of low detection rate and high error detection rate of industrial oil seals in the existing technology, real-time and efficient detection of oil seal surface defects is achieved, and detection accuracy and reliability are improved.
Patent Information
- Application Number
- CN202510411635.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-04-02
AI Technical Summary
When existing machine vision detection technology detects defects such as small scratches, damage and bubbles in industrial oil seals, there are problems such as low detection rate, high error detection rate and insufficient feature extraction capability in complex industrial contexts.
Improve the network framework of the YOLOv5 model, add a micro-object detection layer and a lightweight attention mechanism HWA, and combine it with a multi-scale feature extraction module to form an HT-YOLO model to achieve real-time and efficient detection of oil seal surface defects.
It significantly improves the model's detection ability of small targets, improves the accuracy and reliability of defect recognition, optimizes the calculation amount, and maintains high detection accuracy.
Smart Images

Figure CN119919646B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of defect target detection, and specifically, to a defect target detection method, system, device and medium based on improved YOLOv5. Background Art
[0002] Industrial oil seals, as key sealing components in mechanical equipment, their defects directly affect the safety and service life of the equipment. Therefore, the detection of oil seal defects has become an important link in industrial manufacturing and quality control. Traditional detection methods mainly rely on manual visual inspection and simple mechanical measurement. Although obvious defects can be found to a certain extent, there are problems such as low efficiency, strong subjectivity and low detection accuracy.
[0003] With the development of industrial automation technology, detection methods based on machine vision have gradually been introduced into the detection of oil seal defects. By using image acquisition devices and image processing algorithms to automatically detect the surface of oil seals, the detection efficiency and accuracy can be effectively improved. However, the existing machine vision detection technologies still face some challenges in practical applications, mainly including the following aspects:
[0004] 1) The defects of industrial oil seals usually appear as fine scratches, cracks, bubbles, etc. These defects are small in size and have blurred edges. In a complex industrial background, the detection algorithm has insufficient ability to extract their features, which easily leads to missed detection or false detection.
[0005] 2) Usually, the surface of the oil seal has a high gloss and is easily interfered by environmental light and equipment reflection, resulting in low contrast and a lot of noise in the collected images. Traditional algorithms are difficult to accurately segment the target from the background in this case.
[0006] 3) It is difficult to balance the real-time performance and detection accuracy of existing target detection algorithms. Especially in the high-speed operation scenario of an industrial assembly line, traditional detection algorithms are difficult to meet the requirements of efficient and real-time detection.
[0007] Currently, deep learning technology has made remarkable progress in the field of computer vision. Especially target detection algorithms based on convolutional neural networks (CNNs) have been widely used in the field of industrial detection due to their powerful feature extraction ability and end-to-end detection method. Among them, the YOLO (You Only Look Once) series of algorithms have become important tools for target detection due to their high efficiency in real-time performance and excellent detection performance. However, the standard YOLO model still has the following deficiencies in the task of industrial oil seal defect detection:
[0008] 1) It has limited ability to extract features of small-scale targets (such as tiny scratches or bubbles) and is easily interfered by complex background noise, resulting in a decrease in detection accuracy;
[0009] 2) The model has insufficient generalization ability when facing diverse defect types, and the detection effects on different types of defects are unbalanced;
[0010] 3) Traditional YOLO models have limitations in feature map fusion and multi-scale detection capabilities, making it difficult to balance detailed information and global features, especially when detecting complex defects on the surface of oil seals. Summary of the Invention
[0011] In view of the problems that the existing detection methods have a low detection rate for small-sized targets and are difficult to detect complex defects on the surface, the present invention proposes a defect target detection method, system, device, and medium based on improved YOLOv5. The method first collects image data of the target to be detected during the production process, and performs image preprocessing according to the characteristics of surface defects of the target to be detected. Secondly, key frames are extracted and the defect areas are labeled to construct a defect detection data set. Then, the network framework of the YOLOv5 model is improved to obtain the HT-YOLO model. Finally, the HT-YOLO model is trained to achieve real-time and efficient detection of surface defects of the target to be detected, improving the detection accuracy and speed.
[0012] The specific implementation content of the present invention is as follows:
[0013] A defect target detection method based on improved YOLOv5 specifically includes the following steps:
[0014] Step S1: Collect image data of the target to be detected during the production process and label it to obtain a target defect data set;
[0015] Step S2: Improve the network framework of the YOLOv5 model to obtain the HT-YOLO model;
[0016] Step S3: Train the HT-YOLO model according to the target defect training set to obtain the trained HT-YOLO model;
[0017] Step S4: Evaluate the trained HT-YOLO model according to the target defect test set, and adjust the HT-YOLO model according to the evaluation results to identify the target defect.
[0018] To better implement the present invention, further, the step S1 specifically includes the following steps:
[0019] Step S11: Collect target image data according to the set acquisition time period;
[0020] Step S12: Online search and obtain data related to the target to be detected, and construct the original target data;
[0021] Step S13: Screen and clean the original target data, and divide the cleaned data into a target defect training set, a target defect validation set, and a target defect test set at a ratio of 8:1:1.
[0022] To better implement the present invention, further, the network framework of the improved YOLOv5 model in step S2 specifically includes the following operations:
[0023] Operation 1: Adjust the hierarchical structure of the head part of the YOLOv5 model, fuse shallow semantic information, add a tiny target detection layer and a tiny target anchor box group;
[0024] Operation 2: Replace the C3 module of the YOLOv5 model with an HWAttention attention module;
[0025] Operation 3: When connecting the detection layers, combine the multi-scale feature extraction module with the HWAttention module to form a TC_HWA module.
[0026] To better implement the present invention, further, the specific operation of Operation 1 is as follows: First, adjust the hierarchical structure of the head part of the YOLOv5 model, fuse shallow semantic information, add a group of tiny target detection layers P2 on the basis of the three detection layers, and then add a P2 upsampling fusion module after two rounds of upsampling fusion modules of detection layer P3 and detection layer P4, and connect it to the tiny target detection layer P2; at the same time, set a tiny target anchor box group corresponding to the scale of the tiny target detection layer P2.
[0027] To better implement the present invention, further, step S3 specifically includes the following steps:
[0028] Step S31: Obtain the target shallow features from the acquired target defect feature map according to the tiny target detection layer and the tiny target anchor box group;
[0029] Step S32: Reduce the dimension of the target shallow features, splice them in the channel dimension to obtain the fused target shallow features, generate attention weights according to the fused target shallow features, and weight them to the target shallow feature map;
[0030] Step S33: Generate channel attention weights according to channel attention, reorganize the target shallow feature map, perform residual connection on the reorganized target shallow feature map to obtain the final output feature map, and then convert it into detection result information through the detection layer.
[0031] To better implement the present invention, further, step S32 specifically includes the following steps:
[0032] Step S321: Use average pooling and max pooling to reduce the dimension of the target shallow features respectively, compress and reduce the dimension of the channel dimension, and obtain the dimension-reduced smooth features and dimension-reduced significant features of the target shallow features respectively;
[0033] Step S322: Concatenate the dimension-reduced smooth features and dimension-reduced significant features in the channel dimension to obtain the fused target shallow features;
[0034] Step S323: Keep the spatial dimension unchanged, call a 5×5 convolution to compress the fused target shallow features into an attention score map, and call the Sigmoid activation function to weight the score map to obtain the attention weights;
[0035] Step S324: Weight the attention weights to the target shallow feature map.
[0036] To better implement the present invention, further, the step S33 specifically includes the following steps:
[0037] Step S331: Call the attention module to generate channel attention weights;
[0038] Step S332: Sort the channels according to the channel attention weights to obtain a rearranged channel sequence and the attention weights corresponding to the channel sequence;
[0039] Step S333: Reorganize the target shallow feature map according to the channel sequence and weight the reorganized target shallow feature map;
[0040] Step S334: Divide the weighted target shallow feature map in the channel dimension according to a ratio of 1:2:1, and perform convolution extraction according to the magnitude of the weights;
[0041] Step S335: Fuse the target shallow feature map and the current output feature map through residual connection, and organize the fused features through convolution operation on the channels to obtain the final output feature map, which is then converted into detection result information through the detection layer.
[0042] To better implement the present invention, further, the step S4 specifically includes the following steps:
[0043] Step S41: Train the HT-YOLO model according to the target defect training set;
[0044] Step S42: Input the target defect test set into the trained HT-YOLO model, and obtain the evaluation result according to the set target detection task and key indicators;
[0045] Step S43: Adjust the hyperparameters of the HT-YOLO model according to the evaluation result to obtain the adjusted HT-YOLO model;
[0046] Step S44: Obtain the target defect according to the adjusted HT-YOLO model.
[0047] Based on the above-mentioned defect target detection method based on improved YOLOv5, in order to better implement the present invention, further, a defect target detection system based on improved YOLOv5 is proposed, including an acquisition unit, an improvement unit, a training unit, and an evaluation and detection unit;
[0048] The acquisition unit is used to acquire the image data of the target to be detected during the production process and label it to obtain a target defect data set;
[0049] The improvement unit is used to improve the network framework of the YOLOv5 model to obtain the HT-YOLO model;
[0050] The training unit is used to train the HT-YOLO model according to the target defect training set to obtain the trained HT-YOLO model;
[0051] The evaluation and detection unit is used to evaluate the trained HT-YOLO model according to the target defect test set, and adjust the HT-YOLO model according to the evaluation result to identify the target defect.
[0052] Based on the above-mentioned defect target detection method based on improved YOLOv5, in order to better implement the present invention, further, an electronic device is proposed, including a memory and a processor; a computer program is stored on the memory; when the computer program is executed on the processor, the above-mentioned defect target detection method based on improved YOLOv5 is implemented.
[0053] Based on the above-mentioned defect target detection method based on improved YOLOv5, in order to better implement the present invention, further, a computer-readable storage medium is proposed, and a computer instruction is stored on the computer-readable storage medium; when the computer instruction is executed on the above-mentioned electronic device, the above-mentioned defect target detection method based on improved YOLOv5 is implemented.
[0054] The present invention has the following beneficial effects:
[0055] (1) By introducing a tiny target detection layer, combining with the P2 upsampling fusion block, effectively fusing shallow features, and setting up a special tiny target anchor box group, the present invention significantly enhances the model's detection ability for tiny targets and improves the accuracy and reliability of defect recognition.
[0056] (2) The present invention designs a lightweight attention mechanism HWA. By performing average pooling and max pooling dimensionality reduction processing on the feature map, it extracts the fused features of global smooth features and local significant features, and then generates attention scores and weights the input feature map; effectively improving the expression ability of the feature map, while optimizing the computational complexity and maintaining high detection accuracy; by replacing the original C3 module in YOLO, the computational complexity of the model is significantly reduced without affecting the detection performance.
[0057] (3) The present invention introduces a multi-scale feature extraction module. By performing convolution operations of different sizes on the feature map enhanced and screened by the channel attention mechanism according to the feature weight size; not only effectively enhancing the channel features, but also realizing the extraction of multi-scale features, thus significantly improving the feature expression ability of the model and the accuracy of object detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 It is a flowchart of the defect detection method based on the improved YOLOv5 provided by the present invention.
[0059] Figure 2 It is the model of HT-YOLO provided by the present invention.
[0060] Figure 3 It is the diagram of the HWA module provided by the present invention.
[0061] Figure 4 It is the diagram of the multi-scale feature extraction module provided by the present invention.
[0062] Figure 5 It is the schematic diagram of the interface provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0063] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will combine the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. It should be understood that the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments, and therefore should not be regarded as a limitation of the protection scope. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.
[0064] In the description of the present invention, it should be noted that, unless otherwise clearly specified and defined, the terms "set", "connected", and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can also be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0065] Embodiment 1:
[0066] This embodiment proposes a defect target detection method based on improved YOLOv5, which specifically includes the following steps.
[0067] Step S1: Collect image data of the target to be detected during the production process and label it to obtain a target defect data set;
[0068] The step S1 specifically includes the following steps:
[0069] Step S11: Collect image data of the target to be detected according to the set acquisition time period;
[0070] Step S12: Search and obtain data related to the target to be detected online and construct the original target data;
[0071] Step S13: Screen and clean the original target data, and divide the cleaned data into a target defect training set, a target defect validation set, and a target defect test set at a ratio of 8:1:1.
[0072] Step S2: Improve the network framework of the YOLOv5 model to obtain the HT-YOLO model;
[0073] The improvement of the network framework of the YOLOv5 model in step S2 specifically includes the following operations:
[0074] Operation 1: Adjust the hierarchical structure of the head part of the YOLOv5 model, fuse shallow semantic information, and add a small target detection layer and a small target anchor box group;
[0075] Operation 2: Replace the C3 module of the YOLOv5 model with the HWAttention attention module;
[0076] Operation 3: When connecting the detection layers, combine the multi-scale feature extraction module with the HWAttention module to form the TC_HWA module.
[0077] The specific operation of the above-mentioned Operation 1 is as follows: First, adjust the hierarchical structure of the head part of the YOLOv5 model to fuse shallow semantic information. Add a set of tiny object detection layers P2 on the basis of the three detection layers. Then, after two rounds of upsampling fusion modules of detection layer P3 and detection layer P4, add the P2 upsampling fusion module and connect it to the tiny object detection layer P2. At the same time, set a group of tiny object anchor boxes corresponding to the scale of the tiny object detection layer P2.
[0078] Step S3: Train the HT-YOLO model according to the target defect training set to obtain the trained HT-YOLO model;
[0079] The specific steps of the above-mentioned Step S3 include the following steps:
[0080] Step S31: Obtain the target shallow features according to the tiny object detection layer and the group of tiny object anchor boxes from the acquired target feature map;
[0081] Step S32: Reduce the dimension of the target shallow features, splice them in the channel dimension to obtain the fused target shallow features, generate attention weights according to the fused target shallow features, and weight them to the target shallow feature map;
[0082] The specific steps of the above-mentioned Step S32 include the following steps:
[0083] Step S321: Reduce the dimension of the target shallow features by average pooling and max pooling respectively, compress the dimension of the reduced channels, and obtain the reduced smooth features and reduced significant features of the target shallow features respectively;
[0084] Step S322: Splice the reduced smooth features and reduced significant features in the channel dimension to obtain the fused target shallow features;
[0085] Step S323: Keep the spatial dimension unchanged, call a 5×5 convolution to compress the fused target shallow features into an attention score map, and call the Sigmoid activation function to weight the score map to obtain the attention weights;
[0086] Step S324: Weight the attention weights to the target shallow feature map.
[0087] Step S33: Generate channel attention weights according to the channel attention, reorganize the target shallow feature map, perform residual connection on the reorganized target shallow feature map to obtain the final output feature map, and then convert it into detection result information through the P2, P3, P4, and P5 detection layers; The detection result information includes the position and category information of the finally output detection box, which is the final stage of detection.
[0088] The specific steps of the above-mentioned Step S33 include the following steps:
[0089] Step S331: Invoke the attention module to generate channel attention weights;
[0090] Step S332: Sort the channels according to the channel attention weights to obtain a rearranged channel sequence and the attention weights corresponding to the channel sequence;
[0091] Step S333: Recombine the target shallow feature map according to the channel sequence and weight the recombined target shallow feature map;
[0092] Step S334: Divide the weighted target shallow feature map in a 1:2:1 ratio in the channel dimension and perform convolution extraction according to the magnitude of the weights;
[0093] Step S335: Fuse the target shallow feature map with the current output feature map through residual connection and organize the fused feature map through convolution operation on the channels to obtain the final output feature map set.
[0094] Step S4: Evaluate the trained HT-YOLO model according to the target defect test set, and adjust the HT-YOLO model according to the evaluation results to identify the target defect.
[0095] The specific steps of the said Step S4 include the following steps:
[0096] Step S41: Train the HT-YOLO model according to the target defect training set;
[0097] Step S42: Input the target defect test set into the trained HT-YOLO model and obtain the evaluation results according to the set target detection tasks and key indicators;
[0098] Step S43: Adjust the hyperparameters of the HT-YOLO model according to the evaluation results to obtain the adjusted HT-YOLO model;
[0099] Step S44: Obtain the target defect according to the adjusted HT-YOLO model.
[0100] Working principle: In this embodiment, image data of the target to be detected during the production process is first collected by a high-definition camera. Image preprocessing is performed according to the characteristics of surface defects of the target to be detected, key frames are extracted and the defect areas are marked to construct a target defect detection dataset. Then, the network framework is improved based on the YOLOv5 model: 1) Add a small target detection layer. 2) Use the newly added spatial attention module HWA to enhance the extracted features. 3) Combine the channel attention EffectiveSELayer to construct a multi-scale feature extraction module, and replace the original YOLOv5 module C3 with HWA. Finally, the improved YOLOv5 model is used to train the dataset. Through this embodiment, real-time and efficient detection of surface defects of the target to be detected is achieved, and the detection accuracy and speed are improved simultaneously.
[0101] Embodiment 2:
[0102] Based on the above Embodiment 1, this embodiment is described in detail with a specific example.
[0103] Step S1: Use a high-definition camera to collect image data of industrial oil seals during the production process, and label and enhance the collected data to generate a dataset;
[0104] Step S2: Improve the network framework based on the YOLOv5 model to obtain an improved HT-YOLO model; the specific improvements of the HT-YOLO model are as follows:
[0105] 1) Adjust the hierarchical structure of the head part, fuse shallow semantic information, add a small target detection layer, and increase the small target anchor box group;
[0106] 2) Use the HWAttention attention module to replace the C3 module of the original YOLOv5 model;
[0107] 3) When connecting the detection layer, combine the multi-scale feature extraction module with the HWAttention module to form a TC_HWA module.
[0108] Step S3: Use the HT-YOLO model to train and test the processed oil seal dataset for defect detection.
[0109] Furthermore, based on the original three-layer detection layer, a group of P2 small target detection layers are added. After passing through the original two rounds of upsampling and fusion modules of P3 and P4, a P2 upsampling and fusion module is added to obtain shallow features and connect them to the detection layer. At the same time, set the small target anchor box group corresponding to the P2 scale: [5, 68, 14, 15, 11] to improve the detection accuracy of small defect targets on the oil seal.
[0110] Further, the HWAttention module enhances the target features in the spatial dimension as follows:
[0111] First, the input features are compressed and reduced in the channel dimension through average pooling and max pooling respectively, obtaining the reduced and smoothed features and the reduced and significant features of the input features. Then, these two parts of features are concatenated in the channel dimension, and the fused features contain both global average features and significant features.
[0112] Then, a 5×5 convolution is used to compress the fused features into an attention score map with the spatial dimension unchanged and the number of channels being 1. After that, the score map is weighted through the Sigmoid activation function to obtain the attention weights.
[0113] Finally, the obtained weights are weighted to the input feature map of the module, thereby strengthening specific positions of the input feature map.
[0114] Further, the multi-scale feature extraction module enhances features in the channel dimension, extracts and fuses multi-scale features as follows:
[0115] The input feature map first generates channel attention weights through the channel attention module (EffectiveSELayer). Then, based on these weights, the channels are sorted to obtain the rearranged channel sequence and its corresponding attention weights. According to this sorting order, the channels of the input feature map are reorganized, and the reorganized feature map is weighted to enhance the expression ability of important channels. At this time, the channels of the feature map are sorted according to the weight size;
[0116] Next, the feature map is divided in the channel dimension according to the ratio of 1:2:1. The three divided parts respectively use 1×1, 3×3, and 5×5 convolutions for feature extraction according to the weights from low to high, and the number of channels remains unchanged during the extraction process. The output feature maps of the convolution operations are concatenated in the channel dimension. At this time, the input feature map and the current output feature map are fused through a residual connection to further enhance the expression ability of the model;
[0117] Finally, a 1×1 convolution operation is used to organize the channels of the fused feature map to obtain the final output feature map.
[0118] Further, step S3 specifically includes the following operations:
[0119] 1) Input the divided training set into the improved HT-YOLO based on YOLO for training;
[0120] 2) Input the test set into the trained HT-YOLO model to test its detection ability on new data. Through the object detection task on the test set, evaluate the performance of the HT-YOLO model, and focus on testing key indicators such as accuracy, recall, mean average precision (mAP), and F1 score in industrial oil seal defect detection to ensure that the model can maintain high detection accuracy and reliability in diverse oil seal defect samples;
[0121] 3) According to the evaluation results, adjust the hyperparameters of the HT-YOLO model to improve its detection effect on the test set.
[0122] Working principle: In this embodiment, by optimizing the network structure of YOLOv5, the detection ability of the model for small targets and complex defects is enhanced, and at the same time, the real-time performance and robustness of detection are improved; aiming at the characteristics of the industrial oil seal detection task, a refined object detection process is designed, which can effectively improve the detection accuracy and efficiency of industrial oil seal defects and meet the requirements of high-efficiency and accurate detection in industrial production.
[0123] Other parts of this embodiment are the same as those of the above Embodiment 1, so they will not be elaborated here.
[0124] Embodiment 4:
[0125] Based on any one of the above Embodiment 1 - Embodiment 2, as Figure 5 shown, take an interface for collecting surface defects of industrial oil seals as an example for detailed description.
[0126] Step S1: Data collection and processing.
[0127] First, install high-definition monitoring equipment in the production environment of industrial oil seals. At the same time, the monitoring equipment needs to collect data at different time periods to ensure that the data covers the daily production peak period, low peak period, as well as the situations during equipment maintenance and shutdown. This data includes pictures and videos, and the content should involve normal production status, production anomalies such as equipment failures, oil seal quality problems, and other scenarios that may affect production. By collecting data under different weather and production environments, the robustness and adaptability of the subsequent data analysis model are improved. Secondly, use tools such as web crawlers to search and obtain more pictures, videos, and technical literature data related to industrial oil seals online to enrich the diversity of the dataset. The crawled content includes oil seal products of different brands, models, and materials, as well as their usage effects and failure modes in different application scenarios. Screen and clean the obtained data to remove blurred, duplicate, and irrelevant images and content to ensure the high quality of the collected data. After the data collection is completed, the original data needs to be screened and cleaned. First, remove blurred or damaged pictures and videos to ensure the accuracy of the data. Then, perform data augmentation processing to increase the diversity of data samples by means of rotation, scaling, translation, cropping, adding noise, etc. Then, use annotation tools to annotate the data, such as annotating the damage types of oil seals (such as cracks, aging, oil leakage, etc.) and different working states (such as normal, worn, failed, etc.). Finally, divide the cleaned and augmented dataset into a training set, a validation set, and a test set according to 8:1:1 to ensure the comprehensiveness and balance of the data.
[0128] Step S2: Optimize the YOLO network model.
[0129] Based on the YOLOv5 algorithm, an improved HT-YOLO model is obtained, as Figure 2 . Among them, the specific optimization details are as follows:
[0130] First, since the original YOLOv5 model has a large receptive field in its detection layers (P3, P4, P5) when dealing with tiny defects on the surface of oil seals, it is unable to effectively capture these fine targets. Especially at lower resolutions, tiny defects are usually difficult to distinguish from the background, resulting in missed detections or false detections. To address this challenge, we made targeted improvements to the YOLOv5 architecture by adding a dedicated tiny target detection layer P2 with a smaller receptive field, which can more precisely capture the tiny defects on the surface of oil seals. Correspondingly, in view of the characteristics of tiny defects on the surface of oil seals, we redesigned the anchor boxes to make them more suitable for detecting small-scale and fine defects. These small-sized anchor boxes help to more precisely focus on the target area during the detection process, reducing the problem of decreased detection accuracy caused by the mismatch of anchor box scales. The introduction of this layer effectively enhances the sensitivity to small targets and improves the performance of the model in tiny target detection.
[0131] Meanwhile, in order to reduce the additional parameters and computational cost brought by adding the detection layer, the C3 module of the original model is replaced with the new module HWA. The HWA module is as Figure 3 .
[0132] The C3 module in the original YOLOv5 model is composed of multiple convolutional layers stacked together. The convolutional operation itself has a large computational cost, which will lead to an increase in time and resource consumption during training and inference. Moreover, C3 focuses on global feature extraction and cannot well capture local key features such as target details or tiny changes. This makes the C3 module may not achieve the optimal effect in fine-grained tasks, especially when accurate localization and recognition of small objects are required. Therefore, we introduce the lightweight attention module HWA. By using channel average pooling and channel maximum pooling, it can compress the channel dimension of the feature map to half of the original while extracting smooth features and significant features, reducing the computational cost of subsequent operations. After obtaining the compressed features, a 5×5 convolution is used to extract the attention score map, compressing the feature map to a size of 1×H×W. Then, the score map is weighted through Sigmoid, mapping its value to the range of [0,1] to obtain the spatial attention weight of the feature map. Finally, the obtained weight map is used to weight each channel of the input feature map, increasing the attention to important features in the spatial dimension while suppressing the representation of unimportant features. The lightweight attention HWA module can better capture local features and enhance their representation on the premise of significantly reducing the number of parameters and computational cost, while taking into account the extraction of global features, improving the accuracy of feature extraction.
[0133] Finally, a multi-scale feature extraction module is introduced into the head network. The multi-scale feature extraction module is as Figure 4 . After the HWA module connecting the detection layer, a multi-scale feature extraction module is added to enhance the feature representation of the model and improve the detection effect.
[0134] First, the input feature map is enhanced in the channel dimension through the channel attention module EffectiveSELayer. The weight values calculated by the channel attention mechanism are used to perform weighted screening on each channel, and the feature map is reordered according to the magnitude of the weights. After the feature map is sorted from low to high by weight, it is divided into three parts according to the ratio of 1:2:1 to distinguish the channels of different importance. Then, different sizes of convolution operations are applied to these three groups of feature maps respectively. Among them, a 1×1 convolution kernel is used for feature extraction on the low-weight channels, a 3×3 convolution kernel is used for the medium-weight channels, and a 5×5 convolution kernel is used for the high-weight channels for deep feature extraction. The core of this strategy is that the high-weight channels correspond to more important features. Therefore, using a larger-size convolution kernel can capture richer and more complex context information and improve the model's ability to express important features. Finally, the feature maps processed by different-size convolution kernels are concatenated in the channel dimension and merged into a comprehensive feature map, and then the channel information is further sorted through point convolution, and finally the optimized feature representation is output. This method realizes the fine processing of the important channels of the feature map by combining channel attention and scaled convolution operations, thereby improving the feature extraction ability of the network in complex tasks.
[0135] Step S3: Train the improved YOLO model.
[0136] The training process of the improved YOLO model first depends on the constructed oil seal defect dataset, which contains diverse defect samples and is accurately labeled. During the training process, a phased strategy is adopted. First, the weights of the model are initialized to accelerate the training process and improve the convergence speed of the model. Then, the performance of the model is evaluated in real time using the validation set to ensure that the model can maintain good generalization ability under different data distributions. According to the evaluation results on the validation set, the training hyperparameters such as the learning rate, batch size, and number of iterations are adjusted in real time to ensure that the model is gradually optimized during the training process and avoid overfitting or underfitting. In each cycle of training, the learning rate adopts a dynamic adjustment strategy, such as using methods like CosineAnnealing, and the learning rate is gradually decreased according to the performance of the model, thereby accelerating the convergence of the model and improving the final performance. The batch size and the number of iterations are adjusted according to the scale of the training set and the convergence speed of the model to ensure that the training process is both efficient and robust. In addition, to prevent overfitting, an early stopping strategy is also adopted during the training process. When the performance of the model on the validation set no longer improves significantly, the training is automatically stopped to avoid wasting computing resources and ensure the generalization ability of the model. Through these refined training strategies, it is ensured that the improved YOLO model can achieve the best detection performance in the complex oil seal defect detection task.
[0137] Step S4: Verification and scoring of the model.
[0138] When evaluating the model performance on the validation set, detection accuracy (Precision, P), recall (Recall, R), mean average precision (Mean Average Precision, mAP), and fitness (Fitness, Fit) are mainly used as evaluation metrics.
[0139] Among them, detection accuracy measures the proportion of targets that are actually defective among all the targets predicted as defective by the model, reflecting the reliability of the model in defect recognition. The calculation formula is as follows:
[0140]
[0141] Recall represents the proportion of defective targets successfully detected by the model among all the actually existing defective targets, measuring the comprehensiveness of the model. The calculation formula is as follows:
[0142]
[0143] Mean average precision comprehensively considers the precision and recall performance of the model under different IoU thresholds, and is usually calculated by the area under the precision-recall curve. It is an important metric for evaluating the global performance of the model. The calculation formula is as follows:
[0144]
[0145] In addition, fitness, as a comprehensive metric, comprehensively evaluates the precision and recall of the model, and can measure the applicability of the model in practical applications. Fitness is usually calculated through a weighted combination of precision, recall, and mean average precision at different thresholds to reflect the balance between real-time performance and comprehensiveness of the model in an industrial environment. The calculation formula is as follows:
[0146]
[0147] In the formula, TP represents the number of samples correctly classified as positive samples (actually defective, and the model detects them as defective); FP represents the number of samples misclassified as positive samples (actually normal, and the model detects them as defective); FN represents the number of samples misclassified as negative samples (actually defective, and the model detects them as normal). N represents the number of defect categories. w i represents the weight coefficient of the i-th metric. mAP@50 and mAP@95 represent the mean average precision obtained when the threshold is set to 50% and 95% respectively.
[0148] Through these comprehensive metrics, the actual performance of the improved YOLOv5 model in oil seal defect detection can be comprehensively evaluated to ensure its optimization in terms of precision, recall, and applicability.
[0149] As shown Figure 5 in the figure, where ① is the front detection picture of the industrial oil seal, ② is the back detection picture of the industrial oil seal, ③ is the inner detection picture of the industrial oil seal, ④ is the outer detection picture of the industrial oil seal, and ⑤ is the size detection picture of the industrial oil seal;
[0150] The time consumed for the front detection is 615 ms, the time consumed for the back detection is 240 ms, the time consumed for the inner detection is 151 ms, the time consumed for the outer detection is 397 ms, and the time consumed for the size detection is 109 ms.
[0151] The PPM value of 84159 indicates the number of NG products that may occur per million industrial oil seal products according to the qualification rate.
[0152] Figure 5 The number 53 in the upper right corner represents the overall detection speed, including the glass disk, roller, conveyor belt, and handwheel speed; TG2-40-62.05-9.5 / 12 C7808 is the industrial oil seal model, and batch 1 represents the current batch. The same industrial oil seal may be made by multiple molds. For example, mold 2 can be switched to batch 2.
[0153] Figure 5 For the curve graph in, the X-axis represents the quantity and the Y-axis represents the size. 62.48 is the maximum allowable outer diameter of the industrial oil seal, and 62.34 is the minimum allowable outer diameter of the industrial oil seal; OK, NG, and ER below the curve graph represent the log pictures. Whether the relevant log pictures are enabled during the detection process. For example, if NG is selected, the NG pictures generated during the detection process will be stored in a fixed folder for easy viewing; the bar graph is the normal distribution graph of the industrial oil seal. For example, there are 4437 industrial oil seals with a size of 62.44.
[0154] Other parts of this embodiment are the same as any one of the above-mentioned Embodiment 1 - Embodiment 2, so they will not be elaborated here.
[0155] Embodiment 4:
[0156] Based on any one of the above-mentioned Embodiment 1 - Embodiment 3, this embodiment proposes a defect target detection system based on improved YOLOv5, including a collection unit, an improvement unit, a training unit, and an evaluation and detection unit;
[0157] The collection unit is used to collect image data of the target to be detected during the production process and label it to obtain a target defect data set;
[0158] The improvement unit is used to improve the network framework of the YOLOv5 model to obtain the HT-YOLO model;
[0159] The training unit is used to train the HT-YOLO model according to the target defect training set to obtain the trained HT-YOLO model;
[0160] The evaluation and detection unit is used to evaluate the trained HT-YOLO model according to the target defect test set, and adjust the HT-YOLO model according to the evaluation result to identify the target defect.
[0161] This embodiment also provides an electronic device, including a memory and a processor; a computer program is stored on the memory; when the computer program is executed on the processor, the above-mentioned defect target detection method based on improved YOLOv5 is implemented.
[0162] This embodiment also provides a computer-readable storage medium, on which a computer instruction is stored; when the computer instruction is executed on the above-mentioned electronic device, the above-mentioned defect target detection method based on improved YOLOv5 is implemented.
[0163] Other parts of this embodiment are the same as any one of the above Embodiment 1 - Embodiment 3, so they will not be described in detail.
[0164] The above is only a preferred embodiment of the present invention, and does not impose any form of limitation on the present invention. Any simple modification or equivalent change made to the above embodiments based on the technical essence of the present invention falls within the protection scope of the present invention.
Claims
1. A defect target detection method based on improved YOLOv5, characterized in that: The specific steps include: Step S1: collecting image data of the target to be detected during the production process, annotating to obtain a target defect data set, and dividing the target data set into a target defect training set, a target defect test set, and a target defect verification set; Step S2: Improve the network framework of the YOLOv5 model to obtain the HT-YOLO model; The HT-YOLO model includes a small target detection layer and a small target anchor frame group arranged in the head part, an HWAttention attention module arranged in the neck part, and a TC_HWA module which is a fusion of a multi-scale feature extraction module and the HWAttention attention module; Step S3: training the HT-YOLO model according to the target defect training set to obtain a trained HT-YOLO model; The step S3 specifically comprises the following steps: Step S31: Obtain shallow features of the target according to the acquired target feature map based on the small target detection layer and the small target anchor frame group; Step S32: reduce the dimension of the shallow features of the target, and splice them in the channel dimension to obtain the fused shallow features of the target. According to the fused shallow features of the target, generate attention weights and weight them to the shallow feature map of the target. Step S33: Generate channel attention weights according to channel attention, reorganize the target shallow feature map, and residually connect the reorganized target shallow feature map to obtain the final output feature map; The step S33 specifically includes the following steps: Step S331: calling the attention module to generate channel attention weights; Step S332: Sort the channels according to the channel attention weights to obtain a channel sequence and attention weights corresponding to the channel sequence; Step S333: reorganizing the target shallow feature map according to the channel sequence, and weighting the reorganized target shallow feature map; Step S334: Divide the weighted shallow feature map of the target in the channel dimension according to the ratio of 1:2:1, and perform convolution extraction according to the size of the weight; Step S335: Fusing the target shallow feature map with the current output feature map through residual connection, and sorting the fused feature map through convolution operation channel to obtain the final output feature map, and converting it into detection result information through the detection layer; Step S4: Evaluate the trained HT-YOLO model based on the target defect test set, and adjust the HT-YOLO model parameters according to the evaluation results to obtain the target defect.
2. A defect target detection method based on improved YOLOv5 according to claim 1, characterized in that: The step S1 specifically includes the following steps: Step S11: collecting the image data of the target to be detected according to the set collection time period; Step S12: online search and acquisition of data related to the target to be detected, and constructing and obtaining original target data; Step S13: screening and cleaning the original target data to obtain a cleaned target data set; Step S14: extracting the cleaned target data set key frames and marking the defect areas to construct a target defect data set; Step S15: Divide the target defect data set into a target defect training set, a target defect verification set, and a target defect test set according to a ratio of 8:1:
1.
3. The defect target detection method based on improved YOLOv5 according to claim 1, characterized in that: The network framework of improving the YOLOv5 model described in step S2 specifically includes the following operations: Operation 1: Adjust the hierarchical structure of the head part of the YOLOv5 model, integrate shallow semantic information, add a small target detection layer and a small target anchor frame group; Operation 2: Replace the C3 module of the YOLOv5 model with the HWAttention attention module; Operation three: When connecting the detection layer, the multi-scale feature extraction module is combined with the HWAttention module to form the TC_HWA module.
4. The defect target detection method based on improved YOLOv5 according to claim 3, characterized in that: The specific operation of the operation one is: first adjust the hierarchical structure of the head part of the YOLOv5 model, fuse the shallow semantic information, add a group of small target detection layers P2 on the basis of the three detection layers, and then add a P2 upsampling fusion module after two rounds of upsampling fusion modules of the detection layer P3 and the detection layer P4, and connect it to the small target detection layer P2; at the same time, set a small target anchor frame group corresponding to the scale of the small target detection layer P2.
5. The defect target detection method based on improved YOLOv5 according to claim 1, characterized in that: The step S32 specifically includes the following steps: Step S321: average pooling and maximum pooling are used to reduce the dimension of the shallow features of the target, compress the dimension of the reduced dimension channel, and obtain the reduced dimension smoothing features and reduced dimension significant features of the shallow features of the target respectively; Step S322: splicing the dimension-reduced smooth features and the dimension-reduced significant features in the channel dimension to obtain the fused target shallow features; Step S323: Keeping the spatial dimension unchanged, calling 5×5 convolution to compress the fused target shallow features into an attention score map, and calling the Sigmoid activation function to weight the attention score map to obtain the attention weight; Step S324: Add the attention weight to the target shallow feature map.
6. A defect target detection method based on improved YOLOv5 according to claim 2, characterized in that: The step S4 specifically comprises the following steps: Step S41: training the HT-YOLO model according to the target defect image training set; Step S42: input the target defect image test set into the trained HT-YOLO model, and obtain the evaluation results according to the set target detection tasks and key indicators; Step S43: adjusting the hyperparameters of the HT-YOLO model according to the evaluation results to obtain an adjusted HT-YOLO model; Step S44: Obtain the target defect according to the adjusted HT-YOLO model.
7. A defect target detection system based on improved YOLOv5, used to execute the defect target detection method based on improved YOLOv5 as claimed in claim 1; characterized in that, It includes collection unit, improvement unit, training unit and evaluation and detection unit; The acquisition unit is used to acquire image data of the target to be detected during the production process and annotate it to obtain a target defect data set; The improvement unit is used to improve the network framework of the YOLOv5 model to obtain the HT-YOLO model; The training unit is used to train the HT-YOLO model according to the target defect data set to obtain a trained HT-YOLO model; The evaluation and detection unit is used to evaluate the trained HT-YOLO model according to the target defect feature data set, and adjust the HT-YOLO model according to the evaluation result to identify the target defect.
8. An electronic device, characterized in that: It comprises a memory and a processor; a computer program is stored in the memory; when the computer program is executed on the processor, the defect target detection method based on the improved YOLOv5 as described in any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions; when the computer instructions are executed on the electronic device as described in claim 8, the defect target detection method based on the improved YOLOv5 as described in any one of claims 1-6 is implemented.
Citation Information
Patent Citations
Optical fiber surface defect detection method and device
CN116071294A
Steel surface defect detection method based on improved YOLOv5
CN117036244A