Defect target detection method, system and equipment based on improved YOLOv5 and medium
By improving the YOLOv5 model, adding a tiny object detection layer and attention module, the HT-YOLO model is formed, which solves the problems of low detection rate and high error detection rate in the existing technology, and achieves efficient and real-time detection of industrial oil seal defects.
Patent Information
- Application Number
- CN202510411635.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2045-04-02
AI Technical Summary
When existing machine vision detection technology detects defects such as small scratches, damage and bubbles in industrial oil seals, there are problems such as low detection rate, high error detection rate, and difficult to balance real-time and detection accuracy.
Improve the network framework of the YOLOv5 model, add a micro-object detection layer and a micro-object anchor frame group, and combine the HWAttention attention module and the Tconv module to form an HT-YOLO model for efficient detection of defects on the oil seal surface in real time.
It significantly improves the model's detection ability of small targets, improves the accuracy and reliability of defect recognition, and achieves efficient and real-time detection in industrial assembly line scenarios.
Smart Images

Figure CN119919646A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of defect target detection, and in particular to a defect target detection method, system, device and medium based on improved YOLOv5. Background Art
[0002] Industrial oil seals are key sealing components in mechanical equipment, and their defects directly affect the safety and service life of the equipment. Therefore, the detection of oil seal defects has become an important part of industrial manufacturing and quality control. Traditional detection methods are mainly based on manual visual inspection and simple mechanical measurement. Although obvious defects can be found to a certain extent, they have problems such as low efficiency, strong subjectivity and low detection accuracy.
[0003] With the development of industrial automation technology, machine vision-based detection methods have gradually been introduced into oil seal defect detection. Using image acquisition equipment and image processing algorithms to automatically detect the surface of oil seals can effectively improve detection efficiency and accuracy. However, existing machine vision detection technology still faces some challenges in practical applications, mainly including the following aspects: 1) The defects of industrial oil seals usually appear as small scratches, breakages, bubbles, etc. These defects are small in size and have blurred edges. In a complex industrial background, the detection algorithm is not capable of extracting their features, which can easily lead to missed detection or false detection.
[0004] 2) The surface gloss of oil seals is usually high, which is easily affected by ambient light and equipment reflection, resulting in low contrast and high noise in the collected images. Traditional algorithms find it difficult to accurately segment the target and background in this case.
[0005] 3) Existing target detection algorithms have difficulty in striking a balance between real-time performance and detection accuracy. Especially in high-speed operation scenarios of industrial assembly lines, traditional detection algorithms are difficult to meet the requirements for efficient and real-time detection.
[0006] At present, deep learning technology has made significant progress in the field of computer vision, especially the object detection algorithm based on convolutional neural network (CNN), which has been widely used in the field of industrial inspection due to its powerful feature extraction ability and end-to-end detection method. Among them, the YOLO (You Only Look Once) series of algorithms have become an important tool for object detection due to its efficient real-time performance and excellent detection performance. However, the standard YOLO model still has the following shortcomings in the task of industrial oil seal defect detection: 1) The feature extraction capability of small-scale targets (such as tiny scratches or bubbles) is limited and easily interfered by complex background noise, resulting in reduced detection accuracy; 2) The model’s generalization ability is insufficient when faced with diverse defect types, and the detection effects of different types of defects are uneven; 3) The traditional YOLO model has limitations in feature map fusion and multi-scale detection capabilities, and it is difficult to take into account both detailed information and global features, especially when detecting complex defects on the surface of oil seals. Summary of the invention
[0007] Aiming at the problems that the existing detection methods have low detection rate for small-sized targets and are difficult to detect complex surface defects, the present invention proposes a defect target detection method, system, device and medium based on improved YOLOv5; the method firstly collects image data of the target to be detected in the production process, and performs image preprocessing according to the surface defect characteristics of the target to be detected, then extracts key frames and marks defect areas, and constructs a defect detection data set; then improves the network framework of the YOLOv5 model to obtain the HT-YOLO model; finally, the HT-YOLO model is used for training, so as to realize real-time and efficient detection of surface defects of the target to be detected, and improve the detection accuracy and speed.
[0008] The specific implementation contents of the present invention are as follows: A defect target detection method based on improved YOLOv5 specifically includes the following steps: Step S1: Collect image data of the target to be inspected during the production process and annotate it to obtain a target defect data set; Step S2: Improve the network framework of the YOLOv5 model to obtain the HT-YOLO model; Step S3: training the HT-YOLO model according to the target defect training set to obtain a trained HT-YOLO model; Step S4: Evaluate the trained HT-YOLO model based on the target defect test set, and adjust the HT-YOLO model based on the evaluation results to identify the target defect.
[0009] In order to better implement the present invention, further, the step S1 specifically includes the following steps: Step S11: collecting target image data according to the set collection time period; Step S12: online search and acquisition of data related to the target to be detected, and constructing and obtaining original target data; Step S13: Filter and clean the original target data, and divide the cleaned data into a target defect training set, a target defect verification set, and a target defect test set in a ratio of 8:1:1.
[0010] In order to better implement the present invention, further, the network framework of the improved YOLOv5 model in step S2 specifically includes the following operations: Operation 1: Adjust the hierarchical structure of the head part of the YOLOv5 model, integrate shallow semantic information, add a small target detection layer and a small target anchor frame group; Operation 2: Replace the C3 module of the YOLOv5 model with the HWAttention attention module; Operation 3: When connecting the detection layer, the Tconv module is combined with the HWAttention module to form the TC_HWA module.
[0011] In order to better realize the present invention, further, the specific operation of the operation one is: first adjust the hierarchical structure of the head part of the YOLOv5 model, fuse the shallow semantic information, add a group of small target detection layers P2 on the basis of the three detection layers, and then add a P2 upsampling fusion module after two rounds of upsampling fusion modules of the detection layer P3 and the detection layer P4, and connect it to the small target detection layer P2; at the same time, set a small target anchor frame group corresponding to the scale of the small target detection layer P2.
[0012] In order to better implement the present invention, further, the step S3 specifically includes the following steps: Step S31: Obtaining shallow features of the target based on the acquired target defect feature map according to the small target detection layer and the small target anchor frame group; Step S32: reduce the dimension of the shallow features of the target, and splice them in the channel dimension to obtain the fused shallow features of the target. According to the fused shallow features of the target, generate attention weights and weight them to the shallow feature map of the target. Step S33: Generate channel attention weights according to channel attention, reorganize the target shallow feature map, and residually connect the reorganized target shallow feature map to obtain the final output feature map, which is then converted into detection result information through the detection layer.
[0013] In order to better implement the present invention, further, the step S32 specifically includes the following steps: Step S321: average pooling and maximum pooling are used to reduce the dimension of the shallow features of the target, compress the dimension of the reduced dimension channel, and obtain the reduced dimension smoothing features and reduced dimension significant features of the shallow features of the target respectively; Step S322: splicing the dimension-reduced smooth features and the dimension-reduced significant features in the channel dimension to obtain the fused target shallow features; Step S323: Keeping the spatial dimension unchanged, call 5×5 convolution to compress the fused target shallow features into an attention score map, and call the Sigmoid activation function to weight the score map to obtain the attention weight; Step S324: Add the attention weight to the target shallow feature map.
[0014] In order to better implement the present invention, further, the step S33 specifically includes the following steps: Step S331: calling the attention module to generate channel attention weights; Step S332: Sort the channels according to the channel attention weights to obtain a rearranged channel sequence and attention weights corresponding to the channel sequence; Step S333: reorganizing the target shallow feature map according to the channel sequence, and weighting the reorganized target shallow feature map; Step S334: Divide the weighted target shallow feature map in the channel dimension according to the ratio of 1:2:1, and perform convolution extraction according to the size of the weight; Step S335: The target shallow feature map is fused with the current output feature map through residual connection, and the fused feature map is sorted through the convolution operation channel to obtain the final output feature map, which is then converted into detection result information through the detection layer.
[0015] In order to better implement the present invention, further, the step S4 specifically includes the following steps: Step S41: training the HT-YOLO model according to the target defect training set; Step S42: input the target defect test set into the trained HT-YOLO model, and obtain the evaluation result according to the set target detection task and key indicators; Step S43: adjusting the hyperparameters of the HT-YOLO model according to the evaluation results to obtain an adjusted HT-YOLO model; Step S44: Obtain the target defect according to the adjusted HT-YOLO model.
[0016] Based on the above-mentioned defect target detection method based on improved YOLOv5, in order to better realize the present invention, further, a defect target detection system based on improved YOLOv5 is proposed, including an acquisition unit, an improvement unit, a training unit, and an evaluation and detection unit; The acquisition unit is used to acquire image data of the target to be detected during the production process and annotate it to obtain a target defect data set; The improvement unit is used to improve the network framework of the YOLOv5 model to obtain the HT-YOLO model; The training unit is used to train the HT-YOLO model according to the target defect training set to obtain a trained HT-YOLO model; The evaluation and detection unit is used to evaluate the trained HT-YOLO model according to the target defect test set, and adjust the HT-YOLO model according to the evaluation result to identify the target defect.
[0017] Based on the above-proposed defect target detection method based on improved YOLOv5, in order to better realize the present invention, further, an electronic device is proposed, including a memory and a processor; a computer program is stored on the memory; when the computer program is executed on the processor, the above-proposed defect target detection method based on improved YOLOv5 is implemented.
[0018] Based on the above-proposed defect target detection method based on improved YOLOv5, in order to better realize the present invention, further, a computer-readable storage medium is proposed, on which computer instructions are stored; when the computer instructions are executed on the above-mentioned electronic device, the above-mentioned defect target detection method based on improved YOLOv5 is implemented.
[0019] The present invention has the following beneficial effects: (1) The present invention introduces a small target detection layer and combines it with the P2 upsampling fusion block to effectively fuse shallow features and set up a special small target anchor frame group, which significantly enhances the model's detection capability for small targets and improves the accuracy and reliability of defect recognition.
[0020] (2) The present invention designs a lightweight attention mechanism HWA, which extracts the fusion features of global smooth features and local significant features by performing average pooling and maximum pooling dimensionality reduction processing on the feature map, and then generates an attention score and weights the input feature map; it effectively improves the expressiveness of the feature map, while optimizing the computational complexity and maintaining a high detection accuracy; by replacing the original C3 module of YOLO, the computational complexity of the model is significantly reduced, while the detection performance is not affected.
[0021] (3) The present invention introduces a multi-scale feature extraction module, which applies convolution operations of different sizes to the feature maps that have been enhanced and filtered by the channel attention mechanism according to the feature weights. This not only effectively enhances the channel features, but also realizes the extraction of multi-scale features, thereby significantly improving the feature expression ability of the model and the accuracy of target detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 A flowchart of the defect detection method based on improved YOLOv5 provided by the present invention.
[0023] Figure 2 The HT-YOLO model provided by the present invention.
[0024] Figure 3 This is a diagram of the HWA module provided by the present invention.
[0025] Figure 4 This is a diagram of the Tconv module provided by the present invention.
[0026] Figure 5 A schematic diagram of an interface provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0027] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. It should be understood that the described embodiments are only part of the embodiments of the present invention, not all of the embodiments, and therefore should not be regarded as limiting the scope of protection. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technical personnel in this field without making creative work are within the scope of protection of the present invention.
[0028] In the description of the present invention, it should be noted that, unless otherwise clearly specified and limited, the terms "disposed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be an indirect connection through an intermediate medium, or it can be the internal communication of two elements. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0029] Embodiment 1: This embodiment proposes a defect target detection method based on improved YOLOv5, which specifically includes the following steps.
[0030] Step S1: Collect image data of the target to be inspected during the production process and annotate it to obtain a target defect data set; The step S1 specifically includes the following steps: Step S11: collecting the image data of the target to be detected according to the set collection time period; Step S12: online search and acquisition of data related to the target to be detected, and constructing and obtaining original target data; Step S13: Filter and clean the original target data, and divide the cleaned data into a target defect training set, a target defect verification set, and a target defect test set in a ratio of 8:1:1.
[0031] Step S2: Improve the network framework of the YOLOv5 model to obtain the HT-YOLO model; The network framework of improving the YOLOv5 model described in step S2 specifically includes the following operations: Operation 1: Adjust the hierarchical structure of the head part of the YOLOv5 model, integrate shallow semantic information, add a small target detection layer and a small target anchor frame group; Operation 2: Replace the C3 module of the YOLOv5 model with the HWAttention attention module; Operation 3: When connecting the detection layer, the Tconv module is combined with the HWAttention module to form the TC_HWA module.
[0032] The specific operation of the operation one is: first adjust the hierarchical structure of the head part of the YOLOv5 model, fuse the shallow semantic information, add a group of small target detection layers P2 on the basis of the three detection layers, and then add a P2 upsampling fusion module after two rounds of upsampling fusion modules of the detection layer P3 and the detection layer P4, and connect it to the small target detection layer P2; at the same time, set a small target anchor frame group corresponding to the scale of the small target detection layer P2.
[0033] Step S3: training the HT-YOLO model according to the target defect training set to obtain a trained HT-YOLO model; The step S3 specifically comprises the following steps: Step S31: Obtain shallow features of the target according to the acquired target feature map based on the small target detection layer and the small target anchor frame group; Step S32: reduce the dimension of the shallow features of the target, and splice them in the channel dimension to obtain the fused shallow features of the target. According to the fused shallow features of the target, generate attention weights and weight them to the shallow feature map of the target. The step S32 specifically includes the following steps: Step S321: average pooling and maximum pooling are used to reduce the dimension of the shallow features of the target, compress the dimension of the reduced dimension channel, and obtain the reduced dimension smoothing features and reduced dimension significant features of the shallow features of the target respectively; Step S322: splicing the dimension-reduced smooth features and the dimension-reduced significant features in the channel dimension to obtain the fused target shallow features; Step S323: Keeping the spatial dimension unchanged, call 5×5 convolution to compress the fused target shallow features into an attention score map, and call the Sigmoid activation function to weight the score map to obtain the attention weight; Step S324: Add the attention weight to the target shallow feature map.
[0034] Step S33: Generate channel attention weights based on channel attention, reorganize the target shallow feature map, and residually connect the reorganized target shallow feature map to obtain the final output feature map, which is then converted into detection result information through the P2, P3, P4, and P5 detection layers; the detection result information includes the final output detection box position and category information, which is the final stage of detection.
[0035] The step S33 specifically includes the following steps: Step S331: calling the attention module to generate channel attention weights; Step S332: Sort the channels according to the channel attention weights to obtain a rearranged channel sequence and attention weights corresponding to the channel sequence; Step S333: reorganizing the target shallow feature map according to the channel sequence, and weighting the reorganized target shallow feature map; Step S334: Divide the weighted shallow feature map of the target in the channel dimension according to the ratio of 1:2:1, and perform convolution extraction according to the size of the weight; Step S335: The target shallow feature map is fused with the current output feature map through a residual connection, and the fused feature map is sorted through a convolution operation channel to obtain a final output feature map set.
[0036] Step S4: Evaluate the trained HT-YOLO model based on the target defect test set, and adjust the HT-YOLO model based on the evaluation results to identify the target defect.
[0037] The step S4 specifically comprises the following steps: Step S41: training the HT-YOLO model according to the target defect training set; Step S42: input the target defect test set into the trained HT-YOLO model, and obtain the evaluation result according to the set target detection task and key indicators; Step S43: adjusting the hyperparameters of the HT-YOLO model according to the evaluation results to obtain an adjusted HT-YOLO model; Step S44: Obtain the target defect according to the adjusted HT-YOLO model.
[0038] Working principle: This embodiment first collects image data of the target to be detected during the production process through a high-definition camera, performs image preprocessing based on the surface defect characteristics of the target to be detected, extracts key frames and marks defect areas, and constructs a target defect detection data set. Then, the network framework is improved based on the YOLOv5 model: 1) Add a small target detection layer. 2) Use the newly added spatial attention module HWA to enhance the extracted features. 3) Combined with the channel attention EffectiveSELayer, construct a multi-scale feature extraction module Tconv, and replace the original YOLOv5 module C3 with HWA. Finally, the improved YOLOv5 model is used to train the data set. Through this embodiment, real-time and efficient detection of surface defects of the target to be detected is achieved, and the detection accuracy and speed are improved at the same time.
[0039] Embodiment 2: This embodiment is described in detail with a specific embodiment based on the above embodiment 1.
[0040] Step S1: using a high-definition camera to collect image data of industrial oil seals during the production process, and annotating and enhancing the collected data to generate a data set; Step S2: Improve the network framework based on the YOLOv5 model to obtain an improved HT-YOLO model; the specific improvements of the HT-YOLO model are as follows: 1) Adjust the hierarchical structure of the head part, integrate shallow semantic information, add a small object detection layer, and increase the small object anchor frame group; 2) Use the HWAttention module to replace the C3 module of the original YOLOv5 model; 3) When connecting the detection layer, the Tconv module is combined with the HWAttention module to form the TC_HWA module.
[0041] Step S3: Use the HT-YOLO model to train and test the processed oil seal dataset for defect detection.
[0042] Furthermore, based on the original three-layer detection layer, a group of P2 small target detection layers is added. After the original P3 and P4 upsampling fusion modules, the P2 upsampling fusion module is continued to be added to obtain shallow features and connected to the detection layer. At the same time, the small target anchor frame group corresponding to the P2 scale is set: [5, 68, 14, 15, 11] to improve the detection accuracy of small defect targets on oil seals.
[0043] Furthermore, the HWAttention module enhances the target features in the spatial dimension as follows: First, the input features are compressed and reduced in channel dimension through average pooling and maximum pooling, respectively, to obtain the reduced dimension smoothing features and reduced dimension significant features of the input features, and then these two parts of features are spliced in the channel dimension. The fused features contain both global average features and significant features.
[0044] Then, a 5×5 convolution is used to compress the fused features into an attention score map, with the spatial dimension unchanged and the number of channels being 1. The score map is then weighted using the Sigmoid activation function to obtain the attention weight.
[0045] Finally, the obtained weights are added to the module input feature map, thereby strengthening the input feature map at specific locations.
[0046] Furthermore, the Tconv module performs feature enhancement in the channel dimension and extracts and fuses multi-scale features, as follows: The input feature map first generates channel attention weights through the channel attention module (EffectiveSELayer). Then, the channels are sorted according to these weights to obtain a rearranged channel sequence and its corresponding attention weights. According to this sorting order, the channels of the input feature map are reorganized, and weighting is applied to the reorganized feature map to enhance the expressiveness of important channels. At this time, the channel dimensions of the feature map are sorted according to the weight size; Next, the feature map is divided in the channel dimension in a ratio of 1:2:1. The three parts after division are extracted using 1×1, 3×3 and 5×5 convolutions in order of weight from low to high, and the number of channels remains unchanged during the extraction process. The output feature map of the convolution operation is spliced in the channel dimension. At this time, the input feature map is fused with the current output feature map through residual connection to further enhance the expression ability of the model; Finally, the fused feature map is channel-sorted through a 1×1 convolution operation to obtain the final output feature map.
[0047] Furthermore, step S3 specifically includes the following operations: 1) Input the divided training set into HT-YOLO improved based on YOLO for training; 2) Input the test set into the trained HT-YOLO model to test its detection ability on new data. Through the target detection task on the test set, the performance of the HT-YOLO model is evaluated, focusing on testing its key indicators such as accuracy, recall, mean average precision (mAP), and F1 score in industrial oil seal defect detection, to ensure that the model can maintain high detection accuracy and reliability in a variety of oil seal defect samples; 3) Based on the evaluation results, the hyperparameters of the HT-YOLO model are adjusted to improve the detection effect of the model on the test set.
[0048] Working principle: This embodiment optimizes the network structure of YOLOv5 to enhance the model's detection capabilities for tiny targets and complex defects, while improving the real-time and robustness of detection. According to the characteristics of industrial oil seal detection tasks, a refined target detection process is designed, which can effectively improve the detection accuracy and efficiency of industrial oil seal defects and meet the needs of efficient and accurate detection in industrial production.
[0049] The other parts of this embodiment are the same as those of the above-mentioned embodiment 1, and thus will not be described in detail.
[0050] Embodiment 4: This embodiment is based on any one of the above embodiments 1 to 2. Figure 5 As shown, an interface for collecting surface defects of industrial oil seals is used as an example for detailed explanation.
[0051] Step S1: Data collection and processing.
[0052] First, high-definition monitoring equipment is installed in the production environment of industrial oil seals. At the same time, the monitoring equipment needs to collect data at different time periods to ensure that the data covers the daily production peak period, off-peak period, and equipment maintenance and shutdown conditions. These data include pictures and videos, and the content should cover normal production status, production anomalies such as equipment failure, oil seal quality problems, and other scenarios that may affect production. Through data collection in different weather and different production environments, the robustness and adaptability of subsequent data analysis models are improved. Secondly, using tools such as web crawlers, online search and obtain more pictures, videos and technical literature data related to industrial oil seals to enrich the diversity of data sets. The crawled content includes oil seal products of different brands, models, and materials, and their use effects and failure modes in different application scenarios. The acquired data is screened and cleaned to remove fuzzy, repeated, and irrelevant images and content to ensure the high quality of the collected data. After the data collection is completed, the original data needs to be screened and cleaned. First, remove blurred or damaged pictures and videos to ensure the accuracy of the data. Then, data enhancement processing is performed to increase the diversity of data samples by rotating, scaling, translating, cropping, adding noise, etc. Then, we use annotation tools to annotate the data, such as the damage types of oil seals (such as cracks, aging, oil leakage, etc.), as well as different working conditions (such as normal, wear, failure, etc.). Finally, the cleaned and enhanced data set is divided into training set, validation set and test set according to 8:1:1 to ensure the comprehensiveness and balance of the data.
[0053] Step S2: Optimize the YOLO network model.
[0054] Based on the YOLOv5 algorithm, we can get the improved HT-YOLO model. Figure 2 The specific optimization details are as follows: First, the original YOLOv5 model cannot effectively capture these subtle targets due to the large receptive field of its detection layer (P3, P4, P5) when dealing with tiny defects on the surface of the oil seal. Especially at lower resolutions, tiny defects are usually difficult to distinguish from the background, resulting in missed detection or false detection. To meet this challenge, we made targeted improvements to the YOLOv5 architecture and added a dedicated tiny target detection layer P2, which has a smaller receptive field and can more accurately capture tiny defects on the surface of the oil seal. Relatively speaking, in view of the characteristics of tiny defects on the surface of the oil seal, we redesigned the anchor frame to make it more suitable for detecting small-scale and tiny defects. These small-sized anchor frames help to focus on the target area more accurately during the detection process, reducing the problem of decreased detection accuracy due to mismatch of the anchor frame scale. The introduction of this layer effectively enhances the sensitivity to tiny targets and improves the performance of the model in tiny target detection.
[0055] At the same time, in order to reduce the additional parameters and calculation amount caused by adding the detection layer, the new module HWA is used to replace the C3 module of the original model. Figure 3 .
[0056] The C3 module in the original YOLOv5 model is composed of multiple stacked convolutional layers, and the convolution operation itself is computationally intensive, which will increase the time and resource consumption during training and inference. Moreover, C3 focuses on global feature extraction and cannot capture local key features such as target details or minor changes. This may make the C3 module not achieve optimal results in fine-grained tasks, especially when small objects need to be accurately located and recognized. Therefore, we introduced a lightweight attention module HWA, which can compress the channel dimension of the feature map to half of the original while extracting smooth features and salient features by using channel average pooling and channel maximum pooling, reducing the computational complexity of subsequent operations. After obtaining the compressed features, a 5×5 convolution is used to extract the attention score map, compress the feature map to a size of 1×H×W, and then the score map is weighted by Sigmoid to map its value to the range of [0,1] to obtain the spatial attention weight of the feature map. Finally, the obtained weight map is weighted for each channel of the input feature map, increasing the attention to important features in the spatial dimension while suppressing the representation of unimportant features. The lightweight attention HWA module can better capture local features and enhance their representation while significantly reducing the number of parameters and computations, while taking into account the extraction of global features and improving the accuracy of feature extraction.
[0057] Finally, the Tconv module is introduced into the head network. Figure 4 After connecting the HWA module of the detection layer, add the Tconv module to enhance the feature representation of the model and improve the detection effect.
[0058] First, the channel dimension of the input feature map is enhanced by the channel attention module EffectiveSELayer. The weight values calculated by the channel attention mechanism are used to weightedly screen each channel, and the feature maps are reordered according to the weight. After the feature maps are sorted from low to high by weight, they are divided into three parts in a ratio of 1:2:1 to distinguish channels of different importance. Then, convolution operations of different sizes are applied to the three groups of feature maps, where low-weight channels use 1×1 convolution kernels for feature extraction, medium-weight channels use 3×3 convolution kernels, and high-weight channels use 5×5 convolution kernels for deep feature extraction. The core of this strategy is that high-weight channels correspond to more important features, so using larger-sized convolution kernels can capture richer and more complex contextual information and improve the model's ability to express important features. Finally, the feature maps processed by convolution kernels of different sizes are spliced in the channel dimension and merged into a comprehensive feature map. The channel information is further sorted through point convolution, and the optimized feature representation is finally output. This method combines channel attention with scaled convolution operations to achieve refined processing of important channels of feature maps, thereby improving the network's feature extraction capabilities in complex tasks.
[0059] Step S3: Train the improved YOLO model.
[0060] The training process of the improved YOLO model first relies on the constructed oil seal defect dataset, which contains a variety of defect samples and is accurately annotated. During the training process, a phased strategy is adopted. First, the weight of the model is initialized to accelerate the training process and improve the convergence speed of the model. Then, the performance of the model is evaluated in real time using the validation set to ensure that the model can maintain good generalization ability under different data distributions. According to the evaluation results on the validation set, the training hyperparameters, such as learning rate, batch size, and number of iterations, are adjusted in real time to ensure that the model is gradually optimized during the training process to avoid overfitting or underfitting. In each training cycle, the learning rate adopts a dynamic adjustment strategy, such as using methods such as CosineAnnealing to gradually reduce the learning rate according to the performance of the model, thereby accelerating the convergence of the model and improving the final performance. The batch size and number of iterations are adjusted according to the size of the training set and the convergence speed of the model to ensure that the training process is both efficient and robust. In addition, in order to prevent overfitting, an early stopping strategy is also adopted in the training process. When the performance of the model on the validation set is no longer significantly improved, the training is automatically stopped to avoid wasting computing resources and ensure the generalization ability of the model. These sophisticated training strategies ensure that the improved YOLO model can achieve the best detection performance in complex oil seal defect detection tasks.
[0061] Step S4: Model verification and scoring.
[0062] When evaluating model performance on the validation set, detection accuracy (Precision, P), recall (Recall, R), average precision (MeanAveragePrecision, mAP) and fitness (Fitness, Fit) are mainly used as evaluation indicators.
[0063] Among them, the detection accuracy measures the proportion of targets that are actually defects among all targets predicted by the model to be defects, reflecting the reliability of the model in defect identification. The calculation formula is as follows: The recall rate indicates the proportion of defect targets successfully detected by the model among all actual defect targets, which measures the comprehensiveness of the model. The calculation formula is as follows: The average precision comprehensively considers the precision and recall performance of the model under different IoU thresholds. It is usually calculated by the area under the precision-recall curve. It is an important indicator for evaluating the global performance of the model. The calculation formula is as follows: In addition, fitness, as a comprehensive indicator, comprehensively evaluates the precision and comprehensiveness of the model and can measure the applicability of the model in practical applications. Fitness is usually calculated by a weighted combination of precision, recall, and average precision under different thresholds to reflect the balance between the real-time and comprehensiveness of the model in an industrial environment. The calculation formula is as follows: In the formula, TP represents the number of samples correctly classified as positive (actually defects, model detection results are defects); FP represents the number of samples incorrectly classified as positive (actually normal, model detection results are defects); FN represents the number of samples incorrectly classified as negative (actually defects, model detection results are normal). N represents the number of defect categories. i Represents the weight coefficient of the i-th indicator. mAP@50 and mAP@95 represent the average precision when the threshold is set to 50% and 95%, respectively.
[0064] Through these comprehensive indicators, the actual performance of the improved YOLOv5 model in oil seal defect detection can be comprehensively evaluated to ensure its optimization in terms of precision, recall, and applicability.
[0065] like Figure 5As shown, ① is the front inspection picture of industrial oil seal, ② is the back inspection picture of industrial oil seal, ③ is the inside inspection picture of industrial oil seal, ④ is the outside inspection picture of industrial oil seal, and ⑤ is the size inspection picture of industrial oil seal; The time consumed to complete the front detection is 615ms, the time consumed to complete the back detection is 240ms, the time consumed to complete the inner detection is 151ms, the time consumed to complete the outer detection is 397ms, and the time consumed to complete the size detection is 109ms.
[0066] The PPM value is 84159, which means the number of NG products that may appear per million industrial oil seal products based on the qualified rate.
[0067] Figure 5 The number 53 in the upper right corner indicates the overall inspection speed, including the speed of the glass plate, roller, conveyor belt, and handwheel; TG2-40-62.05-9.5 / 12 C7808 is the industrial oil seal model, and batch 1 indicates the current batch. The same industrial oil seal may be made by multiple molds, for example, mold 2 can be switched to batch 2.
[0068] Figure 5 The X-axis of the curve graph is the quantity, and the Y-axis is the size. 62.48 is the maximum outer diameter allowed for industrial oil seals, and 62.34 is the minimum outer diameter allowed for industrial oil seals. OK, NG, and ER below the curve graph represent log images, indicating whether the relevant log images are enabled during the inspection process. For example, if NG is checked, the NG images generated during the inspection process will be stored in a fixed folder for easy viewing. The bar graph is a normal distribution graph of industrial oil seals. For example, there are 4437 industrial oil seals with a size of 62.44.
[0069] The other parts of this embodiment are the same as any of the above-mentioned embodiments 1 to 2, and thus will not be described in detail.
[0070] Embodiment 4: Based on any one of the above-mentioned embodiments 1 to 3, this embodiment proposes a defect target detection system based on improved YOLOv5, including a collection unit, an improvement unit, a training unit, and an evaluation and detection unit; The acquisition unit is used to acquire image data of the target to be detected during the production process and annotate it to obtain a target defect data set; The improvement unit is used to improve the network framework of the YOLOv5 model to obtain the HT-YOLO model; The training unit is used to train the HT-YOLO model according to the target defect training set to obtain a trained HT-YOLO model; The evaluation and detection unit is used to evaluate the trained HT-YOLO model according to the target defect test set, and adjust the HT-YOLO model according to the evaluation result to identify the target defect.
[0071] This embodiment also proposes an electronic device, including a memory and a processor; a computer program is stored on the memory; when the computer program is executed on the processor, the above-mentioned defect target detection method based on improved YOLOv5 is implemented.
[0072] This embodiment also proposes a computer-readable storage medium, on which computer instructions are stored; when the computer instructions are executed on the above-mentioned electronic device, the above-mentioned defect target detection method based on improved YOLOv5 is implemented.
[0073] The other parts of this embodiment are the same as any one of the above-mentioned embodiments 1 to 3, and thus will not be described in detail.
[0074] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any form. Any simple modification or equivalent change made to the above embodiment based on the technical essence of the present invention shall fall within the protection scope of the present invention.
Claims
1. A defect target detection method based on improved YOLOv5, characterized in that: The specific steps include: Step S1: collecting image data of the target to be detected during the production process, annotating to obtain a target defect data set, and dividing the target data set into a target defect training set, a target defect test set, and a target defect verification set; Step S2: Improve the network framework of the YOLOv5 model to obtain the HT-YOLO model; Step S3: training the HT-YOLO model according to the target defect training set to obtain a trained HT-YOLO model; Step S4: Evaluate the trained HT-YOLO model based on the target defect test set, and adjust the HT-YOLO model parameters according to the evaluation results to obtain the target defect.
2. A defect target detection method based on improved YOLOv5 according to claim 1, characterized in that: The step S1 specifically includes the following steps: Step S11: collecting the image data of the target to be detected according to the set collection time period; Step S12: online search and acquisition of data related to the target to be detected, and constructing and obtaining original target data; Step S13: screening and cleaning the original target data to obtain a cleaned target data set; Step S14: extracting the cleaned target data set key frames and marking the defect areas to construct a target defect data set; Step S15: Divide the target defect data set into a target defect training set, a target defect verification set, and a target defect test set according to a ratio of 8:1:
1.
3. The defect target detection method based on improved YOLOv5 according to claim 1, characterized in that: The network framework of improving the YOLOv5 model described in step S2 specifically includes the following operations: Operation 1: Adjust the hierarchical structure of the head part of the YOLOv5 model, integrate shallow semantic information, add a small target detection layer and a small target anchor frame group; Operation 2: Replace the C3 module of the YOLOv5 model with the HWAttention attention module; Operation 3: When connecting the detection layer, the Tconv module is combined with the HWAttention module to form the TC_HWA module.
4. The defect target detection method based on improved YOLOv5 according to claim 3, characterized in that: The specific operation of the operation one is: first adjust the hierarchical structure of the head part of the YOLOv5 model, fuse the shallow semantic information, add a group of small target detection layers P2 on the basis of the three detection layers, and then add a P2 upsampling fusion module after two rounds of upsampling fusion modules of the detection layer P3 and the detection layer P4, and connect it to the small target detection layer P2; at the same time, set a small target anchor frame group corresponding to the scale of the small target detection layer P2.
5. The defect target detection method based on improved YOLOv5 according to claim 3, characterized in that: The step S3 specifically comprises the following steps: Step S31: Obtain shallow features of the target according to the acquired target feature map based on the small target detection layer and the small target anchor frame group; Step S32: reduce the dimension of the shallow features of the target, and splice them in the channel dimension to obtain the fused shallow features of the target. According to the fused shallow features of the target, generate attention weights and weight them to the shallow feature map of the target. Step S33: Generate channel attention weights according to the channel attention, reorganize the target shallow feature map, and residually connect the reorganized target shallow feature map to obtain the final output feature map.
6. A defect target detection method based on improved YOLOv5 according to claim 5, characterized in that: The step S32 specifically includes the following steps: Step S321: average pooling and maximum pooling are used to reduce the dimension of the shallow features of the target, compress the dimension of the reduced dimension channel, and obtain the reduced dimension smoothing features and reduced dimension significant features of the shallow features of the target respectively; Step S322: splicing the dimension-reduced smooth features and the dimension-reduced significant features in the channel dimension to obtain the fused target shallow features; Step S323: Keeping the spatial dimension unchanged, calling 5×5 convolution to compress the fused target shallow features into an attention score map, and calling the Sigmoid activation function to weight the attention score map to obtain the attention weight; Step S324: Add the attention weight to the target shallow feature map.
7. The defect target detection method based on improved YOLOv5 according to claim 5, characterized in that: The step S33 specifically includes the following steps: Step S331: calling the attention module to generate channel attention weights; Step S332: Sort the channels according to the channel attention weights to obtain a channel sequence and attention weights corresponding to the channel sequence; Step S333: reorganizing the target shallow feature map according to the channel sequence, and weighting the reorganized target shallow feature map; Step S334: Divide the weighted shallow feature map of the target in the channel dimension according to the ratio of 1:2:1, and perform convolution extraction according to the size of the weight; Step S335: The target shallow feature map is fused with the current output feature map through residual connection, and the fused feature map is sorted through the convolution operation channel to obtain the final output feature map, and converted into detection result information through the detection layer.
8. The defect target detection method based on improved YOLOv5 according to claim 2, characterized in that: The step S4 specifically comprises the following steps: Step S41: training the HT-YOLO model according to the target defect image training set; Step S42: input the target defect image test set into the trained HT-YOLO model, and obtain the evaluation results according to the set target detection tasks and key indicators; Step S43: adjusting the hyperparameters of the HT-YOLO model according to the evaluation results to obtain an adjusted HT-YOLO model; Step S44: Obtain the target defect according to the adjusted HT-YOLO model.
9. A defect target detection system based on improved YOLOv5, used to execute the defect target detection method based on improved YOLOv5 as claimed in claim 1; characterized in that, It includes collection unit, improvement unit, training unit and evaluation and detection unit; The acquisition unit is used to acquire image data of the target to be detected during the production process and annotate it to obtain a target defect data set; The improvement unit is used to improve the network framework of the YOLOv5 model to obtain the HT-YOLO model; The training unit is used to train the HT-YOLO model according to the target defect data set to obtain a trained HT-YOLO model; The evaluation and detection unit is used to evaluate the trained HT-YOLO model according to the target defect feature data set, and adjust the HT-YOLO model according to the evaluation result to identify the target defect.
10. An electronic device, characterized in that: It comprises a memory and a processor; a computer program is stored in the memory; when the computer program is executed on the processor, the defect target detection method based on the improved YOLOv5 as described in any one of claims 1 to 8 is implemented.
11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions; when the computer instructions are executed on the electronic device as described in claim 10, the defect target detection method based on the improved YOLOv5 as described in any one of claims 1-8 is implemented.
Citation Information
Patent Citations
Lightweight network model based on single image super-resolution, and processing method
CN113781304A
Water surface floating object target detection system based on unmanned aerial vehicle aerial photography and improved YOLO v3
CN114937195A
Optical fiber surface defect detection method and device
CN116071294A
Steel surface defect detection method based on improved YOLOv5
CN117036244A
Spodoptera frugiperda larva target detection method based on YOLOv7 attention guidance feature optimization mechanism
CN118229956A
Cited By
Defect detection method, system and equipment based on improved YOLOv5 and medium
CN120543543A
Oil seal defect intelligent detection method based on improved YOLOv12
CN120707566A
Improved yolov12-based intelligent detection method for oil seal defects
CN120707566B
Sub-cartridge case surface defect detection method and image training and reasoning integrated platform
CN121304597A
Satellite target detection method and system based on multi-scale feature fusion enhancement
CN122336588A