Industrial part small target defect detection method based on lightweight model

By improving the YOLOv5 network to a lightweight model, the missed detection and precise positioning problems of small target defect detection of industrial parts are solved, efficient detection on embedded devices is achieved, detection performance is improved and computing resource requirements are reduced.

CN120387982APending Publication Date: 2025-07-29ZHEJIANG WANLI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510379452.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-04-01
Filing Date
2025-03-28
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing deep learning-based object detection algorithms are difficult to accurately detect small defects on industrial parts, and the computing resource requirements are high, so they cannot be effectively applied on embedded devices.

Method used

The YOLOv5 network is improved by using a lightweight model, replacing the Conv module with CS-Conv module and the C3 module with LRM module, and adding LRM and CS-Conv modules at a specific level, combining Up-sampling and Concat modules to build a lightweight small-target defect detection model, and improving detection capabilities through feature map upsampling and fusion.

Benefits of technology

It significantly reduces the amount of model parameters and calculations, improves the accuracy and efficiency of detection of small target defects in industrial parts, and is suitable for embedded equipment with limited resources. It improves detection performance by 7.1%, reduces parameter by 18.1%, extends the equipment usage time and reduces energy consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387982A_ABST
    Figure CN120387982A_ABST
Patent Text Reader

Abstract

The invention discloses an industrial part small target defect detection method based on a lightweight model, which is characterized by comprising the following steps: establishing an industrial part small target defect image set, and dividing the industrial part small target defect image set into a training set, a verification set and a test set in proportion; labeling the training set and the verification set to obtain a labeled training set and a labeled verification set; inputting the labeled training set and the labeled verification set into a lightweight small target defect detection model for training to obtain a trained lightweight small target defect detection model; inputting the test set into the trained lightweight small target defect detection model for detection, and outputting a target defect detection result; the method has the advantages that the problems of missing detection and incapability of accurate positioning in small target defect detection of industrial parts are solved, and the quality and the detection efficiency of the parts are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a defect detection method, especially a small target defect detection method for industrial parts based on a lightweight model. Background Art

[0002] The small target defect detection of industrial parts aims to discover the appearance defects of various industrial products, which is an important prerequisite for ensuring product quality and maintaining stable production.

[0003] Object detection can be divided into traditional object detection algorithms and deep learning-based object detection algorithms. Traditional object detection algorithms extract image features and generate candidate regions on the image, and then use a classifier to classify these candidate regions to achieve object detection. However, traditional object detection algorithms have relatively weak adaptability to target changes, occlusion, and complex backgrounds, and are inefficient in processing large-scale data. Therefore, deep learning-based object detection algorithms have emerged.

[0004] Currently, deep learning-based object detection algorithms include Faster R-CNN, YOLO, SSD, etc. They use deep neural networks combined with different detection strategies to achieve efficient and accurate object detection, quickly identifying and locating target objects in images or videos. However, there are still some problems: for example, they cannot meet the detection requirements for high-precision targets, have high requirements for computing resources, and the existing object detection models have a large number of parameters and are not suitable for direct use in embedded devices, so it is difficult to accurately detect small defects on industrial parts. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a small target defect detection method for industrial parts based on a lightweight model, which solves the problems of missed detection and inaccurate positioning in the small target defect detection of industrial parts.

[0006] The technical solution adopted by the present invention to solve the above technical problems is: a small target defect detection method for industrial parts based on a lightweight model, including the following steps:

[0007] Step ①, establish a small target defect image set of industrial parts, and divide the small target defect image set of industrial parts into a training set, a validation set, and a test set according to a ratio;

[0008] Step ②, label the training set and the validation set to obtain the labeled training set and the labeled validation set;

[0009] Step ③, input the labeled training set and the labeled validation set into the lightweight small target defect detection model for training to obtain the trained lightweight small target defect detection model;

[0010] Step ④, input the test set into the trained lightweight small target defect detection model for detection, and output the target defect detection result;

[0011] The described lightweight small target defect detection model is obtained by improving the existing yolov5 network. Replace the Conv module in the existing yolov5 network with the CS-Conv module, replace the C3 module in the existing yolov5 network with the LRM module, and add a module for detecting extremely small targets between the sixteenth and seventeenth layers of the existing yolov5 network;

[0012] The described module for detecting extremely small targets consists of two LRM modules with the same function and structure, two CS-Conv modules with the same function and structure, an Up-sampling module, and two Concat modules with the same function and structure. The input end of the first LRM module is connected to the output end of the sixteenth layer, the input end of the first CS-Conv module is connected to the output end of the first LRM module, the input end of the Up-sampling module is connected to the output end of the first CS-Conv module, the input ends of the first Concat module are respectively connected to the output end of the Up-sampling module and the output end of the second layer, the input end of the second LRM module is connected to the output end of the first Concat module, the input end of the second CS-Conv module is connected to the output end of the second LRM module, the input ends of the second Concat module are respectively connected to the output end of the first CS-Conv module and the output end of the second CS-Conv module, and the output end of the second Concat module is connected to the input end of the twenty-fourth layer;

[0013] Input the extremely large feature map of 128×160×160 output by the twenty-first layer, the large feature map of 256×80×80 output by the twenty-fourth layer, the medium feature map of 512×40×40 output by the twenty-seventh layer, and the small feature map of 1024×20×20 output by the thirtieth layer into the detection head for detection.

[0014] Compared with the prior art, the advantages of the present invention are as follows: The lightweight small-target defect detection model is used to streamline the network structure to the minimum, greatly reducing the number of model parameters and the amount of computation. After the sixteenth layer of the YOLOv5 network, upsampling and fusion operations are continuously performed on the feature map to further expand the feature map, improving the detection ability for small-target defects in industrial parts. At the twentieth layer, the feature maps output by the nineteenth layer and the second layer are fused to obtain a larger feature map, which can better retain detailed information, thus helping to detect smaller targets, solving the problems of missed detection and inaccurate positioning of small-target defects in industrial parts by current object detection algorithms, enabling the small-target defect detection task of industrial parts to be completed quickly and efficiently on devices with limited resources, reducing the energy consumption of the device while ensuring performance, and extending the service life of the device. And through experiments, compared with the YOLOv5 network, the method proposed in the present invention has an mAP@0.5 increase of 7.1% and a reduction in the number of parameters by 18.1%.

[0015] Further, the LRM module is composed of two CSB modules connected in sequence.

[0016] The CSB module includes two CS-Conv modules and an Add layer. After the input feature map is connected in residual with the input feature map after passing through two CS-Conv module operations in sequence, feature fusion is performed in the Add layer, and the output feature map is used as the output of the CSB module. The CSB module not only reduces the amount of computation and the number of parameters of the model, but also helps the network better learn residual information, improving the efficiency and performance of the model, so that the complexity of the model can be significantly reduced while maintaining performance; the CS-Conv module is used to reduce the amount of computation and the number of parameters of the model while maintaining good performance.

[0017] Further, the CS-Conv module includes a Conv1 layer, a Conv2 layer, and a Concat layer. In the Conv1 layer, small convolutions are used to perform convolution operations on the input feature map to generate partial feature maps; in the Conv2 layer, large convolutions are used to perform convolution operations on the partial feature maps to generate additional feature maps; in the Concat layer, the partial feature maps and the additional feature maps are connected in residual, and the output feature map is used as the output of the CS-Conv module. The initial amount of computation and the model size are reduced, and the additionally generated feature maps are fused to maintain good performance.

[0018] Further, the specific operation process of step ① is as follows: Select at least 5000 industrial part images containing small-target defects from the Internet as the industrial part small-target defect image set, and divide the industrial part small-target defect image set into a training set, a validation set, and a test set according to a ratio of 7:2:1.

[0019] Further, in step ②, the LabelMe annotation method is used to annotate the small target defects in each image in the training set and the validation set, and the annotation information is saved to obtain the annotated training set and the annotated validation set.

[0020] Further, the training process in step ③ is as follows:

[0021] Set the training parameters as follows: the batch size is set to 32, the initial learning rate is set to 0.001, the weight decay is set to 5e-4, the number of training epochs is set to 100 epochs, and the cosine annealing strategy is used as the learning rate scheduler;

[0022] Input the augmented and annotated training set into the lightweight small target defect detection model for training;

[0023] According to the loss function L total , use the Adam optimizer to iteratively train the lightweight small target defect detection model. Evaluate the lightweight small target defect detection model using the annotated validation set every 5 epochs. Stop training when the evaluation results are the same for two consecutive times to obtain the trained lightweight small target defect detection model;

[0024] The loss function L total consists of classification loss, localization loss, and confidence loss. L total = λ cls *L cls + λ loc *L loc + λ obj *L obj , where λ cls , λ loc and λ obj respectively represent hyperparameters, L cls represents the classification loss, L loc represents the localization loss, and L obj represents the confidence loss. Description of the Drawings

[0025] Figure 1 is the overall flow schematic diagram of the present invention;

[0026] Figure 2 is the network structure schematic diagram of the existing yolov5 network;

[0027] Figure 3 is the network structure schematic diagram of the lightweight small target defect detection model in the present invention;

[0028] Figure 4 is the structure schematic diagram of the LRM module in the present invention;

[0029] Figure 5 It is a schematic structural diagram of the CSB module in the present invention.

[0030] Figure 6 It is a schematic structural diagram of the CS-Conv module in the present invention;

[0031] Figure 7(a) is the original image for the small target defect detection experiment in this embodiment;

[0032] Figure 7(b) is a schematic diagram of the result of detecting Figure 7(a) using the prior art in this embodiment;

[0033] Figure 7(c) is a schematic diagram of the result of detecting Figure 7(a) using the method proposed in the present invention in this embodiment. Detailed implementation manners

[0034] The present invention will be further described in detail below in conjunction with the embodiments with reference to the drawings.

[0035] As Figure 1 shown, a method for detecting small target defects of industrial parts based on a lightweight model includes the following steps:

[0036] Step ①, establish an image set of small target defects of industrial parts, and divide the image set of small target defects of industrial parts into a training set, a validation set and a test set according to a ratio. Specifically: select at least 5000 industrial part images containing small target defects from the Internet as the image set of small target defects of industrial parts, and divide the image set of small target defects of industrial parts into a training set, a validation set and a test set according to a ratio of 7:2:1;

[0037] Step ②, label the training set and the validation set to obtain the labeled training set and the labeled validation set. Specifically: use the LabelMe labeling method to label the small target defects in each image in the training set and the validation set, save the labeling information, and obtain the labeled training set and the labeled validation set;

[0038] Step ③, input the labeled training set and the labeled validation set into the lightweight small target defect detection model for training to obtain the trained lightweight small target defect detection model. The specific training process is as follows:

[0039] Set the training parameters including: the batch size is set to 32, the initial learning rate is set to 0.001, the weight decay is set to 5e-4, the number of training epochs is set to 100 epochs, the cosine annealing strategy is used as the learning rate scheduler, and the minimum learning rate is 0.0001;

[0040] Input the augmented and labeled training set into the lightweight small target defect detection model for training; where the data augmentation includes existing Mosaic augmentation, random flipping, HSV color space perturbation, and scale transformation (scaling range 0.5 - 1.5);

[0041] According to the loss function L total , use the Adam (Adaptive Moment Estimation) optimizer to iteratively train the lightweight small target defect detection model. Every 5 epochs of training, use the labeled validation set to evaluate the lightweight small target defect detection model once. Adjust the learning rate according to the evaluation results. Stop training when the evaluation results are the same for two consecutive times to obtain the trained lightweight small target defect detection model;

[0042] The loss function L total consists of classification loss, localization loss, and confidence loss. L total = λ cls *L cls + λ loc *L loc + λ obj *L obj , where λ cls , λ loc and λ obj are hyperparameters used to adjust the weights of different parts of the loss. λ cls = 0.5, λ loc = 0.05, λ obj = 1.0, L cls represents the classification loss, which is used to solve class imbalance. The parameters γ = 2.0, α = 0.25, L loc represents the localization loss, which is used to comprehensively consider the overlapping area, the distance of the center point, and the aspect ratio. L obj represents the confidence loss, which is used for the confidence of whether the prediction box contains the target;

[0043] The metrics for the evaluation results include mAP@0.5 (mean average precision when the IoU threshold is 0.5), FPS (frames per second), and the number of parameters (Params), where FPS is used to reflect the inference speed.

[0044] Step ④, input the test set into the trained lightweight small target defect detection model for detection, and output the target defect detection results;

[0045] In this embodiment, the lightweight small target defect detection model is obtained by improving the existing YOLOv5 network. The Conv module in the existing YOLOv5 network is replaced with a CS-Conv module, and the C3 module in the existing YOLOv5 network is replaced with an LRM module. A module for detecting extremely small targets is added between the sixteenth and seventeenth layers of the existing YOLOv5 network. Among them, the network structure of the existing YOLOv5 network is as Figure 2 shown, including an input end, a Conv layer at the 0th layer, a Conv layer at the 1st layer, a C3 layer at the 2nd layer, a Conv layer at the 3rd layer, a C3 layer at the 4th layer, a Conv layer at the 5th layer, a C3 layer at the 6th layer, a Conv layer at the 7th layer, a C3 layer at the 8th layer, an SPPF layer at the 9th layer, a Conv layer at the 10th layer, an Upsampling (upsampling) layer at the 11th layer, a Concat (cross-layer splicing) layer at the 12th layer, a C3 layer at the 13th layer, a Conv layer at the 14th layer, an Upsampling layer at the 15th layer, a Concat layer at the 16th layer, a C3 layer at the 17th layer, a Conv layer at the 18th layer, a Concat layer at the 19th layer, a C3 layer at the 20th layer, a Conv layer at the 21st layer, a Concat layer at the 22nd layer, a C3 layer at the 23rd layer, and three detection output layers;

[0046] The module for detecting extremely small targets consists of two LRM modules with the same functions and structures, two CS-Conv modules with the same functions and structures, an Up-sampling module, and two Concat modules with the same functions and structures. Among them, the input end of the first LRM module is connected to the output end of the sixteenth layer, the input end of the first CS-Conv module is connected to the output end of the first LRM module, the input end of the Up-sampling module is connected to the output end of the first CS-Conv module, the input ends of the first Concat module are respectively connected to the output end of the Up-sampling module and the output end of the second layer, the input end of the second LRM module is connected to the output end of the first Concat module, the input end of the second CS-Conv module is connected to the output end of the second LRM module, the input ends of the second Concat module are respectively connected to the output end of the first CS-Conv module and the output end of the second CS-Conv module, and the output end of the second Concat module is connected to the input end of the twenty-fourth layer (i.e., the seventeenth layer of the original YOLOv5 network);

[0047] The super-large feature map of 128×160×160 output by the 21st layer, the large feature map of 256×80×80 output by the 24th layer, the medium feature map of 512×40×40 output by the 27th layer, and the small feature map of 1024×20×20 output by the 30th layer are input into the detection head for detection;

[0048] As Figure 3 shown, the lightweight small target defect detection model consists of an input module for preprocessing the input image, a backbone network module for feature extraction, a Neck module for constructing ultra-small target detection, a head module for detection, and an output module for outputting feature maps of four different sizes;

[0049] The backbone network module consists of five CS-Conv modules with the same function and structure, four LRM modules with the same function and structure, and one SPPF module;

[0050] The Neck module consists of six CS-Conv modules with the same function and structure, three Up-sampling modules with the same function and structure, six Concat modules with the same function and structure, and six LRM modules with the same function and structure;

[0051] The head module consists of four detection heads;

[0052] As Figure 4 , 5 shown, the LRM module consists of two CSB modules connected in sequence. The CSB module includes two CS-Conv modules and one Add layer. After the input feature map is connected with the input feature map after passing through two CS-Conv modules in sequence by residual connection, feature fusion is performed in the Add layer, and the output feature map is used as the output of the CSB module. The output channel of the first CS-Conv module is c_, the output channel of the second CS-Conv is c2, and the output channel number of the CSB module is c2. The CSB module can not only reduce the computational amount and parameter quantity of the model, but also help the network better learn residual information, improve the efficiency and performance of the model, and thus can significantly reduce the complexity of the model while maintaining the performance.

[0053] As Figure 6As shown in the figure, the CS-Conv module includes a Conv1 layer, a Conv2 layer, and a Concat layer. In the Conv1 layer, small convolutions are used to perform convolution operations on the input feature map to generate partial feature maps. In the Conv2 layer, large convolutions are used to perform convolution operations on the partial feature maps to generate additional feature maps. In the Concat layer, the partial feature maps and the additional feature maps are residually connected, and the output feature map is used as the output of the CS-Conv module. The CS-Conv module is adopted to reduce the computational complexity and the number of parameters of the model while maintaining good performance.

[0054] For example, if the number of input channels of the CS-Conv module is c1 and the number of output channels is c2, then the number of input channels of the Conv1 layer is c1, and the number of output channels is c_, which is half of c2. The number of input channels of the Conv2 layer is c_, and the number of output channels is c_.

[0055] In this embodiment, the industrial part small target defect image set contains 5,000 industrial part images, covering 10 types of small target defects (such as cracks, scratches, holes, etc.). The hardware environment is an NVIDIA RTX 3090 GPU with 32 GB of memory. The software framework is PyTorch 1.10 and CUDA 11.3. The training time is about 12 hours (100 epochs) for single-card training.

[0056] The performance comparison of the object detection models of the method proposed in the present invention and other object detection methods is shown in Table 1. It can be seen that compared with Faster R-CNN, the mAP@0.5 of the present invention is improved by 9.5%, and the number of parameters is reduced by 95.7%. Compared with the baseline YOLOv5, the mAP@0.5 of the present invention is increased by 7.1%, and the number of parameters is reduced by 18.1%, which is more suitable for embedded deployment. And the measured FPS on Jetson Nano (embedded platform) reaches 95, meeting the real-time requirements.

[0057] Table 1 Performance comparison of the object detection models of the method proposed in the present invention and other object detection methods

[0058] Model mAP@0.5 FPS Params(M) Applicable Devices Faster R-CNN 65.8% 8 137.2 Server YOLOv5 (Baseline) 68.2% 120 7.2 Embedded / Server The Model of the Present Invention 75.3% 95 5.9 Embedded Device

[0059] The small target defect detection results of the method proposed in the present invention and other object detection methods are as Figure 7(a)-7(c) shown. Figure 7(a) is an image of the surface of a metal part. The background presents a complex mesh texture and local reflective areas. The micro-crack is located at the lower right edge of the part, with a length of 3 pixels and a width of less than 1 pixel (visible only when magnified to 200%). In Figure 7(b), the background reflective area is misdetected, and the micro-crack is not detected. However, in Figure 7(c), not only does it closely fit the micro-crack area, but also all the micro-cracks are detected, indicating that the method proposed in the present invention has better and higher detection accuracy for small defect targets.

[0060] To verify the contributions of each module in the method proposed by the present invention, ablation experiments as shown in Table 2 were thus conducted. It can be seen that after the single fusion CS-Conv module or LRM module, the mAP@0.5 increased by 3.3% and 2.6% respectively, achieving a significant improvement, which proves that the fusion scheme greatly improves the performance of small target defect detection for industrial parts. The method of the present invention obtained by combining the two types of modules achieved the optimal performance, with mAP@0.5 being 75.3%.

[0061] Table 2 Ablation Experiment

[0062] Experimental Group Model Configuration mAP@0.5 FPS Params(M) Baseline Model Original YOLOv5 68.2% 120 7.2 Experimental Group 1 YOLOv5+CS-Conv 71.5% 115 6.8 Experimental Group 2 YOLOv5+LRM 70.8% 110 6.5 Final Model The Present Invention 75.3% 95 5.9

[0063] Explanation of the terms of the present invention:

[0064] CS-Conv (Cloud Shadow Convolution): Cloud Shadow Convolution

[0065] CSB (Cloud Shadow BottleNeck): Cloud Shadow Bottleneck

[0066] LRM (Lightweight residual module): Lightweight residual module

[0067] SPPF (Spatial Pyramid Pooling-Fast): Fast Spatial Pyramid Pooling

Claims

1. A method for detecting small target defects of industrial parts based on a lightweight model, comprising the following steps: Step ①, establish an image set of small target defects of industrial parts, and divide the image set of small target defects of industrial parts into a training set, a validation set and a test set according to a ratio; Step ②, label the training set and the validation set to obtain the labeled training set and the labeled validation set; Step ③, input the labeled training set and the labeled validation set into a lightweight small target defect detection model for training to obtain a trained lightweight small target defect detection model; Step ④, input the test set into the trained lightweight small target defect detection model for detection, and output the target defect detection result; It is characterized in that the lightweight small target defect detection model is obtained by improving the existing yolov5 network, replacing the Conv module in the existing yolov5 network with a CS-Conv module, replacing the C3 module in the existing yolov5 network with an LRM module, and adding a module for detecting extremely small targets between the sixteenth layer and the seventeenth layer of the existing yolov5 network; The module for detecting extremely small targets is composed of two LRM modules with the same function and structure, two CS-Conv modules with the same function and structure, an Up-sampling module and two Concat modules with the same function and structure. The input end of the first LRM module is connected to the output end of the sixteenth layer, the input end of the first CS-Conv module is connected to the output end of the first LRM module, the input end of the Up-sampling module is connected to the output end of the first CS-Conv module, the input ends of the first Concat module are respectively connected to the output end of the Up-sampling module and the output end of the second layer, the input end of the second LRM module is connected to the output end of the first Concat module, the input end of the second CS-Conv module is connected to the output end of the second LRM module, the input ends of the second Concat module are respectively connected to the output end of the first CS-Conv module and the output end of the second CS-Conv module, and the output end of the second Concat module is connected to the input end of the twenty-fourth layer; Input the extremely large feature map of 128×160×160 output by the twenty-first layer, the large feature map of 256×80×80 output by the twenty-fourth layer, the medium feature map of 512×40×40 output by the twenty-seventh layer, and the small feature map of 1024×20×20 output by the thirtieth layer into the detection head for detection.

2. The industrial part small target defect detection method based on a lightweight model according to claim 1, wherein The LRM module is composed of two CSB modules connected in sequence, The CSB module includes two CS-Conv modules and an Add layer. After the input feature map is connected with the input feature map after being operated by the two CS-Conv modules in sequence in a residual manner, feature fusion is performed in the Add layer, and the output feature map is used as the output of the CSB module.

3. The industrial part small target defect detection method based on a lightweight model according to claim 2, wherein The described CS-Conv module includes a Conv1 layer, a Conv2 layer, and a Concat layer. In the Conv1 layer, small convolutions are used to perform convolution operations on the input feature map to generate partial feature maps; in the Conv2 layer, large convolutions are used to perform convolution operations on the partial feature maps to generate additional feature maps; in the Concat layer, the partial feature maps and the additional feature maps are subjected to residual connection, and the output feature map is used as the output of the CS-Conv module.

4. A method for detecting small target defects of industrial parts based on a lightweight model according to claim 1, characterized in that The specific operation process of step ① is as follows: Select at least 5000 industrial part images containing small target defects from the Internet as the industrial part small target defect image set, and divide the industrial part small target defect image set into a training set, a validation set, and a test set according to the ratio of 7:2:

1.

5. A method for detecting small target defects of industrial parts based on a lightweight model according to claim 1, characterized in that In step ②, the LabelMe annotation method is used to annotate the small target defects in each image in the training set and the validation set, and the annotation information is saved to obtain the annotated training set and the annotated validation set.

6. The industrial part small target defect detection method based on a lightweight model according to claim 1, characterized in that The training process in step ③ is as follows: Set the training parameters including: the batch size is set to 32, the initial learning rate is set to 0.001, the weight decay is set to 5e-4, the number of training epochs is set to 100 epochs, and the cosine annealing strategy is used as the learning rate scheduler; Input the data-augmented annotated training set into the lightweight small target defect detection model for training; According to the loss function L total , the lightweight small target defect detection model is iteratively trained using the Adam optimizer. The lightweight small target defect detection model is evaluated once every 5 epochs using the labeled validation set. When the evaluation results are the same for two consecutive times, the training is stopped to obtain the trained lightweight small target defect detection model; The loss function L total is composed of classification loss, localization loss, and confidence loss, and L total = λ cls * L cls + λ loc * L loc + λ obj * L obj , where λ cls , λ loc and λ obj represent hyperparameters respectively, L cls represents classification loss, L loc represents localization loss, and L obj represents confidence loss.