Industrial defect detection method and system based on YOLOv5 and knowledge optimization auxiliary distillation
Through the lightweight industrial defect detection model based on YOLOv5 and knowledge optimization assisted distillation, the detection problems in limited computing resources and complex environments are solved, and efficient and accurate industrial defect detection is achieved, suitable for embedded systems.
Patent Information
- Application Number
- CN202510310327.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-04
AI Technical Summary
The existing industrial defect detection methods have high computational complexity, insufficient accuracy, poor real-time performance and difficult model deployment on embedded systems with limited computing resources, making it difficult to meet the detection needs of complex industrial environments.
The lightweight industrial defect detection model based on YOLOv5 and knowledge optimization assisted distillation is adopted. By introducing the C3DSConv module, SimAM module and ASPP module, it optimizes feature extraction, and combining the KO assisted model for knowledge distillation, reducing the computational complexity and improving detection accuracy and real-timeness.
It significantly reduces the computational complexity and number of parameters of the model, improves detection accuracy and real-timeness, enhances the adaptability of the model in complex industrial environments, simplifies the deployment process, and is suitable for embedded systems with limited computing resources.
Smart Images

Figure CN120259205A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image detection, and specifically to an industrial defect detection method and system based on YOLOv5 and knowledge optimization-assisted distillation. Background Art
[0002] Industrial defect detection is an important link in industrial production to ensure product quality and improve production efficiency. Traditional defect detection methods mostly rely on manual visual inspection or conventional physical detection techniques (such as ultrasonic detection, magnetic particle detection, ray detection, and infrared thermal imaging). These methods have the following significant defects:
[0003] Manual inspection: Manual inspection is time-consuming, labor-intensive, and vulnerable to human factors, resulting in false detections and missed detections. In addition, the efficiency of manual inspection in a large-scale production environment is low and cannot meet the high-paced industrial demands.
[0004] Physical detection: Physical detection equipment is expensive and has limitations for specific materials or defect types. For example, ray detection requires the use of special equipment and has certain safety risks, while magnetic particle detection is not applicable to non-magnetic materials.
[0005] With the rise of deep learning technology, object detection methods based on convolutional neural networks (CNNs) have gradually replaced traditional detection methods and become one of the mainstream technologies for industrial defect detection.
[0006] In recent years, object detection models such as Faster R-CNN, SSD, and the YOLO series have been widely applied to industrial defect detection scenarios. They extract features in images, locate and classify defect regions, greatly improving the detection efficiency and accuracy. Among them, the YOLO series models have been widely recognized for their real-time advantages (completing object localization and classification in one forward pass) and have gradually developed multiple versions (YOLOv3, YOLOv4, YOLOv5, etc.), continuously optimizing in terms of performance and model complexity. Despite the great progress of deep learning technology, the following main challenges still exist in industrial defect detection:
[0007] (1) Computational resource limitations: Industrial sites usually use embedded devices or resource-constrained devices. The computational complexity of existing object detection models is relatively high, making it difficult to meet real-time requirements. For example, although the YOLOv5 model has optimized the model complexity through the CSPDarkNet architecture, its computational overhead and the number of parameters are still large on embedded devices.
[0008] (2) Trade-off between model lightweight and detection accuracy: Existing lightweight models (such as YOLOv3-tiny, MobileNet-SSD) achieve real-time operation by reducing the number of model parameters, but often at the cost of sacrificing detection accuracy and are difficult to meet high-precision requirements.
[0009] (3) Poor adaptability to complex industrial scenarios: In industrial environments, complex scenarios such as motion blur, small target defects, and multiple defects occurring simultaneously often exist, which pose higher requirements for the feature extraction ability of the model. Existing models are prone to false detections and missed detections in these situations.
[0010] (4) Limitations of knowledge transfer: When there is a large gap in the scale between the teacher model and the student model in current knowledge distillation technology, the knowledge transfer effect is limited, resulting in the performance of the student model being difficult to match that of the teacher model. Summary of the Invention
[0011] The present invention proposes a lightweight industrial defect detection model - YOLO-DSA based on YOLOv5 and knowledge optimization assisted distillation, aiming to solve the following problems existing in existing industrial defect detection methods:
[0012] 1. High computational complexity: Traditional deep learning models usually require a large amount of computing resources, resulting in ineffective deployment on embedded systems or edge devices with limited computing resources.
[0013] 2. Insufficient accuracy: In complex industrial environments, traditional models have low detection accuracy when dealing with small targets, occlusions, blurred images, etc., and cannot meet the actual needs.
[0014] 3. Poor real-time performance: Traditional models cannot achieve real-time performance while ensuring accuracy, resulting in low efficiency in the defect detection process.
[0015] 4. Difficult model deployment: When existing models are applied to industrial sites, due to computing resource limitations, it is often difficult to deploy them in limited hardware environments, affecting the popularization and application of the detection system.
[0016] To achieve the above object, the present invention provides an industrial defect detection method based on YOLOv5 and knowledge optimization assisted distillation, and the steps include:
[0017] Collect a defect data set in the actual production environment;
[0018] Construct an industrial defect detection model based on YOLOv5 and knowledge optimization assisted distillation, and the industrial defect detection model includes: a backbone network, a neck network, and a detector head; in the backbone network, a C3DSConv module is adopted, and at the same time, the combination of a SimAM module and an ASPP module is introduced; the neck network adopts a path aggregation network and a C3DSConv module; the detector head realizes the regression and classification functions of the target box through 1×1 convolution;
[0019] Input the defect data set into the industrial defect detection model to complete the detection of industrial defects.
[0020] Preferably, the defect types in the defect dataset include: chipping, cracking, scratching, grooving, holes, and edge sealing defects; after the defect dataset is collected, data augmentation is performed through preprocessing, and it is divided into a training set and a validation set according to a ratio of 7:3.
[0021] Preferably, the standard convolution in the C3 module in the original YOLOv5 is replaced with DSConv to obtain the C3DSConv module; in the C3DSConv module, DSConv is used in combination with the DSBottleneck module to enhance the feature extraction ability; the stride of the first convolutional layer of the C3DSConv module is 2, which is used to halve the size of the feature map, quickly reduce the resolution of the feature map, and reduce the subsequent computational load; the strides of the second and third convolutional layers are 1, which are used to further extract and refine features.
[0022] Preferably, the SimAM module and the ASPP module are combined to optimize the feature extraction and fusion ability in complex scenarios;
[0023] Among them, the SimAM module generates attention weights to highlight key regions by calculating the similarity between each pixel and its neighboring pixels:
[0024]
[0025] Among them, w t and b t represent the attention weight and bias of the i-th pixel respectively; x i and t represent other neurons and target neurons in a single channel of the input feature ; M = H × W represents the neurons on this channel;
[0026] The ASPP enhances the detection ability for various sizes of defects through parallel operations with multiple dilation convolution sampling rates.
[0027] Preferably, a KO auxiliary model is adopted to improve the accuracy of the industrial defect detection model; the KO auxiliary model includes: a local perception unit, a lightweight multi-head self-attention module, and a feature map generation module; a patch embedding layer composed of depthwise separable convolution DSConv and layer normalization is applied before each stage of the KO model.
[0028] The present invention also provides an industrial defect detection system based on YOLOv5 and knowledge optimization assisted distillation. The system is used for the above method and includes: a collection module, a construction module, and a detection module;
[0029] The collection module is used to collect the defect dataset in the actual production environment;
[0030] The building block is used to build an industrial defect detection model based on YOLOv5 and knowledge-optimized assisted distillation. The industrial defect detection model includes three parts: a backbone network, a neck network, and a detector head. In the backbone network, a C3DSConv module is adopted, and at the same time, the combination of a SimAM module and an ASPP module is introduced. The neck network adopts a path aggregation network and uses a C3DSConv module. The detector head realizes the regression and classification functions of the target box through 1×1 convolution.
[0031] The detection module is used to input the defect data set into the industrial defect detection model to complete the detection of industrial defects.
[0032] Preferably, the defect types in the defect data set include: chipping, crack, scratch, groove, hole, and edge sealing defect. After the defect data set is collected, the acquisition module performs data augmentation through preprocessing and divides it into a training set and a validation set according to a ratio of 7:3.
[0033] Preferably, the building block replaces the standard convolution in the C3 module in the original YOLOv5 with DSConv to obtain the C3DSConv module. In the C3DSConv module, DSConv is combined with the DSBottleneck module to enhance the feature extraction ability. The stride of the first convolutional layer of the C3DSConv module is 2, which is used to halve the size of the feature map, quickly reduce the resolution of the feature map, and reduce the subsequent calculation amount. The strides of the second and third convolutional layers are 1, which are used to further extract and refine features.
[0034] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0035] 1. Improved computational efficiency: Through the introduction of the DSConv module and the design of the C3DSConv module, the computational complexity and the number of parameters of the model are significantly reduced, enabling it to operate efficiently on embedded systems and edge devices with limited computational resources and having strong resource adaptability.
[0036] 2. Enhanced detection accuracy: Through the combination of the SimAM attention mechanism and ASPP, the model's ability to extract multi-scale features and its attention to key features are enhanced, enabling the model to maintain a high detection accuracy in complex industrial environments (such as small object detection, image blur, etc.).
[0037] 3. Optimized real-time performance: Through lightweight design and an efficient knowledge distillation strategy, the present invention ensures the real-time performance of the model while improving the detection accuracy, meeting the requirements for rapid detection in industrial production.
[0038] 4. Enhanced the adaptability of the model: Through two-stage cascaded knowledge distillation based on the KO auxiliary model, the problem of low learning efficiency when there is a large gap in the scale between the teacher and student models in traditional distillation methods is solved, the performance of the lightweight model is improved, and the application scenarios of the model are extended.
[0039] 5. Simplified the model deployment: Due to the significant reduction in the amount of calculation and the number of parameters of the model of the present invention, the deployment and execution are more convenient, and it can adapt to more types of industrial field devices and application environments, having a wide range of application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the technical solutions of the present invention, the drawings required to be used in the embodiments are briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0041] Figure 1 Schematic diagram of the industrial defect dataset for the embodiments of the present invention; among them, (a) represents chipping; (b) represents crack; (c) represents scratch; (d) represents groove; (e) represents hole; (f) represents edge sealing defect;
[0042] Figure 2 Schematic diagram of the YOLO-DSA structure for the embodiments of the present invention;
[0043] Figure 3 Schematic diagram of the C3DSConv module for the embodiments of the present invention;
[0044] Figure 4 Schematic diagram of the ASPP module for the embodiments of the present invention;
[0045] Figure 5 Schematic diagram of the two-stage cascaded knowledge distillation architecture for the embodiments of the present invention;
[0046] Figure 6 Schematic diagram of the KO model for the embodiments of the present invention; among them, (a) represents the overall model; (b) represents the MBConv module therein; (c) represents the KT Block module;
[0047] Figure 7 Schematic diagram of the FGM module for the embodiments of the present invention;
[0048] Figure 8 Schematic diagram of the ABF module for the embodiments of the present invention;
[0049] Figure 9 Schematic diagram of the CL structure for the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0050] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0051] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0052] Embodiment 1
[0053] This embodiment provides an industrial defect detection method based on YOLOv5 and knowledge optimization-assisted distillation. The steps include:
[0054] S1. Collect defect data sets in the actual production environment.
[0055] As Figure 1 shown, the industrial defect data set used in this embodiment comes from the actual production environment and includes various defect types, such as chipping, cracking, scratching, grooving, holes, and edge sealing defects. The data volume of each type of defect is 800 images. After the data is preprocessed, data augmentation is performed to improve the generalization ability of the model. Finally, 6000 images are obtained; and the data set is divided into a training set and a validation set according to a ratio of 7:3. The specific steps are as follows:
[0056] S101 Data preprocessing
[0057] In the data preprocessing stage, first, the images are unified, including format conversion and size adjustment. All original pictures are converted into a consistent format, and the resolution is standardized to the size adapted to the model (640×640 pixels). Subsequently, the images are denoised, and the interference of noise is reduced through Gaussian blur or median filtering, thereby improving the data quality. In addition, the pixel values of the images are normalized, standardized to the range [0,1], and normalized using the mean and standard deviation to enhance the stability of training. For the annotation of defects in each image, a standard annotation format (YOLO format) is adopted to ensure the accuracy of the defect area and provide high-quality supervision information for subsequent model training.
[0058] S102 Data augmentation
[0059] To expand the images of each type of defect from 800 to 6000, a series of data augmentation strategies were adopted. Geometric transformation is the main augmentation method, including random rotation (angle range [-30°, 30°]), scaling (scale range [0.8, 1.2]), translation (pixel range [-20, 20]), and flipping (horizontal or vertical direction). These methods simulate the possible forms of defects at different angles and scales. In addition, by adjusting brightness ([-20%, 20%]), contrast ([-15%, 15%]), hue ([-10, 10]), and gamma value, defect images under various lighting conditions were generated, thus enhancing the robustness of the model.
[0060] To enhance the model's adaptability to complex scenes, Gaussian noise (standard deviation [0, 0.05]) was also added to the images, and the non-defect areas were randomly occluded (such as adding rectangular or circular occluders). In more advanced operations, multi-defect scenes were created by stitching multiple defect images, and new samples were generated by randomly cropping parts of the images.
[0061] These augmentation strategies were combined and applied, and each original image generated approximately 6.5 times the augmented version, thus expanding the data volume to 6000 images for each type. The augmented dataset was divided into a training set (4200 images) and a validation set (1800 images) at a ratio of 7:3, significantly improving the generalization ability of the model and the recognition effect of diverse defects in the actual production environment.
[0062] S2. Construct an industrial defect detection model based on YOLOv5 and knowledge-optimized assisted distillation.
[0063] In this embodiment, YOLOv5 was used as the benchmark algorithm to construct a lightweight improved industrial defect detection model, aiming to reduce the computational amount and complexity of the model. The YOLO-DSA structure is as Figure 2 shown. The model includes three parts: a backbone network, a neck network, and a detector head. In the backbone network, the C3 module in the original YOLOv5 was replaced with the proposed lightweight C3DSConv module, and at the same time, the combination of SimAM and ASPP modules was introduced to achieve the goal of maintaining the lightweight of the model without losing accuracy. The neck network adopts the Path Aggregation Network (PANet), and at the same time, the lightweight C3DSConv module is used to further reduce the model parameters and effectively fuse the shallow and deep features, improving the accuracy of target recognition. The detector head realizes the regression and classification functions of the target box through 1×1 convolution to ensure the accurate prediction of the detection target position and category.
[0064] In image processing tasks in computer vision, in order to capture more semantic information in images, feature extraction networks often need to stack multiple layers of convolutional operations. However, doing so will lead to an increase in parameters, making it complex and difficult to train, and generating redundant features. These redundant features not only occupy additional storage space but also consume a large amount of computing resources when processing them. If a method can be found to reduce or avoid the generation of such redundant features, then the computational amount required for the model to process images can be significantly reduced, thereby improving the overall efficiency. In other words, by optimizing the feature extraction process, unnecessary computational overhead can be reduced, and more efficient image processing can be achieved. Therefore, in this embodiment, DSConv is selected to replace the standard convolution in the original C3 module, and the lightweight C3DSConv module is proposed.
[0065] C3DSConv module: Standard convolutions usually create a large number of similar feature maps, resulting in expensive computations and resource consumption. To solve this problem, the standard convolution is replaced with DSConv in the C3 module. Thus, while maintaining the detection accuracy of the model and increasing its applicability and generalization ability, the computational amount and the number of parameters are greatly reduced, the computational efficiency of the model is improved, and the model training and inference speeds are increased. In the proposed C3DSConv module, DSConv is used in combination with the DSBottleneck module to enhance the feature extraction ability. The stride of the first convolutional layer of the C3DSConv module is 2, which can halve the size of the feature map, thereby quickly reducing the resolution of the feature map and reducing the subsequent computational amount. The strides of the second and third convolutional layers are 1, mainly used to further extract and refine features, which helps the model better capture the detailed information in the image and improve the accuracy of defect detection. At the same time, the step of retaining the original feature map in the original model is removed, thereby simplifying the model structure and reducing the computational amount. The overall structure is as Figure 3 shown.
[0066] In this embodiment, the SimAM module and the ASPP module are combined. This design realizes a comprehensive optimization of the feature extraction and fusion capabilities in complex scenarios. The SimAM module highlights the key regions through the self-similarity mechanism, and the ASPP module uses multi-scale dilated convolutions to enhance the adaptability to defects of different sizes, which is particularly suitable for small target detection and occlusion scenarios. The SimAM module generates attention weights to highlight the key regions by calculating the similarity between each pixel and its neighboring pixels, and the calculation formula is shown in Equation (1). The ASPP module performs parallel operations through multiple dilated convolution sampling rates, Figure 4It shows the structure of the ASPP module, which can extract rich features at different scales, thereby enhancing the detection ability for various sizes of defects. Especially when dealing with small objects and blurred images encountered in industrial scenarios, it can effectively improve the accuracy. The SimAM module preferentially provides a weighted input for the ASPP module, highlighting key features before performing multi-scale information extraction and fusion.
[0067]
[0068] Among them, w t and b t represent the attention weight and bias of the i-th pixel respectively; x i and t represent other neurons and the target neuron in a single channel of the input feature . M = H×W represents the neurons on this channel.
[0069] The combination of SimAM and the ASPP module significantly improves the small target detection accuracy and robustness of the model in complex scenarios through the collaborative optimization of the parameter-free attention mechanism and multi-scale feature extraction, while maintaining the efficient operation of the lightweight design.
[0070] In the previous step, the model was lightweighted by replacing the convolutional module and deleting the steps to retain the original feature map, but this led to a certain loss of accuracy. Therefore, it is necessary to restore the accuracy of the model. Traditional accuracy restoration training methods are similar to ordinary model training, but lack richer prior or supervision information, so it is difficult to achieve better performance. To solve this problem, this embodiment introduces the method of knowledge distillation, that is, extracting distilled knowledge from a high-precision teacher model and transferring it to the lightweight model to imitate the teacher network, which can significantly improve the performance of the lightweight model.
[0071] However, in the current knowledge distillation methods, there is a problem of low learning efficiency of the student model, which is because there are large differences in structure and parameters between the high-precision teacher model and the student model. To solve this problem, this embodiment proposes a new knowledge distillation method based on the KO auxiliary model, and introduces this model to achieve the purpose between the teacher and student models. The architecture of knowledge distillation is as Figure 5 shown.
[0072] The KO model aims to achieve more efficient knowledge transfer by combining the advantages of CNN and Transformer. The overall framework is as Figure 6As shown, the original Transformer model, when processing images, adopts the method of dividing the input image into non-overlapping patches, and only processes the information inside the patches through linear transformation. Although this method is simple and efficient, it has limitations in capturing global image information and local details. To overcome this limitation, three 3×3 convolutional blocks are introduced. Among them, the first convolutional block uses a stride of 2 to reduce the size of the input image, achieving downsampling, thereby effectively expanding the receptive field of subsequent convolutional layers. The next two 3×3 convolutional blocks use a stride of 1 and focus on extracting local features of the image. This design not only retains the ability of the Transformer model to process global information but also enhances the model's perception of image details through the local feature extraction ability of CNN. According to the design concept of modern CNN models, the KO model is further divided into five stages and extracts features and reduces parameters through the MBConv (Mobile Inverted Bottleneck Convolution) module, such as Figure 6 (b) shows, where each stage generates feature maps of different sizes. These feature maps not only enrich the model's representation ability but also provide multi-level supervision information for the knowledge distillation process. During the knowledge distillation process, the teacher model transmits more comprehensive and refined knowledge to the student model through these multi-level feature maps.
[0073] To generate hierarchical representations, the KO model applies a patch embedding layer composed of depthwise separable convolution DSConv and layer normalization (LayerNorm, LN) before each stage. The aim is to reduce the size of intermediate features, specifically by downsampling by a factor of 2 to reduce the resolution, thereby effectively expanding the receptive field of subsequent layers. At the same time, this embedding layer also projects the features into a larger dimension, that is, doubling the size, to increase the model's expressive ability. Inside each stage, multiple KT blocks are stacked in sequence for feature transformation. In each stage, the KT blocks process the features while maintaining the same resolution, so that the model can perform deep feature extraction and transformation while maintaining details. Taking "Stage 3" and "Stage 4" of the KO model as examples, they contain 6 and 9 KT blocks respectively. Through the stacking of these KT blocks, the model can gradually refine more abstract and high-level feature representations, thereby supporting more complex image recognition and understanding tasks. The model ends with a global average pooling layer, a fully connected layer, and a classification layer with softmax.
[0074] The KT module consists of a Local Perception Unit (LPU), a Lightweight Multi-head Self-attention (LMHSA) module, and a Feature Map Generator (FGM). Its structure diagram is as shown in Figure 6 (c).
[0075] In traditional vision transformers, absolute position encoding is used to introduce unique position information for each patch, which destroys translational invariance because it changes the relative relationship between patches. In addition, vision transformers also ignore the local structure and relationship information within patches. To address these issues, the LPU is used, whose goal is to extract local relevant information without affecting the global structure of the patches. The LPU learns to obtain the local context around each patch and uses this information to enhance the model's ability to understand local features and relationships. This approach helps to maintain the invariance of the model to operations such as rotation and translation, while improving performance in complex visual scenarios.
[0076] In the KT block, the input image is divided into non-overlapping tiles and processed by the LMHSA module, and the final output is a sequence of feature vectors. These sequences of feature vectors are abstract representations of the local and global features of the image, rather than directly encoding features spatially like the feature maps in CNN models. For this reason, the FGM is proposed, whose main function is to generate a feature map from the sequence of feature vectors, that is, to convert the abstract feature vectors into a specific two-dimensional feature map. Its structure diagram is as shown in Figure 7 shown.
[0077] In the first stage, an initial knowledge transfer foundation is established through knowledge distillation between the Teacher model and the KO model. In the second stage, knowledge is transferred from the KO model to the Student model, i.e., YOLO-DSA. In the second stage, to improve the performance of knowledge distillation, this embodiment adopts an attention-based fusion (ABF) module, and a comparative loss (CL) function is proposed. The ABF module combines feature maps from different stages in knowledge distillation, thereby improving the distillation performance. The structure is as shown in Figure 8 shown.
[0078] In the first stage of knowledge distillation, the L2 distance is used to compare the gap between two feature maps. However, in the second stage of knowledge distillation, the feature information at different levels is aggregated to learn from the teacher, but it is not sufficient to effectively transfer the feature information at different levels. Therefore, this embodiment proposes CL as the loss function for the second stage. The CL structure is as shown inFigure 9 As shown, the loss function transforms knowledge into feature information at different levels by using spatial pyramid pooling, making it easier to extract information. Then, the L2 distance is used to perform knowledge distillation between these levels respectively.
[0079] S3. Input the defect data set into the industrial defect detection model to complete the detection of industrial defects.
[0080] Embodiment 2
[0081] To verify the accuracy of the present invention, this embodiment is specifically set as a comparative experiment for verification.
[0082] The experimental environment and parameter configuration are shown in Table 1.
[0083] Table 1
[0084]
[0085] In the present invention, to more accurately evaluate the performance of the model, we selected several mature metrics, including precision, recall, F1-score, and average precision, which are used to comprehensively measure the effect of the model in the defect detection task. To evaluate the complexity of the model, we used metrics such as parameter count and floating-point operations (FLOP). The two metrics of precision and recall are calculated based on the confusion matrix. The confusion matrix classifies the prediction results into four types according to the true labels.
[0086] (1) Ablation experiment
[0087] To prove the effectiveness of each module of the model improvement, an ablation experiment was set up in the experiment of the present invention, and the results are shown in Table 2.
[0088] Table 2
[0089]
[0090] As shown in Table 2, the original YOLOv5 has a mAP@0.50 value of 93.7% and a mAP@0.50:0.95 value of 36.1%. At the same time, its GFLOPs and number of parameters are also the worst among the above models. By replacing the standard convolution in the original model with DSConv, the mAP@0.50 value dropped by 3.1%, the mAP@0.50:0.95 value dropped by 2.9%, and the number of model parameters was reduced by 47.7%. By introducing ASPP and SimAM into the YOLOv5s model with DSConv added, the mAP@0.50 value increased by about 0.2%, and the mAP@0.50:0.95 value increased by 0.2%. By replacing ASPP with the original SPPF, the mAP@0.50 value increased by 1.0%, and the mAP@0.50:0.95 value increased by 0.9%. Afterwards, the model of the present invention improved the mAP@0.50 of the YOLOv5s model of ASPP by about 1.2%, and the mAP@0.50:0.95 value by 0.8%. The number of parameters was only slightly increased. Afterwards, after the model of the present invention replaced the SPPF with the ASPP module on the basis of adding the SimAM module, its mAP@0.50 value and mAP@0.50:0.95 value also showed a steady improvement, which was 2.2% and 1.6% higher than the model replaced with DSConv, respectively. On this basis, the model was trained by knowledge distillation. Experiments show that the mAP@0.50 value and mAP@0.50:0.95 value increased by 0.8% and 0.5% respectively compared with the previous model. Compared with the baseline model YOLOv5s, the mAP@0.50 decreased by 0.4%, the mAP@0.50:0.95 value decreased by 0.8%, and the number of parameters decreased by 47.4%. After the introduction of the KO model, the mAP@0.50 value and mAP@0.50:0.95 value increased by 1.5% compared with the baseline model YOLOv5s, which shows the effectiveness of model improvement based on knowledge distillation of the KO model.
[0091] (2) Comparative experiment
[0092] In order to illustrate the advancedness of the model proposed in the present invention and reflect the high lightness of the model, the comparative experiment selected the mainstream target detection algorithm and the Yolov5s algorithm combined with the mainstream lightweight network to compare with the proposed model. The experimental results of the present invention are shown in Table 3.
[0093] Table 3
[0094]
[0095] Table 3 shows that although the model proposed in the present invention is slightly lower than YOLOv8s in terms of the mAP@0.50 value (0.1% lower), its GFLOPs and the number of parameters are significantly lower than those of other mainstream models. For example, the GFLOPs of YOLOv5s-GhostNet is 13.4 and the number of parameters is 6.06M, while those of the model of the present invention are 12.3 and 3.70M respectively. Compared with Faster R-CNN, SSD, and YOLOv3-tiny, the model of the present invention has a significant improvement in mAP@0.50, while achieving a faster running speed and lower parameter requirements.
[0096] Through the above steps, the present invention can achieve efficient and accurate industrial defect detection, which is especially suitable for embedded systems and environments with limited computing resources, and has broad application prospects.
[0097] Embodiment 3
[0098] This embodiment also provides an industrial defect detection system based on YOLOv5 and knowledge optimization assisted distillation, including: an acquisition module, a construction module, and a detection module; the acquisition module is used to acquire a defect data set in the actual production environment; the construction module is used to construct an industrial defect detection model based on YOLOv5 and knowledge optimization assisted distillation, and the industrial defect detection model includes: a backbone network, a neck network, and a detector head; in the backbone network, a C3DSConv module is adopted, and at the same time, the combination of a SimAM module and an ASPP module is introduced; the neck network adopts a path aggregation network and a C3DSConv module; the detector head realizes the regression and classification functions of the target box through 1×1 convolution; the detection module is used to input the defect data set into the industrial defect detection model to complete the detection of industrial defects.
[0099] Next, in combination with this embodiment, it will be detailed how the present invention solves technical problems in real life.
[0100] First, use the acquisition module to acquire the defect data set in the actual production environment.
[0101] As Figure 1 shown, the industrial defect data set adopted in this embodiment comes from the actual production environment and contains various defect types, such as chipping, cracking, scratching, grooving, holes, and edge sealing defects. The data volume of each type of defect is 800 images. After the data is preprocessed, data augmentation is performed to improve the generalization ability of the model, and finally 6000 images are obtained; and the data set is divided into a training set and a validation set according to a ratio of 7:3. The specific process is as follows:
[0102] In the data preprocessing stage, the images are first unified, including format conversion and size adjustment. All original images are converted into a consistent format, and the resolution is standardized to the size adapted to the model (640×640 pixels). Subsequently, the images are denoised, and the interference of noise is reduced through Gaussian blur or median filtering, thereby improving the data quality. In addition, the pixel values of the images are normalized to the range [0,1], and the normalization is performed using the mean and standard deviation to enhance the stability of training. For the annotation of defects in each image, a standard annotation format (YOLO format) is adopted to ensure the accuracy of the defect area and provide high-quality supervision information for subsequent model training.
[0103] To expand the number of images of each type of defect from 800 to 6000, a series of data augmentation strategies are adopted. Geometric transformation is the main augmentation method, including random rotation (angle range [-30°,30°]), scaling (scale range [0.8,1.2]), translation (pixel range [-20,20]), and flipping (horizontal or vertical direction). These methods simulate the possible forms of defects at different angles and scales. In addition, by adjusting the brightness ([-20%,20%]), contrast ([-15%,15%]), hue ([-10,10]), and gamma value, defect images under various lighting conditions are generated, thereby improving the robustness of the model.
[0104] To enhance the model's adaptability to complex scenarios, Gaussian noise (standard deviation [0,0.05]) is also added to the images, and the non-defect areas are randomly occluded (such as adding rectangular or circular occluders). In more advanced operations, multi-defect scenarios are created by stitching multiple defect images, and new samples are generated by randomly cropping part of the images.
[0105] These augmentation strategies are combined and applied, and each original image generates approximately 6.5 times the augmented version, thereby expanding the data volume to 6000 images for each type. The augmented dataset is divided into a training set (4200 images) and a validation set (1800 images) in a ratio of 7:3, significantly improving the generalization ability of the model and the recognition effect of diverse defects in the actual production environment.
[0106] After that, a module is constructed to build an industrial defect detection model based on YOLOv5 and knowledge optimization-assisted distillation.
[0107] In this embodiment, YOLOv5 is used as the benchmark algorithm to construct a lightweight improved industrial defect detection model, aiming to reduce the computational amount and complexity of the model. The YOLO-DSA structure is as Figure 2As shown in the figure. The model includes three parts: a backbone network, a neck network, and a detector head. In the backbone network, the C3 module in the original YOLOv5 is replaced with the proposed lightweight C3DSConv module, and at the same time, the combination of SimAM and ASPP modules is introduced to achieve the goal of maintaining the lightweight of the model without sacrificing accuracy. The neck network adopts the Path Aggregation Network (PANet), and at the same time, the lightweight C3DSConv module is used to further reduce the model parameters and effectively fuse the shallow and deep features, improving the accuracy of object recognition. The detector head realizes the regression and classification functions of the object bounding box through 1×1 convolution to ensure accurate prediction of the detection object position and category.
[0108] In the image processing tasks of computer vision, in order to capture more semantic information in images, feature extraction networks often need to stack multiple layers of convolutional operations. However, this will lead to an increase in parameters, making it complex and difficult to train, and generating redundant features. These redundant features not only occupy additional storage space but also consume a large amount of computing resources when processing them. If a method can be found to reduce or avoid the generation of such redundant features, then the computational amount required for the model to process images can be significantly reduced, thereby improving the overall efficiency. In other words, by optimizing the feature extraction process, unnecessary computational overhead can be reduced to achieve more efficient image processing. Therefore, in this embodiment, DSConv is selected to replace the standard convolution in the original C3 module, and the lightweight C3DSConv module is proposed.
[0109] C3DSConv module: Standard convolution usually creates a large number of similar feature maps, resulting in expensive computational cost and resource consumption. To solve this problem, the standard convolution in the C3 module is replaced with DSConv. Thus, while maintaining the detection accuracy of the model and increasing its applicability and generalization ability, the computational amount and the number of parameters are greatly reduced, improving the computational efficiency of the model and increasing the model training and inference speed. In the proposed C3DSConv module, DSConv is used in combination with the DSBottleneck module to enhance the feature extraction ability. The stride of the first convolutional layer of the C3DSConv module is 2, which can halve the size of the feature map, thereby quickly reducing the resolution of the feature map and reducing the subsequent computational amount. The strides of the second and third convolutional layers are 1, which are mainly used to further extract and refine features, which helps the model better capture the detailed information in the image and improve the accuracy of defect detection. At the same time, the step of retaining the original feature map in the original model is removed, thereby simplifying the model structure and reducing the computational amount. The overall structure is as Figure 3 shown.
[0110] In this embodiment, by combining the SimAM module and the ASPP module, this design realizes a comprehensive optimization of the feature extraction and fusion capabilities in complex scenarios. The SimAM module highlights key regions through a self-similarity mechanism, and the ASPP module uses multi-scale dilated convolutions to enhance the adaptability to defects of different sizes, which is particularly suitable for small target detection and occlusion scenarios. The SimAM module generates attention weights to highlight key regions by calculating the similarity between each pixel and its neighboring pixels, and the calculation formula is shown in Equation (2). The ASPP module performs parallel operations with multiple dilated convolution sampling rates, Figure 4 shows the structure of the ASPP module, which can extract rich features at different scales, thereby enhancing the detection ability for various sizes of defects. Especially when encountering small objects and blurred images in industrial scenarios, it can effectively improve the accuracy. The SimAM module preferentially provides a weighted input for the ASPP module, and then performs multi-scale information extraction and fusion after highlighting the key features.
[0111]
[0112] where, w t and b t respectively represent the attention weight and bias of the i-th pixel; x i and t represent other neurons and target neurons in a single channel of the input feature ; M = H × W represents the neurons on this channel.
[0113] The combination of the SimAM and ASPP modules significantly improves the small target detection accuracy and robustness of the model in complex scenarios through the collaborative optimization of the parameter-free attention mechanism and multi-scale feature extraction, while maintaining the efficient operation of the lightweight design.
[0114] In the previous steps, the model was lightweighted by replacing the convolution module and deleting the steps to retain the original feature map, but this caused a certain loss of accuracy. Therefore, it is necessary to restore the accuracy of the model. Traditional accuracy restoration training methods are similar to ordinary model training, but lack richer prior or supervision information, so it is difficult to achieve better performance. To solve this problem, this embodiment introduces the method of knowledge distillation, that is, extracting distilled knowledge from a high-precision teacher model and transferring it to the lightweight model to imitate the teacher network, which can significantly improve the performance of the lightweight model.
[0115] However, in the current knowledge distillation methods, there is a problem of low learning efficiency of the student model, which is because there are large differences in structure and parameters between the high-precision teacher model and the student model. To solve this problem, this embodiment proposes a new knowledge distillation method based on the KO auxiliary model, and introduces this model to achieve the purpose between the teacher and student models. The architecture of knowledge distillation is as Figure 5as shown
[0116] The KO model aims to achieve more efficient knowledge transfer by combining the advantages of CNN and Transformer. The overall framework is as Figure 6 shown. When the original Transformer model processes images, it uses the method of dividing the input image into non-overlapping patches, and only processes the information inside the patches through linear transformation. Although this method is simple and efficient, it has limitations in capturing global image information and local details. To overcome this limitation, three 3×3 convolutional blocks are introduced. Among them, the first convolutional block uses a stride of 2 to reduce the size of the input image and achieve downsampling, thereby effectively expanding the receptive field of the subsequent convolutional layers. The next two 3×3 convolutional blocks use a stride of 1 and focus on extracting local features of the image. This design not only retains the ability of the Transformer model to process global information, but also enhances the model's perception of image details through the local feature extraction ability of CNN. According to the design concept of modern CNN models, the KO model is further divided into five stages and extracts features and reduces parameters through the MBConv (Mobile Inverted Bottleneck Convolution) module, as Figure 6 (b) shown, where each stage generates feature maps of different sizes. These feature maps not only enrich the model's representation ability, but also provide multi-level supervision information for the knowledge distillation process. During the knowledge distillation process, the teacher model transfers more comprehensive and refined knowledge to the student model through these multi-level feature maps.
[0117] To generate hierarchical representations, the KO model applies a patch embedding layer consisting of depthwise separable convolution DSConv and layer normalization (LayerNorm, LN) before each stage. The aim is to reduce the size of the intermediate features, specifically by downsampling by a factor of 2 to reduce the resolution, thereby effectively expanding the receptive field of the subsequent layers. At the same time, this embedding layer also projects the features into a larger dimension, that is, doubling the size, to increase the model's expressive ability. Inside each stage, multiple KT blocks are stacked in sequence for feature transformation. In each stage, the KT blocks process the features while maintaining the same resolution, so that the model can perform deep feature extraction and transformation while maintaining details. Taking "Stage 3" and "Stage 4" of the KO model as examples, they contain 6 and 9 KT blocks respectively. Through the stacking of these KT blocks, the model can gradually refine more abstract and high-level feature representations, thereby supporting more complex image recognition and understanding tasks. The model ends with a global average pooling layer, a fully connected layer, and a classification layer with softmax.
[0118] The KT module consists of a Local Perception Unit (LPU), a Lightweight Multi-head Self-attention (LMHSA) module, and a Feature Map Generator (FGM). Its structure diagram is as shown in Figure 6 (c).
[0119] In traditional vision transformers, absolute position encoding is used to introduce unique position information for each patch, which destroys translational invariance because it changes the relative relationship between patches. In addition, vision transformers also ignore the local structure and relationship information within patches. To address these issues, the LPU is used, whose goal is to extract local relevant information without affecting the global structure of the patches. The LPU learns to obtain the local context around each patch and uses this information to enhance the model's ability to understand local features and relationships. This approach helps to maintain the invariance of the model to operations such as rotation and translation, while improving its performance in complex visual scenarios.
[0120] In the KT block, the input image is divided into non-overlapping patches and processed by the LMHSA module, and the final output is a sequence of feature vectors. These sequences of feature vectors are abstract representations of the local and global features of the image, rather than directly encoding features spatially like the feature maps in CNN models. For this reason, the FGM is proposed, whose main function is to generate a feature map from the sequence of feature vectors, that is, to convert the abstract feature vectors into a specific two-dimensional feature map. Its structure diagram is as shown in Figure 7 shown.
[0121] In the first stage, an initial knowledge transfer foundation is established through knowledge distillation between the Teacher model and the KO model. In the second stage, knowledge is transferred from the KO model to the Student model, namely YOLO-DSA. In the second stage, to improve the performance of knowledge distillation, this embodiment adopts an attention-based fusion (ABF) module. The ABF module merges feature maps from different stages in knowledge distillation, thereby improving the distillation performance. Its structure is as shown in Figure 8 shown.
[0122] In the first stage of knowledge distillation, the L2 distance is used to compare the gap between two feature maps. However, in the second stage of knowledge distillation, the feature information at different levels is aggregated to learn from the teacher, but it is not sufficient to effectively transfer the feature information at different levels. Therefore, this embodiment proposes CL as the loss function for the second stage. The structure of CL is as shown inFigure 9 As shown, the loss function transforms knowledge into feature information at different levels by using spatial pyramid pooling, making it easier to extract information. Then, the L2 distance is used to perform knowledge distillation between these levels respectively.
[0123] Finally, the detection module inputs the defect data set into the industrial defect detection model to complete the detection of industrial defects.
[0124] The embodiments described above are only descriptions of the preferred embodiments of the present invention and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.
Claims
1. An industrial defect detection method based on YOLOv5 and knowledge optimization assisted distillation, characterized in that the steps Including: Collecting a defect dataset in the actual production environment; Constructing an industrial defect detection model based on YOLOv5 and knowledge-optimized auxiliary distillation. The industrial defect detection model includes three parts: a backbone network, a neck network, and a detector head. In the backbone network, the C3DSConv module is adopted, and at the same time, the combination of the SimAM module and the ASPP module is introduced. The neck network adopts a path aggregation network and the C3DSConv module. The detector head realizes the regression and classification functions of the target box through 1×1 convolution; Inputting the defect dataset into the industrial defect detection model to complete the detection of industrial defects.
2. The industrial defect detection method based on YOLOv5 and knowledge optimization assisted distillation according to claim 1, wherein The defect types in the defect dataset include: chipping, cracks, scratches, grooves, holes, and edge sealing defects. After the collection of the defect dataset is completed, data augmentation is performed through preprocessing, and it is divided into a training set and a validation set according to a ratio of 7:
3.
3. The industrial defect detection method based on YOLOv5 and knowledge-optimized auxiliary distillation according to claim 1, wherein Replacing the standard convolution in the C3 module in the original YOLOv5 with DSConv to obtain the C3DSConv module. In the C3DSConv module, DSConv is used in combination with the DSBottleneck module to enhance the feature extraction ability. The stride of the first convolutional layer of the C3DSConv module is 2, which is used to halve the size of the feature map, quickly reduce the resolution of the feature map, and reduce the subsequent computational amount. The strides of the second and third convolutional layers are 1, which are used to further extract and refine features.
4. The industrial defect detection method based on YOLOv5 and knowledge-optimized auxiliary distillation according to claim 1, wherein By combining the SimAM module and the ASPP module, it is used to optimize the feature extraction and fusion ability in complex scenarios; Among them, the SimAM module generates attention weights to highlight key regions by calculating the similarity between each pixel and its neighboring pixels: Among them, w t and b t represent the attention weight and bias of the i-th pixel respectively; x i and t represent other neurons and the target neuron in a single channel of the input feature ; M = H × W represents the neurons on this channel; ASPP enhances the detection ability for various sizes of defects through parallel operations with multiple atrous convolution sampling rates.
5. The industrial defect detection method based on YOLOv5 and knowledge optimization-assisted distillation according to claim 1, wherein, Adopting a KO auxiliary model to improve the accuracy of the industrial defect detection model. The KO auxiliary model includes: a local perception unit, a lightweight multi-head self-attention module, and a feature map generation module. The KO model applies a patch embedding layer composed of depthwise separable convolution DSConv and layer normalization before each stage.
6. An industrial defect detection system based on YOLOv5 and knowledge optimization assisted distillation, the system is used to implement the method described in any one of claims 1-5, characterized in that, Including: A collection module, a construction module, and a detection module; The collection module is used to collect a defect dataset in the actual production environment; The construction module is used to construct an industrial defect detection model based on YOLOv5 and knowledge-optimized auxiliary distillation. The industrial defect detection model includes three parts: a backbone network, a neck network, and a detector head. In the backbone network, the C3DSConv module is adopted, and at the same time, the combination of the SimAM module and the ASPP module is introduced. The neck network adopts a path aggregation network and the C3DSConv module. The detector head realizes the regression and classification functions of the target box through 1×1 convolution; The detection module is used to input the defect dataset into the industrial defect detection model to complete the detection of industrial defects.
7. The industrial defect detection system based on YOLOv5 and knowledge optimization-assisted distillation according to claim 6, wherein, The defect types in the defect dataset include: chipping, cracks and scratches, grooves, holes, and edge sealing defects; after the defect dataset is collected, the acquisition module performs data augmentation through preprocessing and divides it into a training set and a validation set according to a ratio of 7:
3.
8. The industrial defect detection system based on YOLOv5 and knowledge optimization assisted distillation according to claim 6, characterized in that, The construction module replaces the standard convolution in the C3 module in the original YOLOv5 with DSConv to obtain the C3DSConv module; in the C3DSConv module, DSConv is used in combination with the DSBottleneck module to enhance the feature extraction ability; the stride of the first convolutional layer of the C3DSConv module is 2, which is used to halve the size of the feature map and reduce the resolution of the feature map; the strides of the second and third convolutional layers are 1, which are used to further extract and refine features.