Defect detection method and device

By adding coordinate attention module to the YOLOv5 network model and combining SIoU loss and NWD measurement learning, the accuracy improvement and robustness problems in complex scenarios and small-objective defect detection are solved, and more efficient defect detection performance is achieved.

CN120031884APending Publication Date: 2025-05-23NAT UNIV OF DEFENSE TECH
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510517051.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The prior art has problems such as improving accuracy, robustness and small-objective recognition in complex scenarios and small-objective defect detection.

Method used

The YOLOv5 network model is adopted, and the coordinate attention module is added after the C3 module of the neck network. Combined with SIoU loss and normalized Wastherstan distance (NWD) metric learning, and a hybrid attention mechanism is used to perform adaptive feature extraction across spatial scales.

Benefits of technology

It successfully overcomes the problem of detecting defects in complex scenarios and small targets, significantly improves the stability of bounding boxes and sensitivity to small defects, and improves detection accuracy and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120031884A_ABST
    Figure CN120031884A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of target detection, and relates to a defect detection method and device. The method comprises the following steps: acquiring an original image containing defects, and processing the original image to obtain defect features corresponding to the original image; the defect features comprise defect positions and defect types; generating a data set according to the original image and the corresponding defect features; obtaining a YOLOv5 network model, wherein the network model comprises an input end, a backbone network, a neck network, a head network and an output end which are connected in sequence; adding more than one coordinate attention module behind a C3 module of the neck network to obtain a network model to be trained; inputting the data set into a to-be-trained network model, and training the to-be-trained network model to obtain a trained network model; and performing defect detection by adopting the trained network model. The problem of defect detection of complex scenes and small targets can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of target detection technology, and in particular to a defect detection method and device. Background Art

[0002] The service performance of high-end equipment such as aerospace, medical equipment, and military equipment depends to a large extent on the performance of its components. Most components serve in extremely harsh environments and are required to have strong load-bearing, extreme heat resistance, lightweight, strong corrosion resistance, and high reliability. This poses very severe challenges to the materials, structure, process, and performance of the components.

[0003] Among them, due to serious process instability during the component production process, defects such as pores, poor fusion, spheroidization, cracks and surface roughness are easily generated, resulting in the scrapping of the workpiece and even the failure of the entire product. Therefore, it is of great significance to study defect detection technology.

[0004] Traditional defect detection technologies, such as optical microscopy or ultrasonic flaw detection, are often limited by the operator's experience and skills, making it difficult to achieve large-scale, efficient automated inspection. This is not only time-consuming, but may also lead to quality consistency issues.

[0005] In order to detect the generation of defects in real time and eliminate waste products in time, the development of machine learning and deep neural networks has provided new solutions to complete the detection and classification of defects.

[0006] In the existing technology, in recent years, especially in the object detection framework based on convolutional neural network (CNN) such as YOLO series, the application of machine learning algorithms in the field of defect detection has become increasingly apparent. For example: an improved YOLOv3-Tiny model is proposed, using K-means to improve regression accuracy and enhance the network's ability to extract feature information; the backbone network is improved, a fusion attention mechanism module is embedded, attention is paid to channel and spatial information, the FocalLoss loss function is introduced, the problem of uneven sample distribution is solved, and an improved algorithm of YOLOv4 is proposed.

[0007] However, despite significant progress, existing technologies still have problems in aspects such as accuracy improvement in low-resolution images, robustness in complex backgrounds, and small object recognition. Summary of the invention

[0008] Based on this, it is necessary to provide a defect detection method and device to address the above technical problems, which can overcome the difficulties of defect detection in complex scenes and small targets and perform defect detection.

[0009] A defect detection method, comprising: Acquire an original image containing defects, process the original image, and obtain defect features corresponding to the original image; the defect features include defect locations and defect categories; generate a data set based on the original image and the corresponding defect features; Obtain a YOLOv5 network model, wherein the network model includes: an input end, a backbone network, a neck network, a head network, and an output end connected in sequence; add one or more coordinate attention modules after the C3 module of the neck network to obtain a network model to be trained; Inputting the data set into the network model to be trained, training the network model to be trained, and obtaining a trained network model; Use the trained network model to perform defect detection.

[0010] In one embodiment, the neck network includes 4 C3 modules; One or more coordinate attention modules are added after the C3 module of the neck network to obtain a network model to be trained, including: A coordinate attention module is added after the second C3 module of the neck network to obtain the network model to be trained.

[0011] In one embodiment, the neck network includes 4 C3 modules; One or more coordinate attention modules are added after the C3 module of the neck network to obtain a network model to be trained, including: A coordinate attention module is respectively added after the second C3 module, the third C3 module and the fourth C3 module of the neck network to obtain a network model to be trained.

[0012] In one embodiment, the neck network includes 4 C3 modules; One or more coordinate attention modules are added after the C3 module of the neck network to obtain a network model to be trained, including: A coordinate attention module is added after the second C3 module and the third C3 module of the neck network respectively, and a selective convolution kernel attention module is added after the fourth C3 module of the neck network to obtain a network model to be trained.

[0013] In one embodiment, the loss calculated by the head network includes: calculating a total cost loss; The total cost loss includes: distance cost, angle cost and shape cost; Calculation of total cost loss includes: ; In the formula, is the total cost loss, is the intersection-over-union ratio of the predicted box and the true box, is the distance cost, is the shape cost.

[0014] In one embodiment, the distance cost is: ; in, ; ; ; ; In the formula, is the distance cost, For the direction, here we take and , for Axis direction, for Axis direction, is the angle loss correlation coefficient, which is used to adjust the weight of the distance loss according to the angle difference. is the square of the relative distance, is the angle cost, for exist Pick The value of For the real frame The center position of the axis, For the prediction box The center position of the axis, is the width of the minimum bounding rectangle of the prediction box, for exist Pick The value of For the real frame The center position of the axis, For the prediction box The center position of the axis, is the height of the minimum bounding rectangle of the prediction box.

[0015] In one embodiment, the shape cost is: ; In the formula, is the shape cost, For the direction, here we take and , is the width direction, is the height direction, To calculate the relative difference in height and width, Increase the penalty for the coefficients of the power calculation part.

[0016] In one embodiment, the loss calculated by the head network includes: calculating a total cost loss and then calculating a total distance loss; a weight ratio of the total cost loss to the total distance loss is 0.512:0.488; The total cost loss includes: distance cost, angle cost and shape cost; The total distance loss includes: normalized Wasserstein distance loss.

[0017] In one embodiment, calculating the total distance loss includes: Constructing prediction boxes And the real frame ,in, is the center coordinate of the prediction box value, is the center coordinate of the prediction box value, is the prediction box width, is the predicted box height, is the center coordinate of the real frame value, is the center coordinate of the real frame value, is the actual frame width, is the real frame height; Define the area within the predicted box and the true box as a two-dimensional probability distribution; According to the two-dimensional probability distribution, the Wasserstein distance between the predicted box and the true box is calculated; The Wasserstein distance is normalized and the normalized Wasserstein distance is used as the total distance loss.

[0018] A defect detection device, comprising: An image module is used to obtain an original image containing defects, process the original image, and obtain defect features corresponding to the original image; the defect features include defect locations and defect categories; and generate a data set based on the original image and the corresponding defect features; A model module is used to obtain a YOLOv5 network model, wherein the network model includes: an input end, a backbone network, a neck network, a head network, and an output end connected in sequence; one or more coordinate attention modules are added after the C3 module of the neck network to obtain a network model to be trained; A training module, used for inputting the data set into the network model to be trained, training the network model to be trained, and obtaining a trained network model; The detection module is used to perform defect detection using the trained network model.

[0019] The above defect detection method and device have successfully overcome the difficulties of defect detection in complex scenes and small targets, and proposed a set of systematic and comprehensive solutions, which effectively make up for the shortcomings of current algorithms in these key aspects. Specifically: SIoU loss and normalized Wasserstein distance (NWD) metric learning are combined to enhance the stability of the bounding box and sensitivity to small defects, and a hybrid attention mechanism combining coordinate attention (CA) and selective convolution kernel attention (SK) is adopted to achieve adaptive feature extraction across spatial scales. This application can fundamentally promote the overall leap and upgrade of various industries, especially the additive manufacturing industry, in the two core aspects of quality control and production efficiency improvement, thereby building a solid theoretical foundation and technical support architecture for the long-term and sustainable development of the additive manufacturing field, leading the field to a new height of development and stimulating more innovative applications and research explorations. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 is a schematic flow chart of a defect detection method in an embodiment; Figure 2 A network model architecture diagram of a defect detection method in an embodiment; Figure 3 A specific structural diagram of a network model of a defect detection method in an embodiment; Figure 4 A schematic diagram of a CA attention mechanism of a defect detection method in one embodiment; Figure 5 A schematic diagram of an SK attention mechanism of a defect detection method in one embodiment; Figure 6 A schematic diagram of the architecture of a defect detection method in an embodiment; Figure 7 Graphs showing experimental results of different models in a specific embodiment; Figure 8 is one of the two-dimensional ablation experiment result diagrams of different models in a specific embodiment; Fig. 9 FIG2 is a second diagram of two-dimensional ablation experimental results of different models in a specific embodiment; Fig.10 FIG3 is a three-dimensional ablation experiment result diagram of different models in a specific embodiment; Fig.11 FIG4 is a four-dimensional ablation experiment result diagram of different models in a specific embodiment; Fig.12is a diagram of three-dimensional ablation experimental results of different models in a specific embodiment; Fig.13 This is one of the loss regression graphs when epochs=100 in a specific embodiment; Fig.14 This is the second loss regression graph when epochs=100 in a specific embodiment; Fig.15 is one of the comparison result diagrams of different loss function models in a specific embodiment; Fig.16 FIG2 is a second comparison result diagram of different loss function models in a specific embodiment; Fig.17 FIG3 is a comparison result diagram of different loss function models in a specific embodiment; Fig.18 FIG4 is a fourth diagram of comparison results of different loss function models in a specific embodiment. DETAILED DESCRIPTION

[0021] In order to make the purpose, technical solutions and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application.

[0022] In addition, the descriptions of "first", "second", etc. in this application are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of this application, "multiple groups" means at least two groups, such as two groups, three groups, etc., unless otherwise clearly and specifically defined.

[0023] In this application, unless otherwise clearly specified and limited, the terms "connection", "fixation", etc. should be understood in a broad sense. For example, "fixation" can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection, an electrical connection, a physical connection, or a wireless communication connection; it can be a direct connection, or an indirect connection through an intermediate medium, or it can be the internal connection of two elements or the interaction relationship between two elements, unless otherwise clearly defined. For ordinary technicians in this field, the specific meanings of the above terms in this application can be understood according to specific circumstances.

[0024] In addition, the technical solutions between the various embodiments of the present application can be combined with each other, but it must be based on the fact that ordinary technicians in the field can implement it. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such combination of technical solutions does not exist and is not within the scope of protection required by this application.

[0025] The present application provides a defect detection method, such as Figure 1 The flowchart shown, in one embodiment, includes: Step 101, obtain an original image containing defects, process the original image, and obtain defect features corresponding to the original image; the defect features include defect locations and defect categories; generate a data set based on the original image and the corresponding defect features.

[0026] In this step, the existing technology is used to extract and photograph the defects of the workpiece to obtain the original image containing the defects; the original image is processed, and the defect location information and defect category information in the processed image are marked with Label1Img (existing technology) as defect features, and saved in a txt file.

[0027] Step 102, obtaining a YOLOv5 network model, the network model includes: an input end, a backbone network, a neck network, a head network and an output end connected in sequence; adding one or more coordinate attention modules after the C3 module of the neck network to obtain a network model to be trained.

[0028] Specifically: Obtain the YOLOv5 network model, which includes the following connected in sequence: input end, backbone network, neck network, head network, and output end; The cervical network includes 4 C3 modules; A coordinate attention module is added after the second C3 module of the neck network to obtain the network model to be trained; Alternatively, a coordinate attention module is added after the second C3 module, the third C3 module, and the fourth C3 module of the neck network respectively to obtain the network model to be trained; Alternatively, a coordinate attention module is added after the second C3 module and the third C3 module of the neck network respectively, and a selective convolution kernel attention module is added after the fourth C3 module of the neck network to obtain the network model to be trained.

[0029] The losses calculated by the head network include: calculating the total cost loss, or calculating the total cost loss and then calculating the total distance loss; the total cost loss includes: distance cost, angle cost and shape cost; the total distance loss includes: normalized Wasserstein distance loss (normalized Wasserstein distance loss); when calculating the total cost loss and then the total distance loss, the weight ratio of the total cost loss to the total distance loss is 0.512:0.488, and the sum of the weight ratios of the total cost loss and the total distance loss is 1; Calculation of total cost loss includes: ; In the formula, is the total cost loss, is the intersection-over-union ratio of the predicted box and the true box, is the distance cost, is the shape cost; The distance cost is: ; in, ; ; ; ; In the formula, is the distance cost, For the direction, here we take and , for Axis direction, for Axis direction, is the angle loss correlation coefficient, which is used to adjust the weight of the distance loss according to the angle difference. is the square of the relative distance, is the angle cost, for exist Pick The value of For the real frame The center position of the axis, For the prediction box The center position of the axis, is the width of the minimum bounding rectangle of the prediction box, for exist Pick The value of For the real frame The center position of the axis, For the prediction box The central position of the axis is the height of the minimum bounding rectangle of the prediction box; The shape cost is: ; In the formula, is the shape cost, is the direction, here take and , is the width direction, is the height direction, is to calculate the relative difference between the height and the width, is the coefficient of the power calculation part, increasing the penalty; Calculating the total distance loss includes: Constructing the prediction box and the ground truth box , where, is the center coordinate value of the prediction box, is the center coordinate value of the prediction box, is the width of the prediction box, is the height of the prediction box, is the center coordinate value of the ground truth box, is the center coordinate value of the ground truth box, is the width of the ground truth box, is the height of the ground truth box; Define the regions inside the prediction box and the ground truth box as two-dimensional probability distributions; According to the two-dimensional probability distribution, calculate the Wasserstein distance between the prediction box and the ground truth box; Normalize the Wasserstein distance and use the normalized Wasserstein distance as the total distance loss.

[0030] In this step, the YOLOv5 network model belongs to the prior art and is an efficient single-stage object detection algorithm, and the specific content will not be elaborated here.

[0031] Such as Figure 2 and Figure 3As shown, the network model to be trained obtained in this application is called the SCK-YOLOv5 network model, including: an input end, a backbone network, a neck network, a head network and an output end, which improves the defect detection performance; specifically: the input end uses Mosaic enhancement, adaptive anchor points and dynamic scaling strategies to effectively deal with data imbalance and improve the model running speed and accuracy; the backbone network is composed of a series of complex convolution blocks (such as CBS, C3 and SPPF), which is responsible for deep feature extraction and generating feature maps; the neck network uses a feature pyramid and a path aggregation network to enhance the diversity of feature representation and the robustness of defect detection; the head network works together through three Detect modules; the output end outputs the final prediction result including position, confidence and category.

[0032] The total cost loss in this application is called the SIoULoss loss function, which takes into account the direction of the mismatch between the required true box and the predicted box, and fully considers the vector angle between the required regressions, thereby redefining the penalty indicators, including distance cost, angle cost, and shape cost, thereby improving the convergence rate and efficiency.

[0033] The total distance loss in this application is called the NWDLoss loss function. Its underlying logic is based on the Wasserstein distance (existing technology), a measure of the "distance" between two probability distributions, which takes into account the minimum effort required to convert from one distribution to another. The main idea is to regard the goal of the generative model as a process of making the generated image distribution as close as possible to the real data distribution; compared with traditional losses such as MSE (mean square error) or KL divergence, Wasserstein distance pays more attention to the overall structural consistency and fidelity of details of the generated image; during the calculation process, the model will try to minimize the "transportation cost" between the generated image and the real image. NWDLoss optimizes A dual-input network architecture is constructed in which a generator tries to generate images similar to the training samples, while a discriminator evaluates their differences. The generator is updated to reduce the discriminator's prediction, while the discriminator is updated to maximize its prediction. The two compete with each other to push the generator to gradually approach the real data distribution. For small target defect detection, the Wasserstein distance is used to more accurately measure the similarity of bounding boxes. Therefore, the NWDLoss loss function is introduced, which regards each bounding box as a two-dimensional probability density function, that is, a two-dimensional Gaussian distribution. By calculating the normalized Wasserstein distance between the two Gaussian distributions, the similarity between them can be quantified and evaluated. This comparison strategy based on high-dimensional probability distribution is particularly good at finely distinguishing small changes in small targets, so it can significantly improve the accuracy of defect detection and adaptability to complex scenarios, and enhance the robustness of the system.

[0034] The coordinate attention module in this application integrates the coordinate attention (CA) mechanism; Figure 4 As shown in the figure, it includes two major steps, namely the embedding of coordinate information and the generation of coordinate attention. In the embedding of coordinate information, specifically, the position-related information is embedded in the input feature map, and in the generation of coordinate attention, the attention map is constructed based on the existing position information, and then the attention is applied to the input feature map to highlight the representation content worthy of attention. The main purpose of CA is to encode precise position details into the neural network system, so as to achieve modeling operations for channel associations and long-term dependencies. The advantage is that when in a mobile network environment, it can model large areas and effectively avoid large-scale computing costs. In short, it not only focuses on the interaction between channels and takes into account the directionality of position information, but also has flexibility and efficiency. It can be seamlessly integrated into the key modules of lightweight networks. While paying attention to spatial and channel dimensions, it also solves the problem of long-range dependence. While maintaining high accuracy, it can also effectively control the amount of parameters and computational burden.

[0035] The selective convolution kernel attention module in this application integrates the selective convolution kernel attention (Selective Kernel Attention, SK) mechanism; Figure 5 As shown in the figure, the convolution kernel size is dynamically adjusted through the selective kernel unit SK to adapt to targets of different scales. This mechanism allows the network to capture multi-scale features of the image space more effectively while reducing the waste of computing resources. SK-Net contains three operations: Split, Fuse and Select to achieve the fusion and selection of multi-scale information. It can be seen that the Split operation first processes the input features through convolution kernels of different sizes to capture multi-scale feature information. Then, these different scales are spliced ​​together and enter the Fuse operation for global pooling, dimensionality reduction and dimension increase, and softmax normalization to generate weights of different scales. Finally, the Select operation convolves the weight of each scale with the Split output by weighted summation to select the most effective feature scale, thereby achieving focused detection. In short, this is a selective mechanism that utilizes dynamic convolution kernel size. The core is to allow neurons to adaptively adjust the receptive field according to multi-scale inputs, thereby enhancing the network's ability to process multi-scale features. This attention mechanism can not only improve model performance, but also show good efficiency characteristics.

[0036] Step 103, inputting the data set into the network model to be trained, training the network model to be trained, and obtaining a trained network model.

[0037] In this step, how to train the network model belongs to the existing technology and will not be described in detail here.

[0038] Step 104: Use the trained network model to perform defect detection.

[0039] In this step, the image that needs to be defect detected is input into the trained network model, and the output of the trained network model is the defect detection result.

[0040] In this embodiment, if Figure 6 As shown in the figure, in order to further improve the detection accuracy of defects, this application modifies the network structure of YOLOv5 based on YOLOv5s: (1) The SIoULoss loss function and the NWDLoss loss function are integrated, and a new loss function SNWDLoss is creatively proposed to accelerate network convergence and improve robustness. (2) The CA attention mechanism and the SK attention mechanism are used to complement each other and the two are combined into the neck network. This effectively enhances the algorithm's attention to local features and its ability to organize global sequence information.

[0041] The above defect detection method has the following beneficial effects: (1) A flexible and resource-efficient CA attention mechanism is designed and seamlessly integrated into the neck network in the core architecture of mobile networks. It cleverly integrates channel attention and direction-aware location information, significantly improving the model's accuracy in positioning and recognition, and showing excellent versatility and performance improvement in diverse tasks such as defect detection. In addition, in order to improve the model's computational efficiency, the SK attention mechanism is introduced into the neck network to select and focus on the key parts of the input feature map for the task, better preserve the spatial structure of the image, and enhance the model's adaptability and generalization to various scenarios; at the same time, the SK attention mechanism can be oriented to specific areas such as defects, which can reduce the risk of overfitting.

[0042] (2) SIoULoss loss function is used. As a loss function under semi-supervised learning, SIoULoss guides the model to learn more representative features from the only model under the condition of a small amount of metal defect data set. In addition, NWDLoss loss function is introduced into SIoULoss loss function, which takes into account the influence of noise data and dynamically adjusts the weight of each sample. This mechanism can adaptively process noise points in defect data sets and reduce the impact of noise on model training.

[0043] (3) In order to improve the accuracy and production efficiency of defect detection of additively manufactured parts and improve the training accuracy, this application designs a new defect recognition network structure, namely SCK-YOLOv5. When performing defect detection, the attention mechanism is added, which can make the model ignore the smooth part of the material and pay more attention to the "defects", thereby improving the recognition speed and accuracy of the detection; specifically, the classic YOLOv5 architecture is optimized, and the CA attention mechanism and SK attention mechanism are added in particular; in addition, by adopting the improved version of SIoULoss instead of the traditional CIoULoss, not only the training process is accelerated, but also the accuracy and anti-interference ability of the model are improved, making the positioning more accurate; at the same time, the two loss functions of SIoULoss and NWDLoss are combined, and the SNWDLoss loss function model is innovatively proposed. This loss function based on the Wasserstein distance is essentially a measure of the "migration cost" between two predictions and the true distribution, which especially enhances the model's perception of small sample changes, significantly improves the stability and consistency in the recognition of similar samples, and effectively improves the overall detection performance. This improvement significantly enhances the effect of the model in practical applications.

[0044] This application successfully overcomes the difficulty of defect detection in complex scenes and small targets, and proposes a set of systematic and comprehensive solutions, which effectively make up for the shortcomings of current algorithms in these key aspects. Specifically: SIoU loss and normalized Wasserstein distance (NWD) metric learning are combined to enhance the stability of bounding boxes and sensitivity to small defects, and a hybrid attention mechanism combining coordinate attention (CA) and selective convolution kernel attention (SK) is adopted to achieve adaptive feature extraction across spatial scales. This application can fundamentally promote the overall leap and upgrade of various industries, especially the additive manufacturing industry, in the two core aspects of quality control and production efficiency improvement, thereby building a solid and solid theoretical foundation and technical support architecture for the long-term and sustainable development of the additive manufacturing field, leading the field to a new level of development and stimulating more innovative applications and research explorations.

[0045] It should be understood that although Figure 1 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 1At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.

[0046] In a specific embodiment, the method of the present application is verified.

[0047] 1. Dataset.

[0048] The training dataset used in this application only contains metal additive void defects. The dataset has 1,000 void defect instances, and each instance is annotated with precise bounding boxes that provide comprehensive details about the target location, size, and category.

[0049] In addition, this application takes 100 hole defect instances as a verification set for experimental verification.

[0050] 2. Experimental environment.

[0051] The experimental environment of this application uses the Windows 10 operating system, and all experiments are run on the software platform Python-3.9.19 torch-2.2.2+cu121 CUDA:0 (NVIDIA GeForce RTX 3060 Laptop GPU). The experimental parameters are set as follows: 100 epochs of training rounds, 16 batch size, 640×640 input image size, and other default settings.

[0052] 3. Performance indicators.

[0053] This application uses precision, recall, mean average precision (mAP, including mAP50 and mAP50-95), and regression loss as evaluation criteria to verify the advantages of the SCK-YOLOv5 network model.

[0054] in, TP : The number of true positives, which is positive in the label and also positive in the predicted value; TN : The number of true negatives, which is negative in the label and negative in the predicted value; FP : The number of false positives, which is negative in the label and positive in the predicted value; FN : The number of false negatives, which is positive in the label and negative in the predicted value; TP + TN + FP + FN= total number of pixels, TP + TN = Number of correctly classified pixels.

[0055] 1) Precision Precision is called inspection accuracy, also known as precision, which is the proportion of positive samples that the evaluation model predicts correctly. In defect detection, if the bounding box predicted by the model coincides with the true bounding box, the prediction is considered correct. The formula is as follows: ; 2) Recall The recall rate is also called the inspection rate. It responds to the inspection performance of the model and indicates the percentage of the correct samples predicted as positive to the actual number of positive samples. It is used to evaluate the model's ability to find all true samples. If the true bounding box coincides with the predicted bounding box, the sample is considered to be correctly recalled. The formula is as follows: ; 3) Average Precision Average precision AP is used to calculate the average precision of different categories. AP integrates the changes in precision and recall and is an important indicator for describing the advantages and disadvantages of defect detection models. The formula is as follows, where: is the precision when the recall rate is r: ; 4) Mean Average Precision (mAP) mAP is the average precision of multi-category problems. For binary classification problems, mAP=AP.

[0056] mAP50 represents the mAP value at an IoU threshold of 50%. As a comprehensive indicator, it reflects the mean of precision, recall, and average precision. The higher the mAP value, the more accurate the model. The mAP value ranges from [0,1], and the closer to 1, the better the detection effect. AP and mAP are comprehensive evaluation indicators that reflect the accuracy of the algorithm for a single category and all categories of targets. Higher AP and mAP mean that the model has higher confidence in detecting the target object.

[0057] ; mAP50-95 is a more stringent evaluation indicator that calculates the mAP value within the IoU threshold range of 50%-95% and then takes the average, which can more accurately evaluate the performance of the model at different IoU thresholds.

[0058] 5) Regression Loss train / obj_loss refers to the loss value of objectness prediction during training, i.e., training set loss, which involves the model's ability to determine whether a specific object exists in an image. val / obj_loss refers to the objectness loss value on the validation set, i.e., validation set loss, which evaluates the model's ability to detect the presence of objects on unseen data. Object existence loss usually uses Binary Cross-Entropy (BCE) loss.

[0059] In order to evaluate the accuracy of the training model proposed in this application, a large number of experiments were conducted for comparison, including: quantitative analysis, ablation experiments, loss function comparison experiments, and universal experiments, which proved that this application provides higher detection accuracy and faster speed.

[0060] 4. Quantitative analysis.

[0061] This application is compared with the YOLOv3, YOLOv5, YOLOv6, YOLOV8, and YOLOv10 models in the prior art. The experimental results of different models are shown in Tables 1 and Figure 7 shown.

[0062] Table 1: Experimental results of different models

[0063] It can be clearly seen from Table 1 that compared with other baseline models, the overall performance of the YOLOv5 model is better than other models. The accuracy of YOLOv3 is around 0.7, and the detection effects of YOLOv5m, YOLOv5n, YOLOv6, and YOLOv8 are relatively similar, and their detection accuracy can reach about 0.75-0.8. The recall rate of YOLOv10 is relatively high, reaching 0.903, but its precision and average precision are relatively low, 0.637 and 0.771, respectively, which are 35.2% and 16.3% lower than this application. The detection accuracy of YOLOv5s reaches 0.984, and its recall rate and average precision mean reach 0.873 and 0.916, which is much better than other models, but compared with this application, its accuracy is 0.5% lower, the recall rate is 1.2% lower, and the average precision is 1.8% lower.

[0064] from Figure 7It can be clearly seen that YOLOv3 is not as good as YOLOv5 in the field of visual recognition. Due to the imperfection of their code algorithms, YOLOv6, YOLOv8, and YOLOv10 are not as perfect as YOLOv5 in the field of small target recognition. However, among the YOLO algorithms that have never been improved, YOLOv5s has the best overall performance. This application is based on YOLOv5s and improves its accuracy by 0.5%, its recall rate by 1.2%, and the correctness and coverage of the surface model. At the same time, this application can currently achieve an average accuracy improvement of 1.8%, which shows the correctness and effectiveness of this application.

[0065] In order to more intuitively prove that this application is effective compared with other models, several representative YOLO model detection cases are selected for qualitative analysis, including YOLOv3, YOLOv6, YOLOv5, YOLOv8, YOLOv10 and SCK-YOLOv5 of this application, and specific representative images containing defects are selected from the test set.

[0066] Comparison of the results of the above models: the accuracy of YOLOv3 is mostly around 0.3, the recognition accuracy is low and there are a lot of false detections; the accuracy of YOLOv6 is improved to 0.4 compared with YOLOv3, but there are still a lot of false detections; the accuracy of YOLOv5m and YOLOv5n is around 0.6, although there is no missed detection or false detection, but its recognition accuracy is poor; the accuracy of YOLOv6 is around 0.5, and there are many false detections; the accuracy of YOLOv8 is mostly around 0.4, and there are a lot of false detections; the accuracy of YOLOv10 is also around 0.4, and the false detection rate is less than that of YOLOv8; the accuracy of YOLOv5s is relatively high, and the average accuracy can reach 0.9. However, this application can achieve a recognition result with a confidence level of 1, and 100% of the defects in the image are recognized, without missed detection or false detection, which reflects the superiority of the improved algorithm of this application and reflects the flexibility and accuracy of the improved loss function.

[0067] 5. Ablation experiment.

[0068] In order to verify the effectiveness of the improved model of this application, an ablation experiment was conducted. Ablation experiment is a scientific exploration tool, which analyzes the importance of each variable to the final result by changing the key variables in the experimental design one by one. This technology is mainly used to evaluate the effectiveness of new theories or improvement measures. By comparing the situations with and without a certain factor, it can be identified which parts play a decisive role in the performance of the system. Its advantage is that it provides evidence of causal relationships, helps understand the actual influence of each element in the model, and may promote the simplification and performance improvement of the model. By "eliminating" (ablating) the influencing factors one by one, the best practices of the model can be identified and optimized. This application conducted 9 sets of improved ablation experiments. The ablation test results of different models are shown in Table 2 and Figures 8 to 12 shown.

[0069] Table 2: Ablation test results of different models

[0070] As shown in Table 2, this application is improved based on YOLOv5s (a better version of YOLOv5), and the attention mechanism and loss function are integrated into it, which achieves a certain degree of improvement in accuracy, recall rate and average precision, thereby improving the performance and efficiency of defect detection.

[0071] In the 9 groups of ablation experiments, the performance of precision and recall is relatively similar. After adding a single attention mechanism and loss function, both of them showed different degrees of decline, with precision dropping from 0.98 to 0.97 and recall dropping from about 0.873 to about 0.87. However, after fusing the two loss functions (SNWDLoss), the precision rose to 0.987, which is higher than the original precision of 0.984. Subsequently, when a single attention mechanism was incorporated, the precision dropped again, even falling below 0.712 in one group. After analysis, this may be due to the fact that the CA attention mechanism failed to effectively extract feature information after SIoULoss processing, but after combining with other improvement measures, the recall rate was improved to 0.878. Finally, after fusing the four improvement measures and further optimizing their network structure, both precision and recall reached the highest values, 0.989 and 0.885 respectively.

[0072] like Fig.12 As shown in the figure, the mean average precision (mAP) of the 9 groups of ablation experiments can better reflect the advantages of this application. After adding the attention mechanism and loss function, the mAP has been improved to varying degrees, especially the CA attention mechanism and the two loss functions have a significant effect on improving the mAP. After introducing CA and the two loss functions alone, the mAP is increased to about 0.92, and then the four improvements are integrated, and the mAP reaches 0.934, which is 1.8% higher than before. Overall, the recognition effect has been significantly improved.

[0073] Analyze the improved effect mechanism. Among them, SIoULoss solves the problems in the training process of the classic cross-entropy loss on class-imbalanced datasets. Especially for small objects and difficult-to-classify objects, it has better processing ability. By adjusting the importance of negative samples, it focuses on those "difficult" samples that are difficult to distinguish, reduces the influence of background noise on the model, enables the model to focus more on learning important features, and helps improve the overall detection accuracy. On this basis, the NWDLoss function is introduced. The combination of the two helps optimize the accuracy in defect detection and instance segmentation tasks. Especially for objects with complex shapes or large size variations, the Wasserstein distance provides a more intuitive way to measure the differences in the feature space, which helps model interpretation and visualization. The CA attention mechanism focuses on location information, which allows the network to better understand the location association and context information of objects, enhances the perception of object boundaries, and thus improves the positioning accuracy. Finally, the SK attention mechanism allows the network to dynamically adjust the weight of each pixel's contribution to the target, and can adaptively select and combine different filters according to the input local features, which further improves the flexibility and pertinence of feature extraction.

[0074] Combining these four mechanisms, this application can effectively solve the problems of positioning accuracy, small target defect detection, and robustness in complex scenarios in defect detection, make up for defects such as resolution dependence, false positives, and missed detections that may occur in previous versions, improve the model's global understanding and local detail processing capabilities. They work together to optimize the model's response to various situations, enabling it to significantly improve the performance of defect detection while maintaining speed.

[0075] As Figure 13 to Figure 14 shown, the validation set loss val_Loss (validation loss) value of this application is significantly the smallest (close to 0.015) compared with other models. The lower the validation set loss, the better the understanding and fitting degree of this application for the validation data, the stronger the generalization ability, and a better performance balance is achieved. Related training strategies such as learning rate adjustment and regularization have played a role, enabling the model to maintain consistency for new samples while reducing errors.

[0076] 6. Comparative experiment of loss functions.

[0077] To verify the advantages of the SNWDLoss function proposed in this application, a comparative experiment is conducted. Compare this application with several currently popular improved loss functions (such as CIoU, WIoU, Focal_EIoU, Alpha_IoU). The comparison results of different loss function models are as Figures 15 to 18 shown.

[0078] The SNWDLoss of this application is an innovative loss function that combines SIoULoss and NWDLoss. Compared with traditional loss functions, this application takes into account the accuracy of defect edge detection and the consistency of the overall structure, and shows higher accuracy and completeness in target recognition tasks. Its precision (P), recall rate (R), and mean average precision (mAP) all achieve the best results among the above models. This is because it pays attention to both the center point and boundary information, improves the accuracy of object positioning, reduces the problems of misclassification and missed detection, and helps to reduce inter-class confusion; compared with those loss functions that only focus on a single factor, this application provides a better boundary matching mechanism, showing stronger robustness and generalization ability, and it can handle various complex scenes, including small target defect detection and background interference. Experimental results show that this application can improve the overall performance and stability of the model, improve the performance of the deep learning model in the field of defect detection, enable the model to achieve more accurate predictions in defect detection tasks, and show better performance advantages in practical applications.

[0079] 7. Universal experiment.

[0080] In order to verify that the higher recognition accuracy of this application is not only for the defect data set collected in this application, the SCK-YOLOv5 model and the YOLOv5 model were qualitatively analyzed. In addition, a steel surface defect data set (NEU-DET) was selected, which has a total of 1,800 images.

[0081] After comparison, this application has excellent performance in dealing with missed detection and false detection, and has higher precision recognition and positioning capabilities in crazing and pitted_surface defect detection. It also makes up for the missed detection problem of YOLOv5 in rolled-in_scale detection and significantly reduces the error rate. These show that this application can accurately detect metal surface defects even in complex visual scenes, enhancing the robustness and reliability of visual recognition in practical applications.

[0082] In summary, this application proposes an improved SCK-YOLOV5 defect detection model based on YOLOv5, which solves the persistent limitations of existing technologies in terms of accuracy and environmental robustness in small defect recognition. Specifically: a new SNWDLoss composite loss function is introduced, which combines the angle sensitivity of SIoU and the distribution perception characteristics of NWD, specifically for improving micron-level defect positioning; a dual attention mechanism using CA for global spatial modeling and SK convolution for adaptive feature refinement are used to enable robust multi-scale defect extraction from specular reflective surfaces; this application achieves significant performance improvements, with precision, recall and mAP50 increased by 0.5%, 1.2% and 1.8% respectively. While maintaining real-time processing efficiency, this application also has an excellent improvement in the ability to detect microscopic porosity defects. This work has established a revolutionary paradigm for intelligent quality inspection in high-precision additive manufacturing systems.

[0083] The present application also provides a defect detection device, which in one embodiment includes: an image module, a model module, a training module and a detection module, wherein: An image module is used to obtain an original image containing defects, process the original image, and obtain defect features corresponding to the original image; the defect features include defect locations and defect categories; and generate a data set based on the original image and the corresponding defect features; A model module is used to obtain a YOLOv5 network model. The network model includes: an input end, a backbone network, a neck network, a head network, and an output end connected in sequence; one or more coordinate attention modules are added after the C3 module of the neck network to obtain a network model to be trained; The training module is used to input the data set into the network model to be trained, train the network model to be trained, and obtain a trained network model; The detection module is used to perform defect detection using the trained network model.

[0084] For the specific definition of a defect detection device, please refer to the definition of a defect detection method above, which will not be repeated here. Each module in the above device can be implemented in whole or in part by software, hardware and a combination thereof. Each of the above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.

[0085] The contents not described in detail in this specification belong to the prior art known to professional and technical personnel in this field.

[0086] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0087] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A defect detection method, characterized in that: include: Acquire an original image containing defects, process the original image, and obtain defect features corresponding to the original image; The defect features include defect locations and defect categories; generating a data set based on the original image and the corresponding defect features; Obtain a YOLOv5 network model, wherein the network model includes: an input end, a backbone network, a neck network, a head network, and an output end connected in sequence; add one or more coordinate attention modules after the C3 module of the neck network to obtain a network model to be trained; Inputting the data set into the network model to be trained, training the network model to be trained, and obtaining a trained network model; Use the trained network model to perform defect detection.

2. A defect detection method according to claim 1, characterized in that: The neck network includes 4 C3 modules; One or more coordinate attention modules are added after the C3 module of the neck network to obtain a network model to be trained, including: A coordinate attention module is added after the second C3 module of the neck network to obtain the network model to be trained.

3. A defect detection method according to claim 1, characterized in that: The neck network includes 4 C3 modules; One or more coordinate attention modules are added after the C3 module of the neck network to obtain a network model to be trained, including: A coordinate attention module is respectively added after the second C3 module, the third C3 module and the fourth C3 module of the neck network to obtain a network model to be trained.

4. A defect detection method according to claim 1, characterized in that: The neck network includes 4 C3 modules; One or more coordinate attention modules are added after the C3 module of the neck network to obtain a network model to be trained, including: A coordinate attention module is respectively added after the second C3 module and the third C3 module of the neck network, and a selective convolution kernel attention module is added after the fourth C3 module of the neck network to obtain a network model to be trained.

5. A defect detection method according to any one of claims 1 to 4, characterized in that: The loss calculated by the head network includes: calculating the total cost loss; The total cost loss includes: distance cost, angle cost and shape cost; Calculation of total cost loss includes: ; In the formula, is the total cost loss, is the intersection-over-union ratio of the predicted box and the true box, is the distance cost, is the shape cost.

6. A defect detection method according to claim 5, characterized in that: The distance cost is: ; in, ; ; ; ; In the formula, is the distance cost, For direction, here we take and , for Axis direction, for Axis direction, is the angle loss correlation coefficient, which is used to adjust the weight of the distance loss according to the angle difference. is the square of the relative distance, is the angle cost, for exist Pick The value of For the real frame The center position of the axis, For the prediction box The center position of the axis, is the width of the minimum bounding rectangle of the prediction box, for exist Pick The value of For the real frame The center position of the axis, For the prediction box The center position of the axis, is the height of the minimum bounding rectangle of the prediction box.

7. A defect detection method according to claim 6, characterized in that: The shape cost is: ; In the formula, is the shape cost, For direction, here we take and , is the width direction, is the height direction, To calculate the relative difference in height and width, Increase the penalty for the coefficients of the power calculation part.

8. A defect detection method according to any one of claims 1 to 4, characterized in that: The loss calculated by the head network includes: calculating the total cost loss and then calculating the total distance loss; the weight ratio of the total cost loss to the total distance loss is 0.512:0.488; The total cost loss includes: distance cost, angle cost and shape cost; The total distance loss includes: normalized Wasserstein distance loss.

9. A defect detection method according to claim 8, characterized in that: Calculating the total distance loss includes: Constructing prediction boxes And the real frame ,in, is the center coordinate of the prediction box value, is the center coordinate of the prediction box value, is the prediction box width, is the predicted box height, is the center coordinate of the real frame value, is the center coordinate of the real frame value, is the actual frame width, is the real frame height; Define the area within the predicted box and the true box as a two-dimensional probability distribution; According to the two-dimensional probability distribution, the Wasserstein distance between the predicted box and the true box is calculated; The Wasserstein distance is normalized and the normalized Wasserstein distance is used as the total distance loss.

10. A defect detection device, characterized in that: include: An image module is used to obtain an original image containing defects, process the original image, and obtain defect features corresponding to the original image; The defect features include defect locations and defect categories; generating a data set based on the original image and the corresponding defect features; A model module is used to obtain a YOLOv5 network model, wherein the network model includes: an input end, a backbone network, a neck network, a head network, and an output end connected in sequence; one or more coordinate attention modules are added after the C3 module of the neck network to obtain a network model to be trained; A training module, used for inputting the data set into the network model to be trained, training the network model to be trained, and obtaining a trained network model; The detection module is used to perform defect detection using the trained network model.

Citation Information

Patent Citations

  • Spatial non-cooperative target component identification method based on lightweight and attention mechanism

    CN115205467A

  • Solar cell defect detection method based on attention convolutional neural network

    CN116523820A

  • Non-woven fabric defect detection method based on improved YOLOv5

    CN116912199A

  • Chip defect detection method based on improved YOLOv3 model

    CN117274775A

  • Infrared unmanned aerial vehicle group detection method based on state space model

    CN118506222A