Industrial Surface Defect Detection and Severity Level Prediction Model, Method and Application

By proposing a model including backbone network, defect detection network and defect level prediction network in industrial product surface defect detection, combined with feature processor and Soft Focal Loss loss function, the problem of inability to effectively identify defect severity levels in the prior art is solved, and flexibility and accuracy are achieved when defect standards change.

CN118333955BActive Publication Date: 2025-06-17FITOW (TIANJIN) DETECTION TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410393795.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-02
Publication Date
2025-06-17
Estimated Expiration
2044-04-02

AI Technical Summary

Technical Problem

Existing industrial product defect detection methods cannot effectively identify and predict the severity of defects, resulting in the need to re-label data and re-train the model when defect standards change.

Method used

A surface defect detection and severity level prediction model for industrial products is proposed, including backbone network, defect detection network and defect level prediction network. Multi-scale features are extracted through feature processors and trained using Soft Focal Loss loss function.

Benefits of technology

It realizes the severity level of identifying defects while detecting defect locations, which is better than the traditional multi-branch defect detection classification network, and can effectively deal with the problem of changes in defect standards.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118333955B_ABST
    Figure CN118333955B_ABST
Patent Text Reader

Abstract

The present invention provides an industrial product surface defect detection and severity level prediction model, method and application, belonging to the field of industrial inspection. The model includes: a backbone network, and a defect detection network and a defect level prediction network respectively connected thereto; the defect detection network is connected to the defect level prediction network; the backbone network is used to extract image features; the defect detection network is used to obtain the predicted defect positions according to the image features; the defect level prediction network is used to obtain the severity level of each defect according to the image features and the predicted defect positions. Using the present invention, not only can the positions of defects be detected, but also the severity levels of the defects can be identified. In addition, using the present invention, not only can the features of larger-sized defects be obtained, but also the features of small-sized defects can be effectively retained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of industrial inspection, and particularly relates to a surface defect detection and severity level prediction model, method and application for industrial products. Background Art

[0002] Industrial products are important products or intermediate raw materials in industrial production, and their quality has an important impact on the quality of downstream products and the interests of production enterprises. Common industrial products include: chemical fiber filaments, auto parts, gears, wire meshes, etc. The quality inspection link provides an important guarantee for product quality and is an important link in industrial product production. Initially, industrial quality inspection was mainly completed by manual inspection. With the rise of the concept of Industry 4.0 and the development of artificial intelligence technology, especially deep learning technology, automated quality inspection methods have emerged in an endless stream.

[0003] Most automated quality inspections obtain images of products through industrial cameras and then use image-based defect detection methods. Specifically, methods such as object detection, semantic segmentation, and anomaly detection are used to detect the defective parts in the images. Current inspection methods based on detection or segmentation generally obtain deep learning models through supervised training, and then in the industrial field, information such as the specific location, category, and confidence level of the defect can be output through this model.

[0004] However, defects in industrial products often appear in small areas of the image, and the size of the defective areas varies greatly, the significance level is low, and the semantic concept is vague, resulting in strong subjectivity of the defects. Through data annotation, the algorithm can learn the laws of defects. However, in the actual production process, the standards of defects are not fixed, which leads to the need to re-annotate data and re-train the model when the defect standards change. Therefore, it is necessary to obtain the severity level of defects during industrial defect detection, so that the personnel on the production line can screen the detected defects according to the severity level.

[0005] Currently, vision-based defect detection methods can answer questions such as "Is there a defect?", "Where is the defect?", and "What type of defect is it?", but they cannot answer questions such as "Is the defect serious?". The mainstream defect detection methods are based on the ideas of object detection and semantic segmentation. These methods can output the category and confidence level of the detected defects, but these indicators cannot represent the severity level of the defects. Summary of the Invention

[0006] The purpose of the present invention is to solve the problems existing in the above-mentioned prior art, and provide a surface defect detection and severity level prediction model, method and application for industrial products, which can identify the severity level of defects during the surface defect detection of industrial products, and facilitate quality inspection personnel in the industrial field to screen the detected defects according to different quality inspection standards.

[0007] The present invention is achieved through the following technical solutions:

[0008] In the first aspect of the invention, a detection and severity level prediction model for surface defects of industrial products is provided. The model includes: a backbone network, a defect detection network and a defect severity level prediction network respectively connected thereto; the defect detection network is connected to the defect severity level prediction network;

[0009] The backbone network is used to extract image features;

[0010] The defect detection network is used to obtain the predicted defect positions according to the image features;

[0011] The defect severity level prediction network is used to obtain the severity level of each defect according to the image features and the predicted defect positions.

[0012] Preferably, the defect severity level prediction network includes, connected in sequence: a feature processor, a convolutional layer, a fully connected layer and a SoftMax layer.

[0013] Preferably, the feature processor includes, connected in sequence: a global feature mapping and fusion module, a region of interest pooling module, and a defect size-sensitive multi-scale feature fusion module;

[0014] The global feature mapping and fusion module is used to map the image features F output by the backbone network into the fused features F';

[0015] The region of interest pooling module is used to map the fused features F' into the multi-scale features BF of each defect with a fixed size;

[0016] The defect size-sensitive multi-scale feature fusion module is used to fuse the multi-scale features BF of each defect together to obtain the final defect features BF'.

[0017] Preferably, the global feature mapping and fusion module includes: a mapping and channel number unification sub-module, a fusion sub-module, and is used for the following calculations:

[0018] F' i = cbb(cbf(F i )·Gate(cbf(F i )) + upsample(cbf(F i+1 )·Gate(cbf(F i+1 ))))

[0019] where cbb(X) = Relu(bn(f 3×3×c (X))) represents the fusion sub-module;

[0020] cbf(X) = Relu(bn(f 1×1×c (X))) represents the mapping and channel number unification sub-module;

[0021]

[0022] Gate(X) = Sigmoid(f 1×1×c (X))

[0023] F i is the image feature output by the backbone network, F' i is the fused feature, i = 1, 2, 3, 4;

[0024] Relu represents the activation operation, bn represents the batch normalization operation, f 3×3×c represents the convolution operation with a filter size of 3*3 and an output feature channel number of c; f 1×1×c represents the convolution operation with a filter size of 1*1 and an output feature channel number of c; upsample represents upsampling.

[0025] Preferably, the defect size-sensitive multi-scale feature fusion module includes: a feature fusion weight prediction module and a resampling module, and is used for the following calculations:

[0026]

[0027] wherein, Resample represents resampling;

[0028] The feature fusion weight prediction module obtains the fusion coefficient vector W by the following formula:

[0029] W = sigmoid(fc(flatten_concat(BF1, BF2, BF3, BF4)))

[0030] wherein, the fusion coefficient vector W includes the fusion coefficient ω i , i = 1, 2, 3, 4, fc represents the fully connected layer, and flatten_concat represents flattening the output multi-scale features into one-dimensional features and splicing them together.

[0031] Preferably, the backbone network includes a plurality of residual modules;

[0032] The defect detection network includes: RPN, a defect location prediction head, and a defect classification head.

[0033] In the second aspect of the present invention, a training method for the above model is provided, and the method includes:

[0034] Preparation step: Prepare a data set, and the data set includes: images, the location BBox of defects, and the severity level Label of defects;

[0035] The first - stage training steps: Train the backbone network and the defect detection network to obtain the trained backbone network and defect detection network;

[0036] The second - stage training steps: Freeze the backbone network and train the defect level prediction network to obtain the trained defect level prediction network.

[0037] Preferably, the second - stage training steps include:

[0038] Convert the severity level Label of the defects in the dataset to Soft Label;

[0039] Use the image features F output by the backbone network, the positions BBox of the defects in the dataset, and the Soft Label to train the defect level prediction network.

[0040] Preferably, the loss function MultiLoss used in the second - stage training steps is as follows:

[0041] MultiLoss = w1·SFL+w2·EMD+w3·MSE

[0042] Among them, w1, w2, and w3 are the weights of each loss function;

[0043] SFL is the soft - focus loss function, EMD is the average earth - mover's distance loss function, and MSE is the mean squared error loss;

[0044] The definition of SFL is as follows:

[0045]

[0046]

[0047] Among them, is the Soft Label, representing the probabilities that the sample severity levels are 0, 1, 2, …, C - 1;

[0048] represents the probabilities that the sample severity levels output by the defect level prediction network are 0, 1, 2, …, C - 1;

[0049] γ represents the exponent of the power operation;

[0050] k represents the category serial number;

[0051] m i represents the number of samples with severity level i;

[0052] α k represents the category weight.

[0053] The third aspect of the present invention provides a method for detecting surface defects of industrial products and predicting the severity level, which is implemented by using the above-mentioned model. The method includes:

[0054] Input the collected image into the backbone network;

[0055] Input the image features output by the backbone network into the defect detection network, and the defect detection network outputs the predicted defect positions;

[0056] Input the image features output by the backbone network and the predicted defect positions output by the defect detection network into the defect severity prediction network, and the defect severity prediction network outputs the severity level of each defect corresponding to the predicted defect positions.

[0057] Compared with the prior art, the beneficial effects of the present invention are:

[0058] First of all, using the present invention can not only detect the positions of defects, but also identify the severity levels of defects. Experiments show that the present invention is superior to the traditional defect detection and classification network based on multiple branches;

[0059] Secondly, aiming at the problem of difficult identification of the severity level of defects caused by the huge difference in the size of surface defects of industrial products and semantic ambiguity, the feature processor proposed by the present invention can not only obtain the features of larger-sized defects, but also effectively retain the features of small-sized defects. Experiments show that the features processed by this processor are more conducive to the defect rating task;

[0060] Finally, different from ordinary classification tasks, the defect rating task is an ordered classification task. For this reason, the present invention proposes a method for converting the labeled Label into a Soft Label, so that this task can use both the loss function of the classification task and the loss function of the regression task, and proposes Soft Focal Loss. Experiments show that it is superior to SoftCross Entropy Loss. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 Schematic diagram of the principle when training the model of the method of the present invention;

[0062] Figure 2 Schematic diagram of the principle when testing with the model of the method of the present invention;

[0063] Figure 3-1 Schematic diagram of the structure of the feature processor in the method of the present invention;

[0064] Figure 3-2 Schematic diagram of the structure of the mapping and channel number unification sub-module ConvBlockF in the feature processor of the method of the present invention;

[0065] Figure 3-3 Schematic diagram of the defect size-sensitive multi-scale feature fusion module in the feature processor in the method of the present invention;

[0066] Figure 4 Original image and annotation of the silk ingot in the first embodiment of the present invention;

[0067] Figure 5 Figure 4 Expansion view;

[0068] Figure 6-1 Length distribution curve in the first embodiment of the present invention;

[0069] Figure 6-2 Width distribution curve in the first embodiment of the present invention;

[0070] Figure 7 Red oil mark diagram in the second embodiment of the present invention;

[0071] Figure 8-1 Width distribution diagram in the second embodiment of the present invention;

[0072] Figure 8-2 Height distribution diagram in the second embodiment of the present invention. Detailed implementation manners

[0073] The present invention will be further described in detail below with reference to the accompanying drawings:

[0074] During industrial quality inspection, the standards for defects are not fixed. For example: sometimes only very serious defects (such as the defect severity level is 8) are considered unqualified products, while sometimes the existence of minor defects (such as the defect severity level is 3) is considered an unqualified product. However, the existing methods cannot output the severity level of defects. Therefore, a model that can output the severity level of defects needs to be proposed. Moreover, in the actual quality inspection scenario, the quality inspection standards may change, and the quality inspection standards are related to the severity level of defects. Therefore, it is necessary to screen products that meet the quality inspection standards according to the severity level of defects.

[0075] The characteristics of industrial product defects are that the boundaries are blurred and the sizes of defect areas vary greatly, which poses a great challenge to defect level prediction.

[0076] The present invention provides a two-stage industrial product surface defect detection and severity level prediction model, which can not only detect the location of defects, but also output the severity level of defects. Moreover, the present invention proposes a feature processor, which can not only effectively extract the features of larger defects, but also effectively retain the features of small defects, solving the problem of difficult defect level prediction caused by the huge difference in the sizes of industrial product defects.

[0077] In addition, by converting the labeled Label into Soft Label, the present invention blurs the distinction between the classification task and the regression task, improves the calculation method of Focal Loss, enables it to be applied to the defect rating task, and is superior to using the cross-entropy loss.

[0078] The industrial product surface defect detection and severity level prediction model provided by the present invention includes: a backbone network (Backbone), a defect detection network, and a defect level prediction network; the defect detection network and the defect level prediction network are respectively connected to the backbone network; the defect detection network is connected to the defect level prediction network;

[0079] The backbone network is used to extract image features;

[0080] The defect detection network is used to obtain the predicted defect positions based on the image features;

[0081] The defect level prediction network is used to obtain the severity level (Defect Rank Pred) of each defect based on the image features and the predicted defect positions.

[0082] Specifically, the backbone network uses an existing network, mainly including multiple convolutional blocks (i.e., residual modules), for extracting image features. For example, the backbone network can use the existing ResNet50 (refer to Deep residual learning for image recognition (CVPR2016)). The input of the backbone network is an image (w*h, where w represents the width of the original image and h represents the height of the original image), and the output is the image feature F.

[0083] The defect detection network uses an existing network, including RPN (Region Proposal Network), a defect position prediction head (BBox Head), and a defect classification head (Classification Head) (refer to Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks (NIPS2015)).

[0084] Such as Figure 1As shown, during training, the input to the defect detection network is the image features, BBox, and Label output by the backbone network. After being processed by RPN, the image features are respectively input into the defect location prediction head (BBox Head) and the defect classification head (Classification Head). The defect location prediction head (BBox Head) outputs the predicted defect location BBox Pred, and the defect classification head (Classification Head) outputs the probability Cls Pred that the predicted defect location is a defect. The defect detection network is trained based on the comparison between BBox and BBox Pred, and the comparison between Label and Cls Pred.

[0085] As Figure 2 shown, during prediction, only the image features output by the backbone network need to be input into the defect detection network.

[0086] The defect level prediction network includes, connected in sequence: a feature processor, a convolutional layer, a fully connected layer, and a SoftMax layer (i.e., Figure 1 , Figure 2 the CNN+FC+SoftMax in Figure 1 ). The input to the defect level prediction network is: the image features F1, F2, F3, F4 output by the backbone network, and the location BBox of the defect (the BBox marked in the dataset is used during model training, as Figure 2 shown. The predicted defect location BBoxPred output by the defect detection network is used during prediction, as Figure 1 shown), and the severity level Label of the defect (the Label in the dataset needs to be input during model training, as Figure 2 shown, used to calculate the loss according to the loss function. Label does not need to be input during testing, as

[0087] shown). Figure 3-1 The core of the defect level prediction network is the feature processor. As

[0088] shown, the feature processor includes, connected in sequence: a global feature mapping and fusion module, a region of interest pooling module (RoIPooling), and a multi-scale feature fusion module sensitive to defect size.

[0089] Specifically as follows:

[0090] More specifically, the features used by the defect level prediction network are the image features output by the backbone network. In the industrial defect detection and level prediction scenario, the sizes of defects vary greatly (multi-scale), such as Figure 6-1 and Figure 6-2 shown, and research shows that multi-layer feature fusion is beneficial to solving the multi-scale problem (for details, refer to "Chen Keqi et al. A Review of Deep Learning Research on Multi-scale Object Detection. Journal of Software, 2021, 32(4): 1201–1227"). The backbone network (ResNet50) includes four residual modules connected in sequence, and the image features output by each residual module are F1, F2, F3, F4 (i.e., F i , i = 1, 2, 3, 4), and their sizes and numbers of channels are w1*h1*c1, w2*h2*c2, w3*h3*c3, w4*h4*c4 respectively. Generally, the numbers of channels of these features are different, that is, w1 to w4 decrease in sequence, h1 to h4 decrease in sequence, and c1 to c4 increase in sequence. For the convenience of subsequent processing, it is necessary to unify the numbers of channels of these features. For the above reasons, the present invention proposes a global feature mapping and fusion module.

[0091] The structure of the global feature mapping and fusion module is as shown in Figure 3-2 , and it includes: a mapping and channel number unification sub-module (ConvBlockF) and a fusion sub-module (ConvBlockB). Among them, the input of the global feature mapping and fusion module is: the image feature F i (i = 1, 2, 3, 4) output by the backbone network. The mapping and channel number unification sub-module (ConvBlockF) is used for feature mapping and unifying the channel numbers of features, and the fusion sub-module (ConvBlockB) is used to obtain the fused feature F' i .

[0092] Among them, the functions of the global feature mapping and fusion module are specifically as follows:

[0093] Assuming the input is X, the output of ConvBlockF can be expressed as cbf(X) = Relu(bn(f 1×1×c (X))), where Relu() represents the activation operation, bn() represents the batch normalization operation, and f 1×1×c( ) represents a convolution operation with a filter size of 1*1 and an output feature channel number of c. The sizes of different input features are different. To fuse the features of the upper layer (F4 is the upper layer of F3, F3 is the upper layer of F2, and so on), the features of the upper layer need to be upsampled. When fusing the features of adjacent layers, the upper layer features contain more semantic information and smaller feature sizes, while the lower layer features contain more detailed information and larger feature sizes. The importance of information in different layers is different. Therefore, a gating unit is used to calculate the importance of the features of this layer based on the image features during fusion. The calculation process of the gating unit is as follows:

[0094]

[0095] Gate(X) = Sigmoid(f 1×1×c (X))

[0096] The gating unit is mainly used to enhance the features in the defect area and weaken the features in the non-defect area. The calculation of the gating unit includes a convolution operation f 1×1×c ( ). In this way, through the training process, the response of the gating unit in the defect area is stronger than that in the non-defect area.

[0097] The features of adjacent layers are accumulated element by element and then passed through ConvBlockB to finally obtain the fused feature F' i (F'1, F'2, F'3, F'4), Figure 3-2 The whole process from F to F’ can be expressed as:

[0098] F' i = cbb(cbf(F i )·Gate(cbf(F i )) + upsample(cbf(F i+1 )·Gate(cbf(F i+1 ))))

[0099] Among them, cbb(X) = Relu(bn(f 3×3×c (X))) represents ConvBlockB (similar to ConvBlockF, but the filter size in the convolution is 3*3). Upsample represents upsampling.

[0100] The RoIPooling module is an existing module. For details, please refer to Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks (NIPS2015). According to the input BBox, it separates the features of the defect corresponding to the BBox from the global features. The input of the RoIPooling module is the global feature map and the output F' (F'1, F'2, F'3, F'4) of the fusion module, as well as the location information BBox of the defect (the position and size of the defect on the image are represented by a rectangular box). RoIPooling maps the fused feature F' into multi-scale features BF (BF1, BF2, BF3, BF4) of a fixed size according to the information of each defect.

[0101] The defect size-sensitive multi-scale feature fusion module is used to fuse the multi-scale features BF corresponding to each defect according to the size of the defect, enhance the features of the non-background part in the defect area, and form the final defect feature BF'.

[0102] The structure of the defect size-sensitive multi-scale feature fusion module is as Figure 3-3 shown, and it includes a feature fusion weight prediction module and a resampling module (Resample). The input of the defect size-sensitive multi-scale feature fusion module is the multi-scale features BF (BF1, BF2, BF3, BF4) of each defect, and the output is the feature BF' for defect level prediction. This module is used to fuse the multi-scale features of the defect to form the final feature for defect level prediction.

[0103] The sizes of defects generally vary greatly, and the contributions of features of different scales to the prediction of defect levels of different sizes are different. Therefore, when fusing the multi-scale features of the defect, the feature fusion weight prediction module obtains the fusion coefficient ω of each scale feature according to the multi-scale features BF (BF1, BF2, BF3, BF4) of each defect i (where i = 1, 2, 3, 4), and all ω i constitute the fusion coefficient vector W.

[0104] Specifically, the feature fusion weight prediction module obtains the fusion coefficient vector using the following formula:

[0105] W = sigmoid(fc(flatten_concat(BF1, BF2, BF3, BF4)))

[0106] Among them, the fusion coefficient vector W includes the fusion coefficient ω i, where \(i = 1, 2, 3, 4\), \(fc\) represents the fully connected layer, and \(flatten\_concat\) is an existing function that flattens the output multi-scale features into one-dimensional features and concatenates them together.

[0107] When fusing the multi-scale features of defects, it is necessary to first scale the multi-scale features to the same scale through resampling (Resample) (generally aligning the size of the multi-scale features of each defect with the size of BF1), then multiply by the feature fusion weight, and finally add them together to obtain the final defect feature BF'. The calculation method is as follows:

[0108]

[0109] Resample represents resampling and stands for the Resample module.

[0110] The calculation of the fusion coefficient includes the fully connected layer, so that through the training process, the network outputs reasonable fusion coefficients, thereby achieving the purpose of enhancing the defect information features and weakening the background information features during fusion.

[0111] The above model is a deep learning-based model, so the model needs to be trained first.

[0112] As Figure 1 shown, during training, the two networks are trained independently. First, the defect detection network is trained, and then the defect level prediction network is trained. When training the defect level prediction network, the backbone network needs to be frozen. The model training method specifically includes:

[0113] Preparation step: Prepare the dataset. The dataset refers to a set containing image and annotation data, specifically including: images (\(w*h\), where \(w\) represents the width of the original image and \(h\) represents the height of the original image), the position of the defect BBox (\(w\) b *h b , \(w\) b represents the width of the defect, and \(h\) b represents the height of the defect), and the severity level label of the defect;

[0114] First-stage training step: Train the backbone network and the defect detection network to obtain the trained backbone network and defect detection network;

[0115] Second-stage training step: Freeze the backbone network ( "freezing" is an operation on the neural network, which is an existing technology and will not be elaborated here), and train the defect level prediction network to obtain the trained defect level prediction network.

[0116] After the first-stage training is completed, the model can already distinguish which regions in the image are defects. The main task of the second stage is to enable the model to distinguish the severity level of the defects when the model already has the prior knowledge of distinguishing whether there are defects. During the second-stage training, the prior knowledge is achieved by freezing the backbone network because the backbone network has been trained well in the first stage and it can extract the image features in the image well to distinguish whether the specified region contains defects.

[0117] Specifically, the first-stage training steps include:

[0118] First, the image features output by the backbone network ( Figure 1In it, w1, h1, and c1 respectively represent the feature sizes output by different stages of the backbone network (where w represents width, h represents height, and c represents the feature dimension. The features of an image are a three-dimensional matrix, and these numbers represent the dimensions of the matrix). They are input into the RPN network in the defect detection network. The RPN network locates the approximate position of the defect and outputs candidate boxes (Proposal) (both RPN and Proposal are concepts in Faster R-CNN. For details, please refer to Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks (NIPS2015)). Then, based on the candidate boxes and the image features output by the backbone network, the image features output by the backbone network within the candidate boxes are separated (for details, please refer to Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks (NIPS2015)). Since the sizes of the candidate boxes are generally different, which is not conducive to subsequent convolution operations, RoIPooling (Region of Interest Pooling module) is needed to map the features corresponding to each defect into features of a fixed size. Finally, the above features are respectively input into the defect location prediction head (BBoxHead) and the defect classification head (Classification Head). The defect location prediction head (BBox Head) and the defect classification head (Classification Head) respectively output the position (BBox Pred) of the final defect and the probability (Cls Pred, binary classification) that this position is a defect. For details, please refer to Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks (NIPS2015). In the first-stage training step, using the pre-labeled defect positions BBox in the dataset, the backbone network and the defect detection network are trained using existing neural network training methods, and finally the trained backbone network and defect detection network are obtained.

[0119] In the first-stage training step, the loss functions used are existing classification loss and bounding box regression loss functions. For details, please refer to: Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networ.

[0120] The second-stage training step includes:

[0121] Using the image features F (F1, F2, F3, F4) output by the backbone network, the location BBox of the defects in the dataset (using the standard BBox in the dataset during training and the BBox Pred output by the defect detection network during prediction), and the severity level Label of the defects, the defect level prediction network is trained using an optimization algorithm (for details, refer to Decoupled Weight Decay Regularization (ICLR2019)). The specific process is as follows:

[0122] First, using the feature processor, the features of each defect are extracted from the features F (F1, F2, F3, F4) output by the backbone network (i.e., the "Feature" in Figure 1 which is also BF'), and then after being processed by a convolutional layer, a fully connected layer, and a SoftMax layer (i.e., CNN+FC+SoftMax in Figure 1 ), the severity level (Defect Rank Pred) of each defect is output. This information is a multi-dimensional vector (if the Label in the dataset contains 10 categories, it is a 10-dimensional vector. It can be designed according to the actual situation. For example, if there are 8 levels of defects, the label contains 8 categories, and correspondingly, the severity level of the defect is an 8-dimensional vector). Each dimension of this vector represents the probability of the defect corresponding to the severity level of the defect.

[0123] In the training stage, the present invention converts the Label in the dataset into a Soft Label, and then calculates the loss based on the prediction result of the defect level prediction network and the Soft Label. This loss is used to measure the difference between the prediction result and the labeled Label. Finally, the parameters in the defect level prediction network are updated through an optimization algorithm (for example, the Adam optimizer can be used. For details, refer to Decoupled Weight Decay Regularization (ICLR2019)).

[0124] The method of converting the Label into a Soft Label is an existing method. Refer to Soft Labels for Ordinal Regression (CVPR2019). A brief introduction is as follows:

[0125] The true value of the sample is represented as Y, which is a scalar representing the severity of the defect (a score Score, a score representing the severity of the defect, such as 1, 2, 3, etc.). It can be an integer or a decimal. If not specified otherwise, this value is an integer. If the true value is a decimal, its integer true value is obtained by the rounding method. This value is generally 0, 1, 2,..., C-1.

[0126] The Soft label is a C-dimensional vector, expressed as: where

[0127]

[0128] φ(Y, r i ) = |Y - r i | n

[0129] where r i represents the possible severity level of the defect.

[0130] It can be seen from the above formula that:

[0131] That is, Y S is a probability distribution, representing the probability that the sample belongs to a certain severity level. The core idea of Soft label is to ensure that the value of this probability distribution is the largest at the true value, and the probabilities at other severity levels decrease sequentially with the distance from the true value. Without special instructions, n = 2 here.

[0132] The loss function used in the second-stage training step is Multi Loss, which consists of three parts: SFL (Soft Focal Loss), EMD (Earth Mover's Distance Loss), and MSE (Mean Squared Error Loss), that is:

[0133] MultiLoss = w1·SFL + w2·EMD + w3·MSE

[0134] where w1, w2, and w3 are the weights of each loss function, used to balance the differences in the order of magnitude of each loss function. Without special instructions, w1 = 1, w2 = 10, and w3 = 1.

[0135] If the output of the defect level prediction network is denoted as: (The output after the output of the feature processor is processed by the subsequent CNN + FC + SoftMax is ), representing the probabilities that the severity level of the sample is 0, 1, 2,..., C - 1 (the value of C is determined by the dataset. For example, C is 10, that is, 10 severity levels are labeled). Obviously Then the definition of SoftFocal Loss is as follows:

[0136]

[0137]

[0138] Among them, is the Soft Label, representing the probabilities of the sample severity levels being 0, 1, 2, …, C-1;

[0139] represents the probabilities of the sample severity levels being 0, 1, 2, …, C-1 output by the defect level prediction network;

[0140]

[0141] φ(Y, r i ) = |Y - r i | n

[0142] r i represents the possible defect severity level;

[0143] γ is a parameter, representing the exponent of the power operation, without special meaning. If not specified otherwise, γ takes 2.

[0144] k represents the class serial number;

[0145] m i represents the number of samples with the severity level i;

[0146] α k represents the class weight, which is set according to the number of samples m of each severity level in the dataset k ;

[0147] The design inspiration of Soft Focal Loss (SFL) comes from Focal Loss (FL, Focal Loss for Dense Object Detection (ICCV2017)). FL starts from the perspective of sample classification difficulty, making the loss focus on difficult-to-classify samples, and solving the problem of low classification accuracy for classes with few samples. Of course, difficult-to-classify samples are not limited to classes with few samples, that is, FL not only solves the problem of sample imbalance, but also helps to improve the performance of the model on difficult-to-classify samples, thereby improving the overall performance of the model. The definition of FL is as follows: FL = -α c (1 - p c ) γ log(p c ), where p c represents the probability that the sample is classified into class C, and α c is the weight of this class. The FL loss function can only be used for class-independent classification tasks (classes are independent), and cannot be applied to ordered classification tasks (classes are not independent and have an order, such as the defect level prediction task) and regression tasks.

[0148] The definition of the commonly used loss function Soft Cross Entropy Loss (SCE) is as follows: where y k represents the true value of the sample (i.e., the probability that the sample is considered to be in the k-th class), and p k represents the predicted value. This loss function can be applied to ordinal classification tasks and regression tasks. However, this loss function cannot solve the problems of sample imbalance and difficult classification of some samples.

[0149] The defect severity level prediction task is an ordinal classification task with sample imbalance and difficult classification of some samples. The sample imbalance in this task is reflected in that in the actual industrial production scenario, the number of very severe defect samples is always small. For these defect samples, the difference between the samples with extremely low and extremely high defect severity levels is relatively large, and the classification is relatively simple; for the samples with defect severity levels in the middle part, the difference between the defects is small, and the classification is difficult.

[0150] In summary, the advantage of the SFL proposed in the present invention is that it can be applied to scenarios with sample imbalance and difficult classification of some samples, and can also be applied to ordinal classification tasks. Experiments show that the SFL of the present invention can effectively improve the indicators of the defect severity level prediction task.

[0151] The EMD loss function is a commonly used loss function in regression tasks (for details, refer to Squared Earth Mover's Distance-based Loss for Training Deep Neural Networks, arXiv:1611.05916), which is used to measure the similarity between two distributions and is defined as the minimum value of the average distance that needs to be moved to transform one probability distribution into another probability distribution. Its definition is as follows:

[0152]

[0153]

[0154] Unless otherwise specified, n = 2 in the above formula.

[0155] The MSE loss function is another commonly used loss function in regression tasks, which is used to measure the distance between the predicted value and the true value in the regression task. Its definition is as follows:

[0156]

[0157]

[0158] After training the three networks by the above method, asFigure 2 As shown in Figure 2 , during testing, the collected images are input into the backbone network, and the image features output by the backbone network are input into the defect detection network. The defect detection network outputs the predicted defect positions. Then, the image features output by the backbone network and the predicted defect positions output by the defect detection network are input into the defect severity prediction network, and the defect severity prediction network outputs the severity level of each defect corresponding to the predicted defect position. The connection relationships between the networks during testing are as shown in Figure 2 . Figure 2 as shown.

[0159] The effectiveness of the method of the present invention was verified on two datasets respectively, as follows:

[0160] The evaluation metrics used in the following experiments are as follows:

[0161] The calculation method of the Accuracy metric is as follows.

[0162]

[0163] where N correct represents the number of correctly predicted samples, and N all represents the total number of samples.

[0164] When calculating the accuracy rate, the rules for top(1) and top(3) (representing two evaluation metrics used to compare the advantages and disadvantages of algorithms, which is a method to objectively evaluate the performance of algorithms) to determine whether the sample is correct are as follows (where k = 1 or k = 3 (k represents the value in the parentheses of the above top(1), top(3))):

[0165]

[0166] The above formula means selecting the k categories with the highest probabilities as the predicted values.

[0167] When calculating the accuracy rate, the rules for class(±1) (another evaluation metric) to determine whether the sample is correct are as follows (where k = 1).

[0168]

[0169] Y = k if Y < k

[0170] Y = C - 1 - k if Y > C - 1 - k

[0171] The above formula means that the predicted value is considered correct when it is near the true value (within a distance of k from the true value).

[0172] To ensure that the number of elements in the true value set of each category is the same ((2k + 1) (the evaluation index is defined by ourselves). In the actual application scenario, the defect levels at both ends are either very serious or very minor, and the difference between levels is 1 - 2, not very large. So it is not necessary to divide them very finely at both ends, thus (2k + 1) is chosen.), the conditions are relaxed at both ends. For example, when k = 1 and C = 10, if the true value is 0, predicted values of 0, 1, 2 all represent correct; if the true value is 9, predicted values of 7, 8, 9 all represent correct.

[0173] When calculating the accuracy rate, the rule for score(±1.00) (another evaluation index) to judge whether a sample is correct is as follows (where r = 1.00).

[0174]

[0175] The meaning of the above formula is that the predicted score within the vicinity of the true value (at a distance of r and within from the true value) is considered correct.

[0176] When calculating Regression - related indicators, the true value used is the labeled data, that is, Y; the formula for calculating the predicted score is The calculation methods of MAE (Mean Absolute Error), MSE (Mean Squared Error), MAPE (Mean Absolute Percentage Error), PLCC (Pearson Correlation Coefficient), and SRCC (Spearman Rank Correlation Coefficient) indicators are all calculated according to the original definitions.

[0177] When calculating Classification - related indicators (processed as a single - label (single label) multi - class (multiclass) problem), the true value used is the labeled data, that is, Y; when calculating the evaluation indicators top(k) and class(±k), the predicted value is When calculating the evaluation indicator score(±r), the formula for the predicted value is

[0178] These evaluation indicators mentioned above evaluate the effect of the algorithm from different perspectives. For details, see the following experimental results.

[0179] Example 1:

[0180] Spindle dataset

[0181] During the production process of spindles, they may be contaminated with oil stains, which will affect the quality of the spindles. Therefore, it is necessary to determine the severity level of the oil - stain defects. The original images and annotations are asFigure 4 As shown Figure 4 in the figure, the colored circles and numbers are annotations). In this dataset, the areas other than defects are labeled with polygons, and the severity levels of the defect areas are labeled, with the severity levels being 1, 2, …, 10 respectively, where 1 represents a minor defect and 10 represents the most severe defect.

[0182] The dataset preprocessing method is as follows:

[0183] As Figure 4 shown, the oil stain defects are circularly distributed. If the image determined directly using the Bbox of the defects is used as the input, too many interference factors (including other defects, defect-free areas, and backgrounds) will be introduced.

[0184] Therefore, according to the characteristics of the silk ingot, first, two circles in the image are detected, and then the circular area formed by the two circles is unfolded into a rectangle. When unfolding, try to avoid destroying the original annotated polygons. If it is necessary to destroy them, the defect level annotations of the polygon areas divided into multiple parts remain unchanged and consistent. For the details of the circle detection method, please refer to the paper (C. Lu, S. Xia, M. Shao and Y. Fu, "Arc-Support Line Segments Revisited: An Efficient High-Quality Ellipse Detection," in IEEE Transactions on Image Processing, vol. 29, pp. 768 - 781, 2020, doi: 10.1109 / TIP.2019.2934352.). Figure 4 The unfolded diagram of Figure 5 is as shown

[0185] On this dataset, using the above circle detection method, the correct rate of detecting the circular rings in the image is over 90%.

[0186] After the image is unfolded, the range of the short side of the defect bounding box (BBox) is [9, 633] (the shortest short side of all defect bounding boxes (which is a rectangle) is 9, and the longest is 633.). As Figure 6-1 shown, the horizontal axis represents the side length of the short side of the defect BBox, and the vertical axis represents the proportion (cumulative proportion) of the number of BBoxes less than or equal to this side length in the total number of defects; the range of the long side is [11, 7723], as Figure 6-2 shown, the horizontal axis represents the side length of the long side of the defect BBox, and the vertical axis represents the cumulative proportion. From the above statistical information, it can be seen that the size differences of different BBoxes in this dataset are very large, and it is difficult to extract the features of BBoxes with different sizes. Therefore, the present invention proposes a feature processor module to effectively extract the features of BBoxes with extremely large size differences.

[0187] Statistics show that not only the size ranges of these BBox regions are very wide, but also the size differences of the expanded images are very large, with heights in the range of [256, 688] and widths in the range of [4832, 8740]. Since the original images are too large and have significant size differences, they cannot be directly used for model training. Therefore, the present invention proposes a method for unifying the sizes of defective images.

[0188] Assume that the image size used during training and testing is fixed at w*h. For the labeled polygons (BBoxes) smaller than this size, randomly generate t windows such that these windows can completely enclose these BBoxes; for the labeled polygons larger than this size, fix the window size at w*h and slide this window with a step value s to obtain a series of windows. Calculate the number of pixels count that these windows overlap with the labeled polygon, and sort these windows from smallest to largest based on count, then select the top t windows. The images corresponding to these windows are used as the final images for training and testing. If a window contains other defective regions (BBoxes), and more than 60% of the region of this defect is within this window, then this BBox is retained; otherwise, it is discarded. Unless otherwise specified, t = 10 and s is dynamically selected (ensuring that there are at least 5t steps on the long side). During training, for each defective BBox, always randomly select one from these t images in each epoch; during testing, only use the first of these t images.

[0189] From Figure 5 the statistical information, it can be seen that when the length is around 1000, it can cover about 80% of the data, and when the width is around 250, it can cover about 80% of the data. Therefore, w*h is set to 1024*256.

[0190] 80% of the dataset is used for training and 20% for testing. The model is trained for a total of 200 epochs.

[0191] Using the model and method proposed by the present invention, the test results obtained on this dataset are shown in Table 1.

[0192]

[0193] Table 1

[0194] Accuracy / class(±1) * : The conditions are relaxed at both ends. For example, if the true value is 0, predicted values of 0, 1, or 2 all represent correct. The same applies hereinafter and will not be separately marked.

[0195] Tables 1 and 2 respectively show the test results obtained by training models using different Losses. Table 1 evaluates the performance of the model from the perspective of the classification task accuracy (Accuracy, unit: %), and the main indicators shown include top(1), top(3), class(±1), and score(±1.00). The larger the values of these indicators, the better. top(1) mainly reflects the exact accuracy of the predicted value, and top(3), class(±1), and score(±1.00) reflect the approximate accuracy. Since the severity level annotation has a certain degree of subjectivity, and the number of annotations for each sample in this dataset is only one, the reference significance of the exact accuracy is not great. The approximate accuracy mainly evaluates the accuracy when the distance between the predicted value and the true value is 1. In the real scenario, the prediction of the defect severity level does not need to be too precise, and being able to give a general severity level range can meet the requirements. As can be seen from Table 1, the exact accuracy top(1) is about 40%, while the approximate accuracies top(3) and class(±1) have exceeded 80%, which can meet the requirements of practical applications. The score(±1.00) indicator is close to 70%, slightly lower than the class(±1) indicator, which is because the calculation rule of the score(±1.00) indicator is more stringent than that of class(±1). As shown in Table 1, the indicators of the model using SFL are all higher than those using SCE, indicating that the SFL proposed in the present invention performs better than SCE in this task, proving the superiority of the present invention.

[0196]

[0197] Table 2

[0198] Table 2 evaluates the performance of the model from the perspective of the regression task, and the indicators shown include MAE, MSE, MAPE, PLCC, and SRCC. MAE, MSE, and MAPE reflect the error between the predicted value and the true value, and the lower the value, the better. PLCC and SRCC reflect the correlation between the predicted value and the true value, and the higher the value, the better. MAE represents the mean absolute error. The value less than 1 indicates that the mean absolute error between the predicted value and the true value of the model proposed in the present invention is within 1, that is, the predicted value of the model is near the true value, which can be mutually confirmed with Table 1 and basically meet the application requirements in the real scenario. As can be seen from Table 2, the indicators using SFL are basically better than those using SCE, verifying the superiority of the SFL proposed in the present invention.

[0199] To verify the superiority of the present invention over the existing method (Faster R-CNN), on this dataset, Faster R-CNN and the method of the present invention are respectively used (such as Figure 1)Training and testing were carried out, and the test results are shown in Table 3. If the defect levels marked in the dataset are regarded as the category information of the defects, and this task is regarded as a defect detection and classification task, Faster R-CNN can output the location and category information of the defects. As can be seen from Table 3, the indicators of the method of the present invention are significantly better than those of the existing methods, verifying that the method of the present invention is superior to the traditional multi-branch defect detection and classification network (Faster R-CNN).

[0200]

[0201] Table 3

[0202] Another important innovation of the present invention is to propose a feature processor, which solves the problem of difficult defect level prediction in the scenario where the defect sizes vary greatly. To verify the superiority of the present invention, the feature processor was removed (i.e., "not used" in Table 4), and all the features output by the backbone network were directly fed into the RoIPooling module after being unified in size by the linear interpolation algorithm. The results are shown in Table 4. As can be seen from Table 4, the feature processor proposed by the present invention can effectively improve the indicators of the defect rating task, reflecting the superiority of the present invention.

[0203]

[0204]

[0205] Table 4

[0206] Example 2:

[0207] Red oil dataset

[0208] When assembling automotive screws, to confirm that each screw has been tightened, workers are required to dip the wrench in red oil and then tighten the screw, so that a mark of red oil will be left on the tightened screw, and thus it can be judged whether the screw has been tightened by whether there is red oil on the screw. However, in actual operation, workers may dip the wrench in red oil once and then tighten multiple screws, resulting in uneven depths (amounts) of the red oil marks left on the tightened screws, as Figure 7 shown. This poses a certain challenge to automated quality inspection. The total number of images contained in the dataset is 12163.

[0209] In this dataset, the range of the BBox width is [84, 253], and the range of the height is [61, 257]. As Figure 8-1 shown, the horizontal axis represents the length of the marked BBox width, and the vertical axis represents the proportion (cumulative proportion) of the number of BBoxes less than or equal to this length to the total number of marked BBoxes. As Figure 6-2As shown, the horizontal axis represents the length of the labeled BBox height, and the vertical axis represents the cumulative percentage. According to statistics, the aspect ratio of the image width to height is basically 1.

[0210] This dataset is relatively simple. Each image contains a screw, and the main body (screw cap) in the image is clear with distinct boundaries. From Figure 8-1 and Figure 8-2 statistics, it can be seen that there are differences in the BBox sizes, but compared with the spindle dataset, this difference is already relatively small. And for each sample in this dataset, at least 10 independent manual annotations were carried out, and the final label of each sample is the average of all annotations of this sample (if it is a decimal, it is rounded to an integer). Therefore, the labels of this dataset are more reliable compared with the spindle dataset.

[0211] The preprocessing method of this dataset is similar to that of the spindle dataset (excluding the part of the ring unfolding). As Figure 8-1 and 8-2 shown, when the width and height are 160, it can basically cover more than 90% of the data. Therefore, w*h is set to 160*160.

[0212] Each image in this dataset only contains one screw main body, and there is only a label for the annotation of this main body. Therefore, the existing classification network model, such as ResNet50 (Deep Residual Learning for Image Recognition, CVPR2016), can also be directly used to predict the depth of the red oil mark left on the screw. Due to the differences in image sizes, when using ResNet50, the image size needs to be fixed. Therefore, in the experiment, the image size is directly scaled to 224*224, and the loss function uses the commonly used Cross Entropy Loss in classification tasks. When calculating the Regression-related metrics, the predicted category is regarded as the regression value, so as to calculate the Regression-related metrics. The test results are shown in Table 5.

[0213]

[0214] Table 5

[0215] As shown in Table 5, the accuracy of the method of the present invention on this dataset is close to 90%, the approximate accuracy is close to 100%, and the mean absolute error (MAE) is much less than 1. The performance of the method of the present invention on this dataset is better than that on the ingot dataset. The main reasons are as follows: on the one hand, this dataset is relatively simple; on the other hand, the labeled Labels are also more reliable (the average value is taken from independent labeling by multiple people). This experiment also provides guiding information for the application of the method of the present invention in actual projects, that is, the Label should be reliable enough. However, in actual application scenarios, there may not be so much labeled data, resulting in the Label not being very reliable. The experiment on the ingot dataset also verifies that the method of the present invention can basically meet the application requirements in this case.

[0216] As can be seen from Table 5, compared with the simple classification network ResNet50, the performance indicators of the method of the present invention are significantly improved. On the one hand, it is attributed to the feature processor proposed by the present invention, which can effectively process features of different sizes and ensure that the features of small-size defects can be retained. On the other hand, the SFL loss function proposed by the present invention can better handle the ordered classification task compared with the Cross Entropy Loss. This experiment further illustrates the superiority of the method of the present invention.

[0217] From the above analysis, it can be seen that the method of the present invention can accurately predict the grade of defects.

[0218] In summary, the present invention can obtain the grade of defects while detecting the location of industrial defects, thereby providing a basis for defect screening during industrial on-site production. Experiments show that compared with the grade data manually labeled, the approximate correct rate of the grade data predicted by the method of the present invention is above 80% (according to the top(3) and class(±1) indicators), and the average error rate (MAE) is less than 1 (according to the MAE indicator). It can effectively predict the severity level of defects and has basically reached the standard of manual completion of this task.

[0219] Existing defect detection methods do not have an effective defect screening method and can only let the model learn the rules of defects through labeled data. However, the standards for defects during industrial quality inspection are not fixed. Therefore, when the standards change, it is necessary to relabel the data and retrain the model. When labeling, the method provided by the present invention directly labels the severity level of defects, and the trained model can directly output the severity level of defects. Therefore, it can effectively address the problem of changes in defect standards in the industrial quality inspection field, that is, different defect detection standards correspond to different severity levels of defects, and defects can be screened by the output severity level, providing a parameter for industrial on-site quality inspectors to control the defect standards.

[0220] The above technical solution is only one implementation mode of the present invention. For those skilled in the art, based on the disclosed principle of the present invention, it is very easy to make various types of improvements or deformations, not limited to the technical solution described in the above specific embodiments of the present invention. Therefore, the foregoing description is only preferred and does not have a restrictive meaning.

Claims

1. A model for detecting and predicting the severity of industrial product surface defects, characterized by: The model includes: a backbone network, and a defect detection network and a defect level prediction network respectively connected thereto; the defect detection network is connected to the defect level prediction network; The backbone network is used to extract image features; The defect detection network is used to obtain a predicted defect location based on image features; The defect level prediction network is used to obtain the severity level of each defect based on image features and predicted defect locations; The defect level prediction network includes: a feature processor, a convolutional layer, a fully connected layer and a SoftMax layer connected in sequence; The feature processor includes: a global feature mapping and fusion module, an area of ​​interest pooling module, and a defect size-sensitive multi-scale feature fusion module connected in sequence; The global feature mapping and fusion module is used to map the image feature F output by the backbone network into the fused feature F'; The region of interest pooling module is used to map the fused feature F' into a multi-scale feature BF of each defect of a fixed size; The defect size-sensitive multi-scale feature fusion module is used to fuse the multi-scale features BF of each defect together to obtain the final defect feature BF'; The global feature mapping and fusion module includes: a mapping and channel number unification submodule and a fusion submodule, which are used to perform the following calculations: F' i =cbb(cbf(F i )·Gate(cbf(F i ))+upsample(cbf(F i+1 )·Gate(cbf(F i+1 )))) Where cbb(X)=Relu(bn(f 3×3×c (X))) represents the fusion submodule; cbf(X)=Relu(bn(f 1×1×c (X))) represents the mapping and channel number unified submodule; Gate(X)=Sigmoid(f 1×1×c (X)) F i is the image feature output by the backbone network, F' i is the fused feature, i=1, 2, 3, 4; Relu represents the activation operation, bn represents the batch normalization operation, and f 3×3×c represents a convolution operation with a filter size of 3*3 and an output feature channel number of c; f 1×1×c It indicates a convolution operation with a filter size of 1*1 and an output feature channel number of c; upsample indicates upsampling; The defect size-sensitive multi-scale feature fusion module includes: a feature fusion weight prediction module and a resampling module, which are used to perform the following calculations: Among them, Resample means resampling; The feature fusion weight prediction module uses the following formula to obtain the fusion coefficient vector W: W=sigmoid(fc(flatten_concat(BF1,BF2,BF3,BF4))) Among them, the fusion coefficient vector W includes the fusion coefficient ω i , i = 1, 2, 3, 4, fc represents the fully connected layer, and flatten_concat means flattening the output multi-scale features into one-dimensional features and concatenating them together.

2. The industrial product surface defect detection and severity level prediction model according to claim 1, characterized in that: The backbone network includes a plurality of residual modules; The defect detection network includes: RPN, defect location prediction head and defect classification head.

3. A model training method, the method being used to train the model according to any one of claims 1-2, characterized in that: The method comprises: Preparation step: prepare a data set, the data set includes: an image, a defect location BBox, and a defect severity level Label; The first stage training steps: training the backbone network and defect detection network to obtain the trained backbone network and defect detection network; The second stage training steps: freeze the backbone network, train the defect level prediction network, and obtain the trained defect level prediction network.

4. The method according to claim 3, characterized in that: The second stage training steps include: Convert the severity level Label of the defects in the data set to Soft Label; The defect level prediction network is trained using the image features F output by the backbone network and the defect location BBox and SoftLabel in the data set.

5. The method according to claim 4, characterized in that: The loss function MultiLoss used in the second stage training step is as follows: MultiLoss=w1·SFL+w2·EMD+w3·MSE Among them, w1 w2 w3 are the weights of each loss function; SFL is the soft focus loss function, EMD is the mean bulldozing distance loss function, and MSE is the mean square error loss; The definition of SFL is as follows: k∈[0,1,…,C-1] in, is a soft label, indicating the probability that the severity level of the sample is 0, 1, 2, …, C-1; It represents the probability that the sample severity level output by the defect level prediction network is 0, 1, 2, ..., C-1; γ represents the exponent of the power operation; k represents the category number; m i represents the number of samples with severity level i; α k Represents the category weight.

6. A method for detecting and predicting the severity of industrial product surface defects, characterized in that: The method is implemented using the model according to any one of claims 1 to 2, and the method comprises: Input the collected images into the backbone network; The image features output by the backbone network are input into the defect detection network, and the defect detection network outputs the predicted defect location; The image features output by the backbone network and the predicted defect positions output by the defect detection network are input into the defect level prediction network, and the defect level prediction network outputs the severity level of each defect corresponding to the predicted defect position.

Citation Information

Patent Citations

  • Method for identifying casting DR image loose defects based on improved YOLOv3 network model

    CN111476756A

  • Aluminum material image defect detection method based on self-adaptive anchor frame

    CN112085735A