Intelligent defect identification model and method

Through comparative learning and unsupervised training of the ResNet-50 model, combined with the teacher-student network architecture, the problem of label-free data utilization in industrial defect detection is solved, and efficient defect detection effect is achieved, which is suitable for industrial environments.

CN120339155APending Publication Date: 2025-07-18NANJING STAR ENERGY TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410059615.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-16
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The prior art is difficult to effectively utilize a large amount of label-free data in industrial defect detection, and convolutional neural networks are prone to performance degradation during training, resulting in poor detection results.

Method used

The unsupervised defect detection algorithm of contrast learning is used to pre-train a large number of defect-free samples, combined with the ResNet-50 model and teacher-student network architecture, through contrast learning and image reconstruction technology, the internal features of the data are automatically learned, and defect areas are located through image processing algorithms during the detection stage.

Benefits of technology

It achieves high accuracy and high recall in defect detection, with detection accuracy and recall values reaching 89.72% and 88.64% respectively, which is suitable for lightweight detection in industrial environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339155A_ABST
    Figure CN120339155A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent active learning, in particular to an intelligent defect identification model and method. The invention provides a brand-new unsupervised defect detection algorithm using comparative learning. The algorithm is realized only by using a large amount of defect-free sample data which is easy to obtain; in the training stage, only defect-free samples are used for unsupervised training, a comparative learning method is adopted for data enhancement, and meanwhile, the design of a teacher-student network is beneficial for the model to automatically learn potential features in the data; in a detection stage, an input defect sample is reconstructed by using a pre-trained model, and a defect area can be accurately detected through a traditional image processing algorithm. By implementing the algorithm disclosed by the invention, the problem that defect samples are difficult to obtain in an industrial environment can be solved, and the accuracy rate and the recall value of the algorithm respectively reach 89.72% and 88.64%.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent active learning, and particularly relates to a defect intelligent recognition model and method. Background Art

[0002] Traditional machine learning methods can effectively solve quality inspection problems of various industrial products, such as bearings, mobile phone screens, coils, steel rails, steel beams, etc.; such methods use artificial feature extractors to adapt to specific product image sample datasets, use the features of classifiers and support vector machines, and neural networks to determine whether a product has defects; however, using the method based on a fuzzy neural network to handle problems such as object detection has the disadvantage that it can only handle small dataset problems; using a convolutional neural network for multi-type object detection tasks can identify multiple types of datasets at one time, but the convolutional neural network is prone to a decline in model performance during training; recently, deep learning methods based on convolutional neural network (CNN) have become the mainstream methods for surface defect detection, having better feature representation capabilities than artificially designed feature extraction algorithms; considering that it is difficult to obtain a large number of labeled datasets in actual industrial scenarios and requires a large amount of human and financial costs, therefore, in the future, aiming at the deficiencies of the above methods, a technical solution of pre-training + fine-tuning is adopted. In the pre-training part, the idea of unsupervised contrastive learning is adopted to pre-train a large number of unlabeled datasets, and then in the fine-tuning part, the ResNet model is used as the basic model for defect intelligent recognition, which is a good solution. Summary of the Invention

[0003] The object of the present invention: To solve the deficiencies of existing defect recognition methods, the present invention proposes a brand-new unsupervised defect detection algorithm using contrastive learning, which is implemented only by using a large number of easily obtained defect-free sample data; in the training stage, only defect-free samples are used for unsupervised training, and the method of contrastive learning is adopted for data augmentation. At the same time, the design of the teacher-student network is beneficial for the model to automatically learn the potential features inside the data; in the detection stage, the input defect samples are reconstructed by the pre-trained model, and the defect area can be accurately detected through traditional image processing algorithms.

[0004] Technical solution: To solve the above problems, the present invention provides a defect intelligent recognition model and method, which is characterized in that the method includes two stages: an image pre-training stage and a surface defect area detection stage, wherein:

[0005] Image pre-training stage: It consists of a pre-training module and a downstream task module. Contrastive learning is used for pre-training. A large amount of unlabeled data sets are used, and a teacher-student network architecture is adopted, where the basic model is ResNet-50. The workflow is as follows: Input the actual industrial negative film image into the pre-training part to obtain high-level semantic representations, and then input them into the downstream neural network for training. First, forward propagation is performed, and then backpropagation. After the model converges, the parameter model is saved.

[0006] Surface defect area detection stage: The residual between the reconstructed image and the image to be measured is used as the area where defects may exist, and the final detection result is obtained through conventional image operations.

[0007] In the pre-training module, specifically, first, a sample picture collected on-site is input to generate different view sets V, which includes local views and global views. All views are input into the student network, and only the global view is input into the teacher network, so that our network can learn the context relationship. If a given image input x is provided, the two modules of the student network and the teacher network output k-dimensional probability distribution values Ps and Pt, where the probabilities are obtained by Softmax. The optimization function of the pre-training algorithm model part can be expressed as:

[0008] When updating the gradient, only the student branch is updated, and backpropagation is used for updating. The teacher branch stops gradient update, and the parameter update of the teacher network is obtained from the moving average of the student network. The formula is: θ t ←λθ t +(1 - λ)θ s

[0009] In the downstream task module, there are three requirements for network design reconstruction: The network should be able to adapt to defect areas of different scales, the network needs to identify whether there are defect features in the sample area, and the network model should be reconstructed with as few parameters as possible.

[0010] The reconstruction process should be decomposed into an encoding transformation and a decoding transformation γ, which are defined as follows: I→F γ:F→I In the above formula, I∈R W×H represents the spatial domain of the image sample, which is mapped to the hidden space through the function F represents the image sample features corresponding in the hidden space, which is implemented by the encoding module.

[0011] γ remaps the image sample features F corresponding to the hidden layer space back to the spatial domain of the original image samples, which is implemented by the decoding module, where The encoding and decoding processes are described as: In the above equation, I′ represents the reconstructed image, represents convolution, σ represents the activation function; W and W’ represent the encoding convolution kernel and the decoding convolution kernel respectively; b and b’ represent the encoding bias and the decoding bias.

[0012] The original image is divided into several image patches, usually 16×16, 32×32, and 64×64, as the input to the network; the downstream training model uses three convolution kernels of 1×1, 3×3, and 5×5 to obtain multi-scale features and inputs the multi-scale features into the encoding module.

[0013] The result output by the decoding module is fed into three deconvolution layers of different scales to obtain the final reconstructed image. At the same time, residual connection operations are taken for the input and output to prevent gradient explosion during model training. After obtaining the feature map, it is projected to the same dimension as the input image, that is, the output image O∈R W×H ; finally, the loss is calculated with the true label to achieve model convergence.

[0014] In the training stage of the downstream model, the reconstruction error between the original image and the reconstructed image is used as the loss function to promote network convergence; the present invention provides a loss function combining L1 loss and L SSIM as the loss function of the downstream network model, as follows: L ResNet-DE = αL1+(1-α)L SSIM In the above formula, α is the weight coefficient, and its value range is (0, 1), which is used to balance the ratio of L1 loss and L SSIM ; the present invention will experimentally compare the effects of different weight coefficients and loss functions on the detection results of the downstream model; where:

[0015] The L1 loss is the mean absolute error loss: L1 = ||I src - I rec ||1 + λ||ω|| F In the above formula, I src represents the input original image, I rec represents the image reconstructed by the model, ω represents the set of weights in the reconstruction network, λ represents the penalty factor of the regularization term, 0 < λ < 1;

[0016] LSSIM is the loss function; for the image pair (x, y) of the model input and output, SSIM can be defined as: SSIM(x, y) = (1(x, y))α(c(x, y)) β (s(x, y)) γ where α > 0, β > 0, γ > 0, 1(x, y) represents the luminance ratio, c(x, y) represents the contrast comparison, s(x, y) represents the structure comparison; Ux and Uy are the average values of x and y respectively, σx and σy are the standard deviations of x and y, and σxy is the covariance of x and y; C1, C2, C3 are non-zero constants, usually α = β = γ = 1, C3 = C2 / 2; then the loss function can be defined as follows: L SSIM (x, y) = 1 - SSIM(x, y)

[0017] In the surface defect area detection stage, after training the defect image input to reconstruct the network, the network will output an approximate defect-free image; that is, the reconstructed network will "repair" the defect area into a normal area while maintaining the defect-free area; according to this feature, the pixel-level difference between the output image and the input image, through conventional image processing techniques, can accurately locate the defect area, and the specific processing process is as follows:

[0018] Obtaining the residual map: The input image and the downstream model reconstructed image are used to make a difference image to obtain the reconstruction error of the network for the defect area, and the residual map is obtained, which contains the location information of the abnormal area; among them, the residual map V is as follows: v(i, j) = (I src (i, j) - I rec (i, j)) 2

[0019] Denoising processing: The residual map shows a large amount of noise, forming pseudo-defects, which affects the judgment of the real defect area. Mean filtering can be used for denoising processing.

[0020] Threshold segmentation and defect location: Further processing operations are performed on the basis of the above-mentioned defect image processing;

[0021] Finally, the adaptive threshold method is used to obtain the most ideal result, that is, the defect location.

[0022] Implementing the algorithm of this invention patent can solve the problem of difficult to obtain defect samples in the industrial environment, and its precision and recall rate reach 89.72% and 88.64% respectively. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 It is a schematic diagram of the unsupervised contrastive learning model process of an embodiment of the present invention;

[0024] Figure 2 It is a schematic diagram of the network structure model architecture of an embodiment of the present invention;

[0025] Figure 3 It is a position diagram of the residual graph processing flow in an embodiment of the present invention. DETAILED DESCRIPTION Example 1

[0026] The specific implementation modes of the present invention are described in detail below in conjunction with the accompanying drawings.

[0027] The present invention provides a defect intelligent recognition model and method, characterized in that the method includes two stages: an image pre-training stage and a surface defect area detection stage, wherein:

[0028] Image pre-training stage: It consists of a pre-training module and a downstream task module. It uses a contrastive learning method for pre-training, uses a large number of unlabeled data sets, and adopts a teacher-student network architecture. The basic model is ResNet-50. The workflow is as follows: input the actual industrial film image into the pre-training part to obtain high-level semantic representation, and then input it into the downstream neural network for training, forward propagation, and then back propagation, and save the parameter model after the model converges;

[0029] Surface defect area detection stage: The residual between the reconstructed image and the image to be tested is used as the area where defects may exist, and the final detection result is obtained through conventional image operations.

[0030] In the pre-training module (see Figure 1 Specifically, we first input a sample image collected on site to generate different perspective sets V, which include local views and global views. All views are input into the student network, and only the global view is input into the teacher network, so that our network can learn the contextual relationship. If the image input x is given, the two modules of the student network and the teacher network output k-dimensional probability distribution values Ps and Pt, where the probabilities are obtained by Softmax. The optimization function of the pre-training algorithm model can be expressed as:

[0031] When the gradient is updated, only the student branch is updated, and back propagation is used for update. The teacher branch stops updating the gradient. The parameter update of the teacher network is derived from the moving average of the student network. The formula is: θ t ←λθ t+(1 - λ)θ s

[0032] In the downstream task module, due to the multi-scale characteristics of industrial product surface defects, similarity to background textures, and complex shapes, high requirements are imposed on the accuracy and operation time of the detection algorithm. Therefore, three requirements are put forward for the network design reconstruction: 1) The network should be able to adapt to defect regions of different scales; 2) The network needs to identify whether there are defect features in the sample area; 3) Reconstruct the network model with as few parameters as possible.

[0033] See Figure 2 As shown, the reconstruction process is usually decomposed into an encoding transformation and a decoding transformation γ, defined as follows: I → F γ: F → I In the above formula, I ∈ R W×H represents the spatial domain of the image sample, which is mapped to the hidden space through the function and F represents the image sample features corresponding in the hidden space, which is implemented by the encoding module.

[0034] γ remaps the image sample features F corresponding to the hidden layer space back to the spatial domain of the original image sample, which is implemented by the decoding module, where the encoding and decoding processes are described as: In the above equation, I' represents the reconstructed image, represents convolution, σ represents the activation function; W and W' represent the encoding convolution kernel and the decoding convolution kernel respectively; b and b' represent the encoding bias and the decoding bias.

[0035] To adapt to larger images, the original image is divided into several image blocks, usually 16×16, 32×32, and 64×64, as the input to the network; the downstream training model uses three convolution kernels of 1×1, 3×3, and 5×5 to obtain multi-scale features and inputs the multi-scale features into the encoding module.

[0036] The result output by the decoding module is fed into three deconvolution layers of different scales to obtain the final reconstructed image. At the same time, residual connection operations are taken at the input and output to prevent gradient explosion during model training. After obtaining the feature map, it is projected to the same dimension as the input image, that is, the output image O ∈ R W×H ; finally, the loss is calculated with the real label to achieve model convergence.

[0037] During the training phase of the downstream model, the reconstruction error between the original image and the reconstructed image is used as the loss function to promote network convergence; the present invention provides a loss function combining L1 loss and L SSIM as the loss function of the downstream network model, as follows: L ResNet-DE = αL1+(1-α)L SSIM In the above formula, α is the weight coefficient, and its value range is (0, 1), which is used to balance the proportion of L1 loss and L SSIM ; where:

[0038] The L1 loss is the mean absolute error loss: L1 = ||I src - I rec ||1+λ||ω|| F In the above formula, I src represents the input original image, I rec represents the image reconstructed by the model, ω represents the set of weights in the reconstruction network, λ represents the penalty factor of the regularization term, 0 < λ < 1; the reconstruction algorithm model with MSE as the loss function is applicable to image samples with regular texture backgrounds. Because, the texture backgrounds of most industrial products are irregular, abnormal features are easily incorporated into the texture background, and the difference between abnormal features and normal texture background features is very small.

[0039] L SSIM is the loss function; for the image pair (x, y) of the model input and output, SSIM can be defined as: SSIM(x, y) = (l(x, y))α(c(x, y)) β (s(x, y)) γ where, α > 0, β > 0, γ > 0, 1(x, y) represents the luminance ratio, c(x, y) represents the contrast comparison, s(x, y) represents the structure comparison; Ux and Uy are the averages of x and y respectively, σx and σy are the standard deviations of x and y, and σxy is the covariance of x and y; C1, C2, C3 are non-zero constants, usually α = β = γ = 1, C3 = C2 / 2; then the loss function can be defined as follows: L SSIM (x, y) = 1 - SSIM(x, y)

[0040] In the surface defect area detection stage (see Figure 3 as shown), after training the input of the defect image to reconstruct the network, the network will output an approximate defect-free image; that is, the reconstructed network will "repair" the defect area into a normal area while maintaining the defect-free area; according to this feature, the pixel-level difference between the output image and the input image can accurately locate the defect area through conventional image processing techniques. The specific processing process is as follows:

[0041] Obtaining the residual map: The input image (as shown in Figure 3 (a)) and the downstream model reconstructed image (as shown in Figure 3 (b)) are used to make a difference image to obtain the reconstruction error of the network for the defect area. The obtained residual map is as shown in Figure 3 (c), which contains the location information of the abnormal area; among them, Figure 3 (a) is the original image input to the model, Figure 3 (b) is the ResNet-DE reconstruction map, Figure 3 (d) is the filtering of the residual map, Figure 3 (e) is the defect location, Figure 3 (c) is the residual map V as follows: v(i, j) = (I src (i, j) - I rec (i, j)) 2

[0042] Denoising processing: Figure 3 The residual map of (c) shows a large amount of noise, forming pseudo-defects, which affects the judgment of the true defect area. Mean filtering is used for denoising processing to obtain Figure 3 (d);

[0043] Threshold segmentation and defect location: Further processing operations are performed on the basis of the above-mentioned defect image processing;

[0044] Finally, the adaptive threshold method is used to obtain the most ideal result, that is, the defect location Figure 3 (e).

[0045] The present invention provides an unsupervised defect detection algorithm using contrastive learning. This algorithm is implemented only by using a large number of easily obtained defect-free sample data, and it can solve the problem of difficult access to defect samples in the industrial environment. In the training stage, only defect-free samples are used for unsupervised training, and the method of contrastive learning is adopted for data augmentation. At the same time, the design of the teacher-student network is beneficial for the model to automatically learn the potential features inside the data. In the detection stage, the pre-trained model is used to reconstruct the input defect samples, and the traditional image processing algorithm can be used to accurately detect the defect area. The results show that the detection algorithm proposed by the present invention has achieved good results, and its precision rate and recall value have reached 89.72% and 88.64% respectively. At the same time, due to the lightness of the downstream network, it is suitable for transplantation into the industrial detection environment.

[0046] The present invention is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative rather than restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the purpose of the present invention and the scope protected by the claims, and these all belong to the protection scope of the present invention.

Claims

1. A defect intelligent recognition model and method, characterized in that, The method includes two stages: the image pre-training stage and the surface defect area detection stage, where: In the image pre-training stage, it consists of a pre-training module and a downstream task module. The pre-training is carried out by using the contrastive learning method, using a large amount of unlabeled data sets, and adopting the teacher-student network architecture, where the basic model adopts ResNet-50. The workflow is as follows: Input the actual industrial negative film image into the pre-training part to obtain high-level semantic representations, and then input them into the downstream neural network for training. First, forward propagation is carried out, and then backward propagation. After the model converges, the parameter model is saved; In the surface defect area detection stage, the residual between the reconstructed image and the image to be measured is used as the area where defects may exist, and the final detection result is obtained through conventional image operations.

2. The defect intelligent recognition model and method according to claim 1, wherein In the pre-training module, first, a sample picture collected on-site is input to generate a set of different perspectives V, which includes local views and global views. All views are input into the student network, and only the global view is input into the teacher network, so that the network can learn the context relationship. If a given image input x is given, the two modules of the student network and the teacher network output the probability distribution values Ps and Pt in k dimensions, where the probabilities are both obtained by Softmax. The optimization function of the pre-training algorithm model part can be expressed as: When updating the gradient, only the student branch is updated, and the update is carried out by backpropagation. The teacher branch stops the gradient update, and the parameter update of the teacher network is obtained from the moving average of the student network. The formula is: θ t ←λθ t +(1 - λ)θ s 。 3. The defect intelligent recognition model and method according to claim 1, characterized in that In the downstream task module, there are three requirements for network design reconstruction: the network should be able to adapt to defect regions of different scales, the network needs to identify whether there are defect features in the sample area, and reconstruct the network model with as few parameters as possible; and the reconstruction process should be decomposed into an encoding transformation and a decoding transformation γ, which are defined as follows: γ:F→I In the above formula, I ∈ R W×H represents the spatial domain of the image sample, and is mapped to the hidden space through the function F represents the corresponding image sample features in the hidden space. is implemented by the encoding module, and γ remaps the image sample features F corresponding to the hidden layer space back to the spatial domain of the original image sample, which is implemented by the decoding module, where The encoding and decoding processes are described as: In the above equation, I′ represents the reconstructed image, represents convolution, σ represents the activation function; W and W' represent the encoding convolution kernel and the decoding convolution kernel respectively; b and b' represent the encoding bias and the decoding bias; At the same time, the original image is divided into several image patches of 16×16, 32×32, and 64×64 as the input to the network; The downstream training model uses three convolutional kernels of 1×1, 3×3, and 5×5 to obtain multi-scale features, and the multi-scale features are input into the encoding module; The results output by the decoding module are fed into transposed convolutional layers of three different scales to obtain the final reconstructed image. At the same time, residual connection operations are performed on the input and output. After obtaining the feature maps, projection is carried out to the same dimension as the input image, that is, the output image O ∈ R W×H ; finally, loss calculation is performed with the true label to achieve model convergence.

4. The defect intelligent recognition model and method according to claim 1, characterized in that, During the training phase of the downstream model, the reconstruction error between the original image and the reconstructed image is used as a loss function to promote network convergence; the present invention provides a loss function combining L1 loss and L SSIM as the loss function of the downstream network model, as follows: L ResNet-DE = αL1 + (1 - α)L SSIM In the above formula, α is the weight coefficient, and its value range is (0, 1), which is used to balance the proportion of the L1 loss and L SSIM of Where: The L1 loss is the mean absolute error loss: L1 = ||I src -I rec ||1 + λ||ω|| F In the above formula, I src represents the original input image, and I rec represents the image reconstructed by the model. ω represents the set of weights in the reconstruction network, λ represents the penalty factor of the regularization term, and 0 < λ < 1; L SSIM is the loss function; for the image pair (x, y) of the model input and output, SSIM can be defined as: SSIM(x, y) = (l(x, y)) α (c(x, y)) β (s(x, y)) γ Where α>0, β>0, γ>0, l(x,y) represents the brightness ratio, c(x,y) represents the contrast comparison, s(x,y) represents the structure comparison; Ux and Uy are the averages of x and y respectively, σx and σy are the standard deviations of x and y, and σxy is the covariance of x and y; C1, C2, and C3 are non-zero constants. Usually, α = β = γ = 1, and C3 = C2 / 2; then the loss function can be defined as follows: L SSIM (x, y) = 1 - SSIM(x, y).

5. The defect intelligent recognition model and method according to claim 1, wherein In the surface defect area detection stage, after training the input of the defect image to reconstruct the network, the network will output an approximate defect-free image; that is to say, the reconstructed network will "repair" the defect area into a normal area while maintaining the defect-free area. According to this feature, the pixel-level difference between the output image and the input image, through conventional image processing techniques, can accurately locate the defect area. The specific processing process is as follows: Obtaining the residual map: The input image and the downstream model reconstructed image are used to make a difference image to obtain the reconstruction error of the network for the defect area to obtain the residual map, which contains the position information of the abnormal area; where the residual map V is as follows: v(i, j) = (I src (i, j) - I rec (i, j)) 2 Denoising processing: The residual map shows a large amount of noise, forming pseudo-defects, which affects the judgment of the true defect area. Mean filtering can be used for denoising processing; Threshold segmentation and defect localization: Further processing operations are performed on the basis of the aforementioned defect image processing; Finally, the adaptive threshold method is used to obtain the most ideal result, that is, the defect position.