Method, System and Device for Constructing a Few-Shot Industrial Image Defect Detection Model

The proposed U-Net-based image segmentation method with attention layers and dilation convolutions addresses the limitations of traditional and deep learning methods, achieving efficient and accurate industrial defect detection with minimal samples and reduced computational load.

CN114782391BActive Publication Date: 2025-07-15HANGZHOU YIDALONG TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210475852.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-29
Publication Date
2025-07-15
Estimated Expiration
2042-04-29

AI Technical Summary

Technical Problem

The existing industrial image defect detection methods are poorly robust in various detection scenarios, and deep learning-based methods require a large amount of data and computing resources, making it difficult to achieve real-time detection.

Method used

A small sample industrial image defect detection model based on image segmentation is adopted. Through the combination of segmentation network and decision network, a small number of defect samples are used for training, including preprocessing, random flip, attention layer and hollow convolution, and defect detection is carried out in combination with local features and global information.

Benefits of technology

It significantly improves the accuracy and speed of defect detection, can be used in multiple detection scenarios, and can achieve better classification performance under a small number of samples, reducing the amount of calculation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114782391B_ABST
    Figure CN114782391B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, system and device for constructing a few-shot industrial image defect detection model, including: dividing an industrial image data set into a training set and a test set; preprocessing the training set and the test set to enhance the image contrast; randomly vertically flipping and horizontally flipping the preprocessed training set to obtain training sample data; inputting the training sample data into a segmentation network for training to complete the first step of constructing the detection model; correcting the mask image to obtain the industrial image defect area to complete the second step of constructing the detection model; splicing the mask image and the industrial image defect area to obtain a two-channel feature map, inputting the two-channel feature map into a decision network for training to complete the third step of constructing the detection model; evaluating the detection model according to the defect classification results of the industrial images with different defect types in the test set. The present invention can improve the detection accuracy and speed of defect images and evaluate the constructed model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of constructing few-shot industrial image defect detection models, and in particular to a method, system and device for constructing few-shot industrial image defect detection models. Background Art

[0002] At present, the methods for industrial image defect detection can be divided into two categories: methods based on traditional machine vision and methods based on deep learning.

[0003] The methods of traditional machine vision mainly use algorithms such as threshold segmentation, morphological processing, wavelet transform, and edge detection to achieve defect detection. In deep learning, convolutional neural networks are mainly used to extract features from images. For example, defect detection is achieved by using classification networks, segmentation networks, object detection networks, etc. Commonly used models include: ResNet, YOLO, U-Net, etc.

[0004] The methods based on traditional machine vision require numerous preprocessing steps, and the preprocessing methods have strong pertinence. A large number of parameters need to be set manually, and they only have good performance in a single detection scenario. Therefore, the methods based on traditional machine vision have poor robustness, and the same detection method is difficult to be applied to multiple detection scenarios. The methods based on deep learning rely on the excellent feature extraction ability of convolutional neural networks, which can significantly improve the detection accuracy. However, this method requires a large amount of image data to train the network, and it also takes a lot of manpower and time to make labels for the data. In actual industrial production, a large number of defect-free images can usually be obtained, while the cost of obtaining defective images is relatively high. On the other hand, the industrial application field has extremely high requirements for the real-time performance and accuracy of algorithms. In order to achieve high accuracy, the mainstream network models use relatively deep networks, which will cause large computational overhead and it is difficult to achieve real-time detection. Summary of the Invention

[0005] The purpose of the present invention is to provide a method, system and device for constructing few-shot industrial image defect detection models, aiming to solve the construction of few-shot industrial image defect detection models.

[0006] The present invention provides a method for constructing a few-shot industrial image defect detection model based on image segmentation, including:

[0007] S1. Obtain an industrial image dataset, and divide the industrial image dataset into a training set and a test set;

[0008] S2. Preprocess the training set and the test set to enhance the image contrast;

[0009] S3. Randomly vertically flip and horizontally flip the preprocessed training set to obtain training sample data;

[0010] S4. Input the training sample data into the segmentation network for training. After the training is completed, a mask image is obtained, completing the first step of constructing the detection model.

[0011] S5. Correct the mask image to obtain the defective area of the industrial image, completing the second step of constructing the detection model.

[0012] S6. Stitch the mask image and the defective area of the industrial image to obtain a two-channel feature map. Input the two-channel feature map into the decision network for training. After the training is completed, obtain the defect probability of the industrial image in the training set. Determine whether the industrial image in the training set has defects according to the defect probability, completing the third step of constructing the detection model.

[0013] S7. Input the preprocessed test set into the detection model to obtain the classification result of whether there are defects in the test set, and evaluate the detection model according to relevant evaluation indicators.

[0014] The present invention also provides a few-shot industrial image defect detection model construction system based on image segmentation, including:

[0015] Partition module: used to obtain the industrial image data set and partition the industrial image data set into a training set and a test set;

[0016] Preprocessing module: used to preprocess the training set and the test set to enhance the image contrast;

[0017] Training sample data module: used to randomly vertically flip and horizontally flip the preprocessed training set to obtain the training sample data;

[0018] Segmentation network module: used to input the training sample data into the segmentation network for training. After the training is completed, a mask image is obtained, completing the first step of constructing the detection model.

[0019] Correction module: used to correct the mask image to obtain the defective area of the industrial image, completing the second step of constructing the detection model.

[0020] Decision network module: used to stitch the mask image and the defective area of the industrial image to obtain a two-channel feature map. Input the two-channel feature map into the decision network for training. After the training is completed, obtain the defect probability of the industrial image in the training set. Determine whether the industrial image in the training set has defects according to the defect probability, completing the third step of constructing the detection model.

[0021] Evaluation module: used to input the preprocessed test set into the detection model to obtain the classification result of whether there are defects in the test set, and evaluate the detection model according to relevant evaluation indicators.

[0022] An embodiment of the present invention further provides a method for constructing a few-shot industrial image defect detection model, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the steps of the above method are implemented when the computer program is executed by the processor.

[0023] An embodiment of the present invention further provides a computer-readable storage medium, on which an implementation program for information transmission is stored, and the steps of the above method are implemented when the program is executed by a processor.

[0024] By using the embodiment of the present invention, the detection accuracy and speed of defect images can be significantly improved under the condition of training with a small number of defect samples, and it can be used in a variety of detection scenarios, and the constructed model can also be evaluated.

[0025] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it is implemented in accordance with the content of the description, and in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the following specifically describes the embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0027] Figure 1 is a flowchart of a method for constructing a few-shot industrial image defect detection model based on image segmentation according to an embodiment of the present invention;

[0028] Figure 2 is a model flowchart of a method for constructing a few-shot industrial image defect detection model based on image segmentation according to an embodiment of the present invention;

[0029] Figure 3 is a schematic diagram of a segmentation network of a method for constructing a few-shot industrial image defect detection model based on image segmentation according to an embodiment of the present invention;

[0030] Figure 4 is a schematic diagram of an attention module of a method for constructing a few-shot industrial image defect detection model based on image segmentation according to an embodiment of the present invention;

[0031] Figure 5 is a schematic diagram of a decision network module of a method for constructing a few-shot industrial image defect detection model based on image segmentation according to an embodiment of the present invention;

[0032] Figure 6 It is a schematic diagram of a few-shot industrial image defect detection model construction system based on image segmentation according to an embodiment of the present invention;

[0033] Figure 7 It is a schematic diagram of a few-shot industrial image defect detection model construction device based on image segmentation according to an embodiment of the present invention.

[0034] Description of reference numerals:

[0035] 610: Division module; 620: Preprocessing module; 630: Training sample data module; 640: Segmentation network module; 650: Correction module; 660: Decision network module; 670: Evaluation module. Detailed implementation manners

[0036] Next, the technical solutions of the present invention will be clearly and completely described in conjunction with the embodiments. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0037] Method embodiment

[0038] According to an embodiment of the present invention, a method for constructing a few-shot industrial image defect detection model based on image segmentation is provided. Figure 1 It is a flowchart of a method for constructing a few-shot industrial image defect detection model based on image segmentation according to an embodiment of the present invention, as Figure 1 shown, and specifically includes:

[0039] S1. Obtain an industrial image dataset, and divide the industrial image dataset into a training set and a test set;

[0040] S2. Preprocess the training set and the test set to enhance the image contrast;

[0041] S2 specifically includes:

[0042] S21. Calculate the gray-scale histograms of the training set and the test set;

[0043] S22. Perform normalization processing on the gray-scale histograms;

[0044] S23. Calculate the minimum pixel point and the maximum pixel point of the normalized histogram;

[0045] S24. Gradually shift the minimum pixel point and the maximum pixel point to the right and left, and calculate the sum of the probability densities on both sides of the histogram;

[0046] S25. After the sum of the probability densities exceeds a given threshold, record the minimum pixel point index and the maximum pixel point index;

[0047] S26. Perform gray-scale stretching on the training set and the test set, and its calculation formula is:

[0048] output = (input - min_index) * 255 / (max_index - min_index)

[0049] where input is the input image, min_index is the minimum pixel point index described above, max_index is the maximum pixel point index described above, and output is the output image;

[0050] S27. Traverse the output image pixel by pixel. If the pixel value is greater than 255, modify the pixel value of this point to 255. If the pixel value is less than 0, modify the pixel value of this point to 0.

[0051] S3. Randomly vertically flip and horizontally flip the preprocessed training set to obtain training sample data;

[0052] S4. Input the training sample data into the segmentation network for training. After the training is completed, a mask image is obtained, and the first step of constructing the detection model is completed;

[0053] S4 specifically includes: inputting the training sample data into the segmentation network for training. Among them, the segmentation network is built through the PyTorch deep learning framework, and the segmentation network consists of a set of left-right symmetric encoders, decoders, and two attention blocks;

[0054] The encoder consists of three layers. The structure of the first layer is a convolutional block formed by cascading two groups of 3 * 3 convolutions + batch normalization + ReLU activation. The structures of the latter two layers are dilated convolutional blocks formed by cascading two groups of dilated convolutions + batch normalization + ReLU activation; A max-pooling downsampling layer is connected between two adjacent layers of the encoder;

[0055] The decoder consists of three layers. The first two layers are both formed by cascading two groups of 3 * 3 convolutions + batch normalization + ReLU activation. The last layer is formed by cascading two groups of 3 * 3 convolutions + batch normalization + Sigmoid activation. An upsampling layer is connected between adjacent decoders;

[0056] The calculation formula of the attention block is:

[0057] Y(X) = X * (Sigmoid(Conv2(ReLU(Conv1(X)))))

[0058] Where X is the result of global average pooling on the output feature map of the first or second layer of the encoder, with a size of C*1*1, where C is the number of channels; Conv1 represents changing the number of channels of the feature map to C / r through a 1*1 convolution, and Conv2 represents restoring the number of channels of the feature map to C through a 1*1 convolution; ReLU and Sigmoid are the activation functions after the Conv1 and Conv2 operations respectively;

[0059] After the training is completed, a mask image is obtained, completing the first step of constructing the detection model.

[0060] S5. Correct the mask image to obtain the defective area of the industrial image, completing the second step of constructing the detection model;

[0061] S5 specifically includes:

[0062] S51: Binarize the mask image;

[0063] S52: Perform dilation processing on the binarized mask image;

[0064] S53: Calculate the contours of the dilated mask image and calculate the area of each contour;

[0065] S54: Compare the contour areas to obtain the contour with the largest area, calculate the minimum bounding rectangle of the contour with the largest area, and record the coordinate information of the minimum bounding rectangle;

[0066] S55: Calculate the ratio of the longest side to the shortest side of the minimum bounding rectangle according to the rectangle coordinate information, determine the defect type, and extract the defective area on the original training set image according to the defect type.

[0067] S6. Stitch the mask image and the defective area of the industrial image to obtain a two-channel feature map, input the two-channel feature map into the decision network for training, and after the training is completed, obtain the defect probability of the industrial image in the training set. Judge whether the industrial image in the training set is defective, completing the third step of constructing the detection model;

[0068] S6 specifically includes:

[0069] Stitch the output result of the segmentation network and the defective area to obtain a two-channel feature map. Use the PyTorch deep learning framework to build a decision network. The decision network includes: four groups of 3*3 depthwise separable convolutions, a Sigmoid activation layer, and a global average pooling layer. Each group of depthwise separable convolutions is followed by a max-pooling downsampling layer. Train the decision network according to the two-channel feature map. After the training is completed, obtain the defect probability of the industrial image. Judge whether the industrial image in the training set is defective, completing the third step of constructing the detection model.

[0070] S7. Input the preprocessed test set into the detection model to obtain the classification results of whether there are defects in the test set, and evaluate the detection model according to relevant evaluation indicators.

[0071] S7 specifically includes:

[0072] S71. Input the preprocessed test set into the segmentation network to obtain the test mask image;

[0073] S72. Correct the test mask image to obtain the defective area of the test industrial image;

[0074] S73. Stitch the test mask image and the defective area of the test industrial image to obtain a test dual-channel feature map, input the test dual-channel feature map into the decision network to obtain the defect probability of the industrial image, and judge whether the industrial image of the test set is defective according to the defect probability;

[0075] S74. Obtain the number of images correctly classified by the decision network on the test set and the total number of images in the test set, and calculate the accuracy index;

[0076] S75. Obtain the number of images correctly classified as defective by the decision network on the test set and the total number of images classified as defective images in the test set, and calculate the precision rate index;

[0077] S76. Obtain the number of images correctly classified as defective by the decision network on the test set and the total number of defective images in the test set, and calculate the recall rate evaluation index;

[0078] S77. Obtain the precision rate and recall rate evaluation indicators, and calculate the harmonic mean evaluation indicator of the detection model.

[0079] The following specific implementation manners take the magnetic tile defect data set as an example, and the steps include:

[0080] Figure 2 It is the model flow chart of the method for constructing a few-shot industrial image defect detection model based on image segmentation in the embodiment of the present invention; as Figure 2 shown:

[0081] Step 1: Obtain the magnetic tile image data set and divide it into a training set and a test set;

[0082] Specifically, the training set images include: 30 images of hole defects, 30 images of crack defects, and 600 defect-free images. The test set images include: 112 remaining defect images and 342 defect-free images;

[0083] Step 2: Preprocess the images to enhance the contrast of the images. Step S2 specifically includes the following steps:

[0084] Step 2.1: Calculate the grayscale histogram of the image and normalize the histogram;

[0085] Step 2.2: Shift the minimum pixel index of the histogram one unit to the right and the maximum pixel index one unit to the left, and calculate the sum of the probability densities on the left side of the minimum pixel index and the sum of the probability densities on the right side of the maximum pixel index at this time;

[0086] Step 2.3: Repeat Step S2.2 until the sum of the probability densities exceeds a given threshold, and record the minimum pixel index and the maximum pixel index at this time;

[0087] Step 2.4: Perform grayscale stretching on the image, and its calculation formula is:

[0088] output=(input - min_index)*255 / (max_index - min_index)

[0089] where input is the input image, min_index and max_index are the minimum pixel index and the maximum pixel index respectively, and output is the output image after grayscale stretching;

[0090] Step 2.5: Traverse each pixel of the image. If the pixel value is greater than 255, modify the pixel at this point to 255. If the pixel value is less than 0, modify the pixel at this point to 0;

[0091] Figure 3 is the schematic diagram of the segmentation network of the method for constructing a few - shot industrial image defect detection model based on image segmentation in an embodiment of the present invention; as Figure 3 shown:

[0092] Step 3: Optionally, refer to Figure 3 Use the PyTorch deep learning framework to build an image segmentation network;

[0093] Specifically, the segmentation network architecture is a U - shaped structure. The U - shaped structure includes an encoder, two attention layers, and a decoder. The structure of the encoder is:

[0094] The encoder has three convolutional blocks. The convolutional block is composed of two consecutive 3*3 convolutional layers + batch normalization layer + ReLU activation layer. A max - pooling downsampling layer is connected after the first two convolutional blocks;

[0095] Figure 4 is the schematic diagram of the attention module of the method for constructing a few - shot industrial image defect detection model based on image segmentation in an embodiment of the present invention; as Figure 4 shown:

[0096] Refer to Figure 4, the attention layer is used for the outputs of the first two convolutional blocks of the encoder, and its calculation formula is:

[0097] Y(X) = X * (Sigmoid(Conv2(ReLU(Conv1(X)))))

[0098] where X is the result after global average pooling of the output feature map of the first or second layer of the encoder, with a size of C * 1 * 1, and C is the number of channels. Conv1 represents changing the number of channels of the feature map to C / r through a 1*1 convolution, where r is the reduction ratio of the number of channels. Conv2 represents restoring the number of channels of the feature map to C through a 1*1 convolution. ReLU and Sigmoid are the activation functions after the Conv1 and Conv2 operations respectively;

[0099] The decoder consists of three convolutional blocks. The first two convolutional blocks are composed of two sets of consecutive 3*3 convolutional layers + batch normalization layers + ReLU activation layers. The last convolutional block replaces the above ReLU activation layer with a Sigmoid activation layer. The input of each convolutional block of the decoder is the output of the attention layer after 2x upsampling of the feature map of the previous layer;

[0100] Step 3.1: Use the images in the training set to train the segmentation network, and its loss function is:

[0101]

[0102] where n is the number of samples, y n is the label mask image, and x n is the output of the segmentation network;

[0103] Step 4: Perform correction processing on the output result of the segmentation network, which includes:

[0104] Step 4.1: Perform binarization processing on the image to eliminate darker segmentation points;

[0105] Step 4.2: Perform morphological processing on the binarized image, and use a 5*5 dilation kernel to perform dilation processing on it;

[0106] Step 4.3: Calculate the contour of the dilated image, and retain the contour with the largest area to remove interference;

[0107] Step 4.4: Calculate the minimum bounding rectangle of the contour and record the rectangle coordinate information;

[0108] Step 4.5: Make a defect type judgment according to the ratio of the longest side to the shortest side of the rectangle, and extract the defect area on the original image according to the rectangle coordinates;

[0109] Figure 5It is a schematic diagram of the decision network module of the method for constructing a few-shot industrial image defect detection model based on image segmentation according to an embodiment of the present invention; as Figure 5 shown:

[0110] Step 5: Optionally, according to the defect type, with reference to Figure 5 Use the deep PyTorch deep learning framework to build a decision network for making decision classification on images of different defect types;

[0111] Specifically, the decision network is composed of four groups of 3*3 depth separable convolutions, a Sigmoid activation layer and a global average pooling layer, and each group of depth separable convolutions is followed by a max pooling downsampling layer;

[0112] Step 6: Send the images in the training set into the trained segmentation network, perform the refinement processing on the output results, and obtain the defect region and the segmentation mask image;

[0113] Step 7: Concatenate the defect region and the segmentation mask image on the channel to obtain a two-channel feature map;

[0114] Step 8: Use the two-channel feature map to train the decision network, and the loss function is:

[0115]

[0116] where n is the number of samples, x n is the probability value predicted by the decision network that the image has a defect, and y n is the label value, the label 0 means no defect, and the label 1 means there is a defect;

[0117] Step 9: Optionally, after the industrial image defect detection model is trained, evaluate the test set in the magnetic tile defect dataset according to the evaluation metrics. The evaluation metrics include accuracy, precision, recall, and weighted harmonic mean (F-Measure). The formulas are as follows:

[0118]

[0119]

[0120]

[0121]

[0122] Among them, TN is the number of images in which negative samples (images without defects) are correctly classified, TP is the number of images in which positive samples (defective images) are correctly classified, FN is the number of images in which positive samples are misclassified, and FP is the number of images in which negative samples are misclassified. is the weight.

[0123] In summary, the present invention proposes a three-stage industrial image defect detection method. The model of the present invention is implemented based on a convolutional neural network. First, in the segmentation network module, the present invention proposes to add two attention layers before the U-Net skip connection, which can emphasize the feature maps of important channels and suppress the feature maps of unimportant channels, thereby improving the segmentation performance of the model. Secondly, two dilated convolutional layers are also introduced in the segmentation network to increase the receptive field of the model to capture defects of different sizes. In addition, the present invention also proposes a refinement processing module for further correcting the segmentation result and extracting the defective area in the original image. Finally, the decision network proposed by the present invention makes a classification decision by combining the local features of the defective area and the global segmentation information, and uses depthwise separable convolutions with less computational complexity to speed up the calculation speed.

[0124] Compared with the prior art, the present invention has the following advantages:

[0125] 1. Good classification performance can be achieved by only using 30 images of each defect type for training;

[0126] 2. Attention layers and dilated convolutional layers are added to the segmentation network, which can emphasize useful channel features and suppress useless channel features during skip connection, thereby improving the segmentation performance. In addition, two dilated convolutional layers are introduced to adapt to defects of different sizes and enhance the feature capture ability of the model;

[0127] 3. The refinement processing module further corrects the segmentation result, making the extraction of the defective area more accurate;

[0128] 4. The decision network uses two inputs: the segmentation mask image and the defective area image, makes a decision by combining the information of both, improves the classification accuracy, and uses depthwise separable convolutions to greatly reduce the computational complexity and speed up the decision-making speed.

[0129] System embodiment

[0130] According to an embodiment of the present invention, a few-shot industrial image defect detection model construction system based on image segmentation is provided. Figure 6 is a schematic diagram of the few-shot industrial image defect detection model construction system based on image segmentation according to an embodiment of the present invention, as Figure 6 shown, specifically including:

[0131] Partitioning module 610: It is used to obtain an industrial image dataset and partition the industrial image dataset into a training set and a test set;

[0132] Preprocessing module 620: It is used to preprocess the training set and the test set to enhance the image contrast;

[0133] Training sample data module 630: It is used to randomly flip the preprocessed training set vertically and horizontally to obtain training sample data;

[0134] Segmentation network module 640: It is used to input the training sample data into the segmentation network for training. After the training is completed, a mask image is obtained, and the first step of constructing the detection model is completed;

[0135] Calibration module 650: It is used to calibrate the mask image to obtain the defective area of the industrial image, and the second step of constructing the detection model is completed;

[0136] Decision network module 660: It is used to splice the mask image and the defective area of the industrial image to obtain a two-channel feature map. The two-channel feature map is input into the decision network for training. After the training is completed, the defect probability of industrial images of different defect types is obtained. Whether the industrial images in the training set are defective is judged according to the defect probability, and the third step of constructing the detection model is completed;

[0137] Evaluation module 670: It is used to input the preprocessed test set into the detection model to obtain the classification result of whether there are defects in the test set, and evaluate the detection model according to relevant evaluation indicators.

[0138] The preprocessing module 620 is specifically used for:

[0139] Calculate the grayscale histograms of the training set and the test set;

[0140] Normalize the grayscale histograms;

[0141] Calculate the minimum pixel point and the maximum pixel point for the normalized histogram;

[0142] Gradually shift the minimum pixel point and the maximum pixel point to the right and left, and calculate the sum of the probability densities on both sides of the histogram;

[0143] After the sum of the probability densities exceeds a given threshold, record the minimum pixel point index and the maximum pixel point index;

[0144] Perform grayscale stretching on the training set and the test set, and its calculation formula is:

[0145] output=(input - min_index)*255 / (max_index - min_index)

[0146] Where input is the input image, min_index is the index of the minimum pixel point, max_index is the index of the maximum pixel point, and output is the output image;

[0147] Traverse the output image pixel by pixel. If the pixel value is greater than 255, modify the pixel value of this point to 255. If the pixel value is less than 0, modify the pixel value of this point to 0;

[0148] The segmentation network module is specifically used for: inputting training sample data into the segmentation network for training. Among them, the segmentation network is built through the PyTorch deep learning framework, and the segmentation network consists of a set of left-right symmetric encoders, decoders, and two attention blocks;

[0149] The encoder consists of three layers. The structure of the first layer is a convolutional block formed by cascading two groups of 3*3 convolutions + batch normalization + ReLU activation. The structures of the latter two layers are dilated convolutional blocks formed by cascading two groups of dilated convolutions + batch normalization + ReLU activation; A max-pooling downsampling layer is connected between two adjacent layers of the encoder;

[0150] The decoder consists of three layers. The first two layers are both formed by cascading two groups of 3*3 convolutions + batch normalization + ReLU activation. The last layer is formed by cascading two groups of 3*3 convolutions + batch normalization + Sigmoid activation. An upsampling layer is connected between adjacent decoders;

[0151] The calculation formula of the attention block is:

[0152] Y(X) = X * (Sigmoid(Conv2(ReLU(Conv1(X)))))

[0153] Where X is the result after global average pooling of the output feature map of the first or second layer of the encoder, and its size is C*1*1, where C is the number of channels; Conv1 represents changing the number of channels of the feature map to C / r through a 1*1 convolution, and Conv2 represents restoring the number of channels of the feature map to C through a 1*1 convolution; ReLU and Sigmoid are the activation functions after the Conv1 and Conv2 operations respectively;

[0154] After training is completed, a mask image is obtained, completing the first step of constructing the detection model;

[0155] The correction module is specifically used for:

[0156] Perform binarization processing on the mask image;

[0157] Perform dilation processing on the binarized mask image;

[0158] Calculate the contours of the dilated mask image and calculate the areas of each contour;

[0159] Compare the contour areas to obtain the contour with the largest area, calculate the minimum bounding rectangle of the contour with the largest area, and record the coordinate information of the minimum bounding rectangle;

[0160] Calculate the ratio of the longest side to the shortest side of the minimum bounding rectangle according to the rectangular coordinate information, determine the defect type, and extract the defect area on the original training set image according to the defect type;

[0161] The decision network module is specifically used for:

[0162] Concatenate the output result of the segmentation network with the defect area to obtain a two-channel feature map. Build a decision network using the PyTorch deep learning framework. The decision network includes: four groups of 3*3 depthwise separable convolutions, a Sigmoid activation layer, and a global average pooling layer. Each group of depthwise separable convolutions is followed by a max-pooling downsampling layer. Train the decision network according to the two-channel feature map. After training is completed, obtain the defect probabilities of industrial images of different defect types. Judge whether the industrial images in the training set are defective according to the defect probabilities, and complete the third step of constructing the detection model;

[0163] The evaluation module is specifically used for:

[0164] Input the preprocessed test set into the segmentation network to obtain a test mask image;

[0165] Correct the test mask image to obtain the defect area of the test industrial image;

[0166] Concatenate the test mask image with the defect area of the test industrial image to obtain a test two-channel feature map. Input the test two-channel feature map into the decision network to obtain the defect probability of the industrial image. Judge whether the industrial image in the test set is defective according to the defect probability;

[0167] Obtain the number of images correctly classified by the decision network on the test set and the total number of images in the test set, and calculate the accuracy index;

[0168] Obtain the number of images correctly classified as defective by the decision network on the test set and the total number of images classified as defective in the test set, and calculate the precision index;

[0169] Obtain the number of images correctly classified as defective by the decision network on the test set and the total number of defective images in the test set, and calculate the recall rate evaluation index;

[0170] Obtain the number of images correctly classified as defective by the decision network on the test set and the total number of defective images in the test set, and calculate the recall rate evaluation index;

[0171] Obtain the precision rate and recall rate evaluation metrics, and calculate the harmonic mean evaluation metric of the detection model.

[0172] The embodiments of the present invention are system embodiments corresponding to the above method embodiments. The specific operations of each module can be understood with reference to the description of the method embodiments, and will not be elaborated here.

[0173] Device Embodiment 1

[0174] The embodiments of the present invention provide an apparatus for constructing a few-shot industrial image defect detection model for image segmentation, as Figure 7 shown, including: a memory 70, a processor 72, and a computer program stored on the memory 70 and executable on the processor 72. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0175] Device Embodiment 2

[0176] The embodiments of the present invention provide a computer-readable storage medium, on which an implementation program for information transmission is stored. When the program is executed by the processor 72, the steps in the above method embodiments are implemented.

[0177] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the technical solutions of the embodiments of the present invention deviate from the scope of the present solution.

Claims

1. A method for constructing a few-shot industrial image defect detection model based on image segmentation, characterized in that Including: S1. Obtain an industrial image dataset and divide the industrial image dataset into a training set and a test set; S2. Preprocess the training set and the test set to enhance the image contrast; S3. Randomly vertically flip and horizontally flip the preprocessed training set to obtain training sample data; S4. Input the training sample data into a segmentation network for training. After the training is completed, a mask image is obtained, and the first step of constructing the detection model is completed; S5. Correct the mask image to obtain the defective area of the industrial image, and complete the second step of constructing the detection model; S6. Stitch the mask image and the defective area of the industrial image to obtain a two-channel feature map. Input the two-channel feature map into a decision network for training. After the training is completed, obtain the defect probabilities of industrial images of different defect types in the training set. Determine whether the industrial images in the training set are defective according to the defect probabilities, and complete the third step of constructing the detection model; S7. Input the preprocessed test set into the detection model to obtain the classification result of whether there are defects in the test set, and evaluate the detection model according to relevant evaluation indicators.

2. The method according to claim 1, wherein The specific content of S2 includes: S21. Calculate the grayscale histograms of the training set and the test set; S22. Perform normalization processing on the grayscale histograms; S23. Calculate the minimum pixel point index and the maximum pixel point index of the normalized histogram; S24. Shift the minimum pixel point index of the histogram one unit to the right and the maximum pixel point index one unit to the left, and calculate the sum of the probability densities on the left side of the minimum pixel point index and the sum of the probability densities on the right side of the maximum pixel point index; S25. After the sum of the probability densities exceeds a given threshold, record the minimum pixel point index and the maximum pixel point index; S26. Perform grayscale stretching on the training set and the test set, and its calculation formula is: output=(input - min_index)*255 / (max_index - min_index) where input is the input image, min_index is the minimum pixel point index mentioned above, max_index is the maximum pixel point index mentioned above, and output is the output image; S27. Traverse each pixel of the output image. If the pixel value is greater than 255, modify the pixel value of this point to 255. If the pixel value is less than 0, modify the pixel value of this point to 0.

3. The method according to claim 2, wherein The specific content of S4 includes: Input the training sample data into a segmentation network for training. Among them, the segmentation network is built through the PyTorch deep learning framework, and the segmentation network consists of a group of left-right symmetric encoders, decoders, and two attention blocks; The encoder consists of three layers. The structure of the first layer is the concatenation of 3*3 convolution*2 + batch normalization + ReLU activation. The structures of the latter two layers are to replace the above 3*3 convolution*2 with dilated convolution*2 to expand the receptive field of the convolution. A max-pooling downsampling layer is connected between two adjacent layers of the encoder; The decoder consists of three layers. The first two layers are both connected in series with 3*3 convolution * 2 + batch normalization + ReLU activation. The last layer replaces the above ReLU activation with Sigmoid activation. An upsampling layer is connected between adjacent decoders; The calculation formula of the attention block is: Y(X) = X * (Sigmoid(Conv2(ReLU(Conv1(X))))) where X is the result of global average pooling of the output feature map of the first or second layer of the encoder, with a size of C * 1 * 1, where C is the number of channels; Conv1 represents changing the number of feature map channels to C / r through 1*1 convolution, where r is the reduction ratio of the number of channels, and Conv2 represents restoring the number of feature map channels to C through 1*1 convolution; ReLU and Sigmoid are the activation functions after the Conv1 and Conv2 operations respectively; After training is completed, a mask image is obtained, completing the first step of constructing the detection model.

4. The method according to claim 3, wherein The specific steps of S5 are as follows: S51: Binarize the mask image; S52: Dilate the binarized mask image; S53: Calculate the contours of the dilated mask image and calculate the area of each contour; S54: Compare the contour areas to obtain the contour with the largest area, calculate the minimum bounding rectangle of the contour with the largest area, and record the coordinate information of the minimum bounding rectangle; S55: Calculate the ratio of the longest side to the shortest side of the minimum bounding rectangle according to the rectangle coordinate information, determine the defect type, and extract the defect area from the original training set image according to the defect type.

5. The method according to claim 4, wherein The specific steps of S6 are as follows: Concatenate the output result of the segmentation network with the defect area to obtain a two-channel feature map. Use the PyTorch deep learning framework to build a decision network. The decision network includes: four groups of 3*3 depthwise separable convolution, Sigmoid activation layer, and global average pooling layer. A max-pooling downsampling layer is connected after each group of depthwise separable convolution. Train the decision network according to the two-channel feature map. After training is completed, obtain the defect probability of the industrial image. Judge whether the industrial image in the training set is defective according to the defect probability, completing the third step of constructing the detection model.

6. The method according to claim 5, wherein The specific steps of S7 are as follows: S71: Input the preprocessed test set into the segmentation network to obtain a test mask image; S72: Correct the test mask image to obtain the defect area of the test industrial image; S73: Concatenate the test mask image with the defect area of the test industrial image to obtain a test two-channel feature map. Input the test two-channel feature map into the decision network to obtain the defect probability of the industrial image. Judge whether the industrial image in the test set is defective according to the defect probability; S74: Obtain the number of images correctly classified by the decision network on the test set and the total number of images in the test set, and calculate the accuracy index; S75: Obtain the number of images correctly classified as defective by the decision network on the test set and the total number of images classified as defective in the test set, and calculate the precision index; S76: Obtain the number of images correctly classified as defective by the decision network on the test set and the total number of defective images in the test set, and calculate the recall evaluation index; S77. Obtain the precision rate and recall rate evaluation metrics, and calculate the harmonic mean evaluation metric of the detection model.

7. A system for constructing a few-shot industrial image defect detection model based on image segmentation, characterized in that Including: Partitioning module: used to obtain an industrial image dataset and partition the industrial image dataset into a training set and a test set; Preprocessing module: used to preprocess the training set and the test set to enhance the image contrast; Training sample data module: used to randomly vertically flip and horizontally flip the preprocessed training set to obtain training sample data; Segmentation network module: used to input the training sample data into the segmentation network for training. After the training is completed, a mask image is obtained, completing the first step of constructing the detection model; Calibration module: used to calibrate the mask image to obtain the defective area of the industrial image, completing the second step of constructing the detection model; Decision network module: used to splice the mask image and the defective area of the industrial image to obtain a two-channel feature map, input the two-channel feature map into the decision network for training. After the training is completed, obtain the defect probability of the industrial image in the training set, and determine whether the industrial image in the training set is defective according to the defect probability, completing the third step of constructing the detection model; Evaluation module: used to input the preprocessed test set into the detection model to obtain the classification result of whether there are defects in the test set, and evaluate the detection model according to relevant evaluation metrics.

8. The system according to claim 7, wherein Including: Specifically, the preprocessing module is used for: Calculate the gray-level histograms of the training set and the test set; Perform normalization processing on the gray-level histograms; Calculate the minimum pixel point index and the maximum pixel point index for the normalized histogram; Shift the minimum pixel point index of the histogram one unit to the right and the maximum pixel point index one unit to the left, and calculate the sum of the probability densities on the left side of the minimum pixel point index and the sum of the probability densities on the right side of the maximum pixel point index; After the sum of the probability densities exceeds a given threshold, record the minimum pixel point index and the maximum pixel point index; Perform gray-level stretching on the training set and the test set, and its calculation formula is: output=(input - min_index)*255 / (max_index - min_index) where input is the input image, min_index is the minimum pixel point index mentioned above, max_index is the maximum pixel point index mentioned above, and output is the output image; Traverse each pixel of the output image. If the pixel value is greater than 255, modify the pixel value of this point to 255. If the pixel value is less than 0, modify the pixel value of this point to 0; Specifically, the segmentation network module is used for: input the training sample data into the segmentation network for training. Among them, the segmentation network is built through the PyTorch deep learning framework, and the segmentation network consists of a group of left-right symmetric encoders, decoders, and two attention blocks; The encoder consists of three layers. The structure of the first layer is 3*3 convolution*2 + batch normalization + ReLU activation in series. The structures of the latter two layers are to replace the above 3*3 convolution*2 with dilated convolution*2 to expand the receptive field of the convolution. A max-pooling downsampling layer is connected between two adjacent layers of the encoder; The decoder consists of three layers. The first two layers are both concatenated with 3*3 convolution*2 + batch normalization + ReLU activation. The last layer replaces the above ReLU activation with Sigmoid activation. An upsampling layer is connected between adjacent decoders; The calculation formula of the attention block is: Y(X) = X * (Sigmoid(Conv2(ReLU(Conv1(X))))) where X is the result of global average pooling of the output feature map of the first or second layer of the encoder, with a size of C*1*1, and C is the number of channels; Conv1 represents changing the number of channels of the feature map to C / r through 1*1 convolution, where r is the reduction ratio of the number of channels, and Conv2 represents restoring the number of channels of the feature map to C through 1*1 convolution; ReLU and Sigmoid are the activation functions after the Conv1 and Conv2 operations respectively; After training is completed, a mask image is obtained, completing the first step of constructing the detection model; The correction module is specifically used for: Performing binarization processing on the mask image; Performing dilation processing on the binarized mask image; Calculating the contours of the dilated mask image and calculating the area of each contour; Comparing the contour areas to obtain the contour with the largest area, calculating the minimum bounding rectangle of the contour with the largest area, and recording the coordinate information of the minimum bounding rectangle; Calculating the ratio of the longest side to the shortest side of the minimum bounding rectangle according to the rectangular coordinate information, determining the defect type, and extracting the defect area from the original training set image according to the defect type; The decision network module is specifically used for: Concatenating the output result of the segmentation network with the defect area to obtain a two-channel feature map. Using the PyTorch deep learning framework to build a decision network, the decision network includes: four groups of 3*3 depthwise separable convolutions, a Sigmoid activation layer, and a global average pooling layer. A max pooling downsampling layer is connected after each group of depthwise separable convolutions. Training the decision network according to the two-channel feature map. After training is completed, obtaining the defect probability of the industrial image, and judging whether the industrial image in the training set is defective, completing the third step of constructing the detection model; The evaluation module is specifically used for: Inputting the preprocessed test set into the segmentation network to obtain a test mask image; Correcting the test mask image to obtain the defect area of the test industrial image; Concatenating the test mask image with the defect area of the test industrial image to obtain a test two-channel feature map, inputting the test two-channel feature map into the decision network to obtain the defect probability of the industrial image, and judging whether the industrial image in the test set is defective according to the defect probability; Obtaining the number of images correctly classified by the decision network on the test set and the total number of images in the test set, and calculating the accuracy index; Obtaining the number of images correctly classified as defective by the decision network on the test set and the total number of images classified as defective in the test set, and calculating the precision index; Obtaining the number of images correctly classified as defective by the decision network on the test set and the total number of defective images in the test set, and calculating the recall evaluation index; Obtaining the precision and recall evaluation indexes of the decision network on the test set, and calculating the harmonic mean evaluation index of the detection model.

9. An apparatus for constructing a few-shot industrial image defect detection model based on image segmentation, characterized in that Including: A memory, a processor, and a computer program stored on the memory and executable on the processor, where when the computer program is executed by the processor, the steps of the method for constructing a few-shot industrial image defect detection model based on image segmentation according to any one of claims 1 to 6 are implemented.

10. A computer-readable storage medium, characterized in that, An implementation program for information transmission is stored on the computer-readable storage medium, and when the program is executed by the processor, the steps of the method for constructing a few-shot industrial image defect detection model based on image segmentation according to any one of claims 1 to 6 are implemented.