A Bottle Mouth Defect Detection Method Based on Dual-Stream Semi-Mask Reconstruction

Through the dual-stream semi-mask reconstruction method, the automatic encoder network with Transformer structure is used to reconstruct bottle mouth images and defect detection, which solves the problems of high computing resource consumption and insufficient model characterization capabilities in the prior art, and achieves efficient and accurate bottle mouth defect detection.

CN116071302BActive Publication Date: 2025-06-17XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211641123.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-20
Publication Date
2025-06-17
Estimated Expiration
2042-12-20

AI Technical Summary

Technical Problem

The existing bottle mouth defect detection method based on convolutional neural networks has problems such as high computing resource consumption and insufficient model characterization capabilities, resulting in low defect detection accuracy.

Method used

Using a dual-stream semi-mask reconstruction method, the bottle mouth image is converted into two mask images with complementary masks through a random mask module, and the automatic encoder network with Transformer structure is used for reconstruction, and defects in the image are detected through the exception scoring module.

Benefits of technology

More efficient and more accurate bottle mouth defect detection is achieved, which reduces computing resource consumption, improves model characterization capabilities, reduces labor costs and improves detection efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116071302B_ABST
    Figure CN116071302B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for detecting bottle mouth defects based on dual-stream semi-mask reconstruction, including: constructing a bottle mouth defect detection system, which includes a random mask module, an autoencoder network, and an anomaly scoring module connected in sequence; inputting a training data set composed of defect-free bottle mouth images into the bottle mouth defect detection system to train the autoencoder network and obtain a trained bottle mouth defect detection system; using the random mask module to convert the bottle mouth image to be detected into two mask images with complementary masks; inputting the two mask images into the trained autoencoder network respectively to obtain reconstructed images; using the anomaly scoring module to compare the bottle mouth image to be detected with the reconstructed images to confirm whether there are bottle mouth defects in the bottle mouth image to be detected. The present invention completes the screening service of bottle mouth defect products by training a defect detector, reduces labor costs, and overcomes the problems of slow manual sorting speed and easy errors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of bottle mouth surface defect detection, and particularly relates to a bottle mouth defect detection method based on dual-stream semi-mask reconstruction. Background Art

[0002] In recent years, industrial production requires enterprises to strengthen supply-side structural reform, improve product quality, continuously optimize the production structure, and promote the development of industrial production towards the intelligent direction. Quality problems of finished products often occur in the industrial production process. Some products can be quality-tested through their apparent conditions. Therefore, surface defect detection has important research value in industrial production. With the progress of industrial technology, more and more attention has been paid to the abnormal parts in natural image data. For example, defects in roads, bridges, and railways are detected to maintain infrastructure. In the biomedical field, early lesion detection is carried out through medical imaging processing. However, due to the large variety of defects in some industrial bottle mouth images and the fact that defective products are often rare, it is difficult to collect enough samples.

[0003] With the rapid development of deep neural networks, deep learning has shown great capabilities in characterizing the features of complex data such as high-dimensional data, time data, spatial data, and graphical data, making defect detection algorithms based on deep learning increasingly popular and applied to various tasks. Since defective image samples are difficult to obtain while normal image samples are easy to obtain, semi-supervised learning methods have become a hot topic in the research of image surface defect detection, and among them, reconstruction-based methods have also been explored by many researchers. The reconstruction-based method is to train the neural network structure only for the reconstruction of normal training images. Abnormal images are easily detected because they cannot be well reconstructed, and abnormal images are determined by defining an abnormal score. At present, most of the reconstruction-based methods are further explorations based on convolutional autoencoders and generative adversarial networks. In recent years, the rapid rise of Vision Transfomer in the field of computer vision has shown learning capabilities not weaker than convolutional neural networks, opening up a new path for surface defect detection.

[0004] However, although the existing technical solutions based on convolutional neural network autoencoders and generative adversarial networks have solved the problem of the small number of defect samples, the training based on generative adversarial networks requires expensive computing resources, and the model representation ability of the autoencoder-based model is insufficient, and the reconstructed image is quite different from the input image, resulting in low defect detection accuracy. Summary of the Invention

[0005] In order to solve the above problems existing in the prior art, the present invention provides a bottle mouth defect detection method based on dual-stream semi-mask reconstruction. The technical problems to be solved by the present invention are realized through the following technical solutions:

[0006] The present invention provides a method for detecting bottle mouth defects based on dual-stream semi-mask reconstruction, including:

[0007] S1: Construct a bottle mouth defect detection system, which includes a random mask module, an autoencoder network, and an anomaly scoring module connected in sequence;

[0008] S2: Input a training data set composed of defect-free bottle mouth images into the bottle mouth defect detection system to train the autoencoder network, and obtain a trained bottle mouth defect detection system;

[0009] S3: Use the random mask module to convert the bottle mouth image to be detected into two mask images with complementary masks;

[0010] S4: Input the two mask images into the trained autoencoder network respectively to obtain reconstructed images;

[0011] S5: Use the anomaly scoring module to compare the bottle mouth image to be detected with the reconstructed image to confirm whether there are bottle mouth defects in the bottle mouth image to be detected.

[0012] In an embodiment of the present invention, the random mask module is specifically used for:

[0013] Adjust the input bottle mouth image to a predetermined size, divide the resized image into regular and non-overlapping image blocks, and perform position encoding on each image block;

[0014] Randomly mask the multiple image blocks with a preset mask rate, and embed the position encoding information of the visible image blocks into the visible image blocks.

[0015] In an embodiment of the present invention, the autoencoder network includes an encoder and a decoder connected, where

[0016] The encoder includes a plurality of first Transformer structural units connected in sequence for extracting feature information of visible blocks;

[0017] The decoder includes a plurality of second Transformer structural units connected in sequence for performing position embedding on the output of the encoder and the masked part of the input bottle mouth image, and completing the reconstruction of the masked part of the input bottle mouth image.

[0018] In one embodiment of the present invention, the first Transformer structural unit and the second Transformer structural unit have the same structure, and respectively include a first normalization layer, a fully connected layer, a multi-head attention sub-unit, a first fusion sub-unit, a second normalization layer, a multi-layer perceptron, and a second fusion sub-unit, where,

[0019] The first normalization layer is used to perform normalization processing on the input visible image blocks;

[0020] The fully connected layer is used to perform a linear transformation on the normalized visible image blocks, and decompose them into three parts, namely query Q, key K, and value V;

[0021] The multi-head attention sub-unit is used to process each decomposed part through multiple attention heads respectively, obtain the feature maps of each part and splice them to obtain the spliced feature map;

[0022] The first fusion sub-unit is used to combine the spliced feature map with the input visible image blocks;

[0023] The second normalization layer is used to perform normalization processing on the combined feature map;

[0024] The multi-layer perceptron includes two fully connected layers and a non-linear activation function arranged between the two fully connected layers, and is used to improve the ability of the encoder to fit non-linearity by using the fully connected layer and the non-linear activation function;

[0025] The second fusion sub-unit is used to combine the output result of the first fusion sub-unit with the output result of the multi-layer perceptron and output a feature image.

[0026] In one embodiment of the present invention, S2 includes:

[0027] S2.1: Obtain multiple defect-free bottle mouth images, adjust them to a predetermined size, and divide each adjusted image into a plurality of regular and non-overlapping image blocks;

[0028] S2.2: Randomly mask the multiple image blocks with a preset masking rate, and embed the position encoding information of the visible image blocks into the visible image blocks;

[0029] S2.3: Input the visible image blocks embedded with position encoding information into the encoder, extract the feature information of the visible image blocks, and then input the output of the encoder and the masked part of the current defect-free bottle mouth image into the decoder together to reconstruct the masked part of the current defect-free bottle mouth image, and update the parameters in the encoder and the masker by using a loss function;

[0030] S2.4: When the value of the loss function tends to be stable, the training ends, and the trained encoder and decoder are obtained.

[0031] In an embodiment of the present invention, the loss function is:

[0032] loss = ∑(loss' × mask) / ∑mask

[0033] loss' = ||I in - I rec ||2

[0034] where I rec is the reconstructed image output by the decoder, I in is the image input to the encoder, and mask is the masked part of the input defect-free bottle mouth image.

[0035] In an embodiment of the present invention, the S3 includes:

[0036] Using the random masking module to convert the bottle mouth image to be detected into two masked images with complementary masks, and the masking rate of the two masked images is 50%.

[0037] In an embodiment of the present invention, the S4 includes:

[0038] S4.1: Input the two masked images into the trained autoencoder network respectively, and perform masked part reconstruction on each masked image in the two masked images, and output two masked reconstructed images;

[0039] S4.2: Combine the two masked reconstructed images to obtain a complete reconstructed image of the bottle mouth image to be detected.

[0040] In an embodiment of the present invention, the S5 includes:

[0041] S5.1: Calculate the absolute value of the pixel difference between the corresponding pixel points of the bottle mouth image to be detected and the reconstructed image;

[0042] S5.2: Add up all the absolute values of the pixel differences in the image and divide by the total number of pixel points, and use the calculation result as the anomaly score;

[0043] S5.3: Compare the anomaly score with a set threshold. When the anomaly score is greater than the threshold, it is determined that the bottle mouth image to be detected is a defective image.

[0044] Compared with the prior art, the beneficial effects of the present invention are:

[0045] 1. The present invention transforms the process of manually screening products with bottle mouth defects. By training a defect detector, a more intelligent robotic arm can be used to complete the business of screening products with bottle mouth defects, reducing labor costs and overcoming the problems of slow manual sorting speed and easy errors. By training a deep model with a large number of easily collectible normal bottle mouth images, the process of requiring a large amount of manual data annotation is eliminated, reducing costs. The present invention can also timely perceive its abnormal state and remind relevant staff to take a series of more accurate and forward-looking measures to ensure the quality and efficiency of products during the production process.

[0046] 2. The present invention is based on computer vision and uses image data for training and testing. Based on the semi-supervised learning method and the Transformer structure, the model is trained with 50% masked images, and the complete defect-free images are reconstructed through dual-stream masked images. The differences between the input images and the reconstructed images are analyzed to detect their abnormal states. This method has low cost, is simple and easy to implement, and realizes more efficient and accurate defect detection.

[0047] 3. Considering the deficiencies of expensive computing resources required for training based on generative adversarial networks and the insufficient representation ability of the autoencoder model based on convolutional neural networks, the present invention proposes an autoencoder network based on Transformer. The encoder only maps the visible blocks from the image space to the latent space, avoiding the process of the traditional generative adversarial network updating the search input vector to capture the normal sample manifold, reducing computing resources and strengthening the model representation ability. At the same time, the decoder mainly focuses on the reconstruction of the masked part, further reducing computing resources.

[0048] The following will further elaborate on the present invention in conjunction with the drawings and embodiments. Description of the Drawings

[0049] Figure 1 is a schematic flow chart of a method for detecting bottle mouth defects based on dual-stream semi-masked reconstruction provided by an embodiment of the present invention;

[0050] Figure 2 is a schematic block diagram of a bottle mouth defect detection system provided by an embodiment of the present invention;

[0051] Figure 3 is a schematic diagram of a processing system of a method for detecting bottle mouth defects based on dual-stream semi-masked reconstruction provided by an embodiment of the present invention;

[0052] Figure 4 is a schematic structural diagram of a Transformer structure unit provided by an embodiment of the present invention;

[0053] Figure 5 is a schematic diagram of the training process of a bottle mouth defect detection system provided by an embodiment of the present invention;

[0054] Figure 6 It is a schematic diagram of the test process of a bottle mouth defect detection system provided by an embodiment of the present invention;

[0055] Figure 7 It is an ROC curve diagram provided by an embodiment of the present invention;

[0056] Figure 8 It is a schematic diagram of some images in a training set and a test set provided by an embodiment of the present invention;

[0057] Figure 9 It is a bottle mouth defect detection result diagram obtained by using the method of an embodiment of the present invention. Detailed implementation manners

[0058] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the intended invention purpose, the following, in combination with the accompanying drawings and specific implementation manners, details a processing system of a bottle mouth defect detection method based on dual-stream semi-mask reconstruction proposed according to the present invention.

[0059] The foregoing and other technical contents, features and effects of the present invention can be clearly presented in the following detailed description in conjunction with the accompanying drawings. Through the description of the specific implementation manners, a more in-depth and specific understanding of the technical means and effects adopted by the present invention to achieve the intended purpose can be obtained. However, the accompanying drawings are only for reference and illustration, and are not used to limit the technical solution of the present invention.

[0060] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant is intended to cover non-exclusive inclusion, so that an article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the article or device including the said element.

[0061] Embodiment 1

[0062] Please refer to Figure 1 , Figure 1 It is a bottle mouth defect detection method based on dual-stream semi-mask reconstruction provided by an embodiment of the present invention. The bottle mouth defect detection includes:

[0063] S1: Construct a bottle mouth defect detection system, which includes a random masking module, an autoencoder network, and an anomaly scoring module connected in sequence. The autoencoder network further includes an encoder and a decoder, as Figure 2 and Figure 3 shown. First, use the random masking module to mask 50% of the original bottle mouth image. Add position encoding to the remaining visible blocks and perform patch embedding on them, then input them into the encoder, and the encoder performs feature extraction. Concatenate the encoder output and the masked part through position embedding, the decoder completes the reconstruction of the masked part, and the anomaly scoring module scores the difference between the input image and the reconstructed image. When performing defect detection, set a threshold artificially, and determine the image corresponding to the anomaly score higher than this threshold as a defective image. The embodiments of the present invention aim to be able to discover bottle mouth defects more efficiently and accurately. The embodiments of the present invention combine surface defect detection, anomaly detection, and semi-supervised learning techniques, use completely normal samples to train the model, and use samples with defects to test the ability of the model to detect defects.

[0064] Furthermore, the random masking module is specifically used for: adjusting the input bottle mouth image to a predetermined size, dividing the resized image into multiple regular and non-overlapping image patches and performing position encoding on each image patch; randomly masking the multiple image patches with a preset masking rate, and embedding the position encoding information of the visible image patches into the visible image patches. The specific patch embedding operation is implemented by performing convolution with a convolution kernel of size p and a stride of size p at the same time.

[0065] Furthermore, the autoencoder network includes an encoder and a decoder connected, where,

[0066] the encoder includes a plurality of first Transformer structural units connected in sequence, which are used to extract the feature information of the visible blocks; the decoder includes a plurality of second Transformer structural units connected in sequence, which are used to perform position embedding on the output of the encoder and the masked part of the input bottle mouth image, and complete the reconstruction of the masked part of the input bottle mouth image.

[0067] Furthermore, please refer to Figure 4 , Figure 4It is a schematic structural diagram of a Transformer structure unit provided by an embodiment of the present invention. The first Transformer structure unit and the second Transformer structure unit of this embodiment have the same structure, and respectively include a first normalization layer, a fully connected layer, a multi-head attention sub-unit, a first fusion sub-unit, a second normalization layer, a multi-layer perceptron, and a second fusion sub-unit. Among them, the first normalization layer is used to perform normalization processing on the input visible image patches, and the fully connected layer is used to perform a linear transformation on the normalized visible image patches, which is decomposed into three parts, namely query key value The multi-head attention sub-unit is used to process each decomposed part through multiple attention heads respectively, obtain the feature maps of each part and splice them to obtain the spliced feature map; the first fusion sub-unit is used to combine the spliced feature map with the input visible image patches; the second normalization layer is used to perform normalization processing on the combined feature map; the multi-layer perceptron includes two fully connected layers and a non-linear activation function arranged between the two fully connected layers, and is used to utilize the fully connected layer and the non-linear activation function to improve the ability of the encoder to fit non-linearity; the second fusion sub-unit is used to combine the output result of the first fusion sub-unit with the output result of the multi-layer perceptron and output a feature image.

[0068] S2: Input the training data set composed of defect-free bottle mouth images into the bottle mouth defect detection system to train the autoencoder network and obtain the trained bottle mouth defect detection system.

[0069] Specifically, please refer to Figure 5 , Figure 5 It is a schematic diagram of the training process of a bottle mouth defect detection system provided by an embodiment of the present invention. Step S2 of this embodiment includes:

[0070] S2.1: Obtain multiple defect-free bottle mouth images, adjust them to a predetermined size, and divide each resized image into multiple regular and non-overlapping image patches.

[0071] S2.2: Randomly mask the multiple image patches with a preset masking rate, and embed the position encoding information of the visible image patches into the visible image patches.

[0072] S2.3: Input the visible image patches embedded with position encoding information into the encoder, extract the feature information of the visible image patches, and then input the output of the encoder and the masked part of the current defect-free bottle mouth image into the decoder to reconstruct the masked part of the current defect-free bottle mouth image, and use the loss function to update the parameters in the encoder and the masker;

[0073] S2.4: When the value of the loss function tends to be stable, the training ends, and the trained encoder and decoder are obtained.

[0074] Specifically, in the training stage, first, a large number of defect-free images obtained are resized to 224×224 in size using a data processing method. Each defect-free image is segmented into 196 16×16 blocks. With a masking rate of 50%, 98 of the 16×16 blocks are randomly masked. Segmentation and masking operations are performed on each image, thus forming a training set. Each image in the training set includes visible blocks with position encoding and thus a masked part with position encoding. That is to say, an image is segmented into regular and non-overlapping blocks, and 50% of the blocks are masked by random sampling. After random masking, redundant information between adjacent blocks can be largely eliminated, improving the network's reconstruction ability for the learnable masked part.

[0075] Then, after position encoding of the visible block positions (recording the position of each block in the original image), block embedding is performed, which serves as the input to the encoder. The encoder uses the Vision Transformer structure to map the visible blocks into the latent space. The output of the encoder and the masked part of the current image are input into the decoder after position embedding. The decoder mainly reconstructs the masked part, ultimately prompting the decoder to reconstruct the masked blocks to be infinitely close to the corresponding parts of the input image.

[0076] As Figure 4 shown, the encoder of this embodiment takes the visible block feature information with added position embedding, first performs normalization processing, that is, for a sample, calculates the mean and variance of all the feature maps of the sample, and finally normalizes the sample. Then, based on the multi-head self-attention mechanism, the dependence relationship between two blocks in the input image is described. The input I in after linear transformation is decomposed into three parts, namely query key value In the multi-head self-attention mechanism, Q, K, and V use different linear projections, respectively:

[0077] Q i =QW i Q , K i =KW i K , V i =VW i V (i = 1, 2...h)

[0078] where, W i Q , W iK , W i V are the weight parameter matrices for the three parts of the linear projection, and h is the number of attention heads. After the attention operation, we have:

[0079]

[0080] where d k represents the dimension of Q and K. Let H i = Attention(Q i , K i , V i ). Concatenate each attention head to get the final output:

[0081] MultiHead(Q, K, V) = Concat(H1, H2,... H h )W O

[0082] Subsequently, combine the feature map obtained after multi-head self-attention processing with the feature map obtained from patch embedding, and then perform normalization. Then, use a multi-layer perceptron, which consists of two fully connected layers and a GELU activation function. The input feature layer passes through the first fully connected layer, and the number of channels is expanded to 4 times the original. After calculation by the GELU activation function, the second fully connected layer restores the number of channels to the original number. Therefore, the output feature dimension is the same as the input feature dimension. Finally, combine the feature map obtained by combining multi-head self-attention processing and patch embedding with the feature map processed by the multi-layer perceptron as the final output. Further, the decoder performs a position embedding operation on the output of the encoder and the masked part of the current defect-free bottle mouth image, and also completes the reconstruction of the masked part through the Transformer blocks structure.

[0083] In this embodiment, the difference between the masked block reconstructed by the autoencoder and the corresponding block in the original image is used as the loss function to train the model. The L2 loss is used, and the loss function is:

[0084] loss = ∑(loss' × mask) / ∑mask

[0085] loss' = ||I in - I rec ||2

[0086] where I rec is the reconstructed image output by the decoder, I in is the image input to the encoder, and mask is the masked part of the input defect-free bottle mouth image.

[0087] Based on the above training method, the optimal reconstruction model is saved.

[0088] In this embodiment, after the training phase, it further includes a testing process for the trained bottle mouth defect detection system. Specifically, please refer to Figure 6 , Figure 6 which is a schematic diagram of the testing process of a bottle mouth defect detection system provided by an embodiment of the present invention. Since a masking rate of 50% is adopted, theoretically, when the number of training times is sufficient, the network can complete each 50% masking training process. Therefore, by using the optimal reconstruction model saved in the training phase with dual-stream masks (the two parts of the masks are complementary) as inputs, the complementary parts can be reconstructed respectively. Combining the reconstructions of the two parts together is the complete reconstruction of the original input image. The specific process is as follows in steps S3 and S4. For defect images that have not been seen before, the reconstructed images are more inclined to normal images, so there will be significant differences. Thus, the anomaly score is obtained through the difference between the reconstructed image and the input image. By setting a threshold, when the anomaly score is greater than this threshold, it is determined as a defect image.

[0089] S3: Use the random masking module to convert the bottle mouth image to be detected into two masked images with complementary masks.

[0090] In this embodiment, the random masking module is used to convert the bottle mouth image to be detected into two masked images with complementary masks, and the masking rate of both of the two masked images is 50%.

[0091] S4: Input the two masked images into the trained autoencoder network respectively to obtain reconstructed images;

[0092] Specifically, S4 includes:

[0093] S4.1: Input the two masked images into the trained autoencoder network respectively, and perform masked part reconstruction on each masked image in the two masked images, and output two masked reconstructed images;

[0094] S4.2: Combine the two masked reconstructed images to obtain the complete reconstructed image of the bottle mouth image to be detected.

[0095] S5: Use the anomaly scoring module to compare the bottle mouth image to be detected with the reconstructed image to confirm whether there are bottle mouth defects in the bottle mouth image to be detected.

[0096] Further, S5 includes:

[0097] S5.1: Calculate the absolute value of the pixel difference between the corresponding pixel points of the bottle mouth image to be detected and the reconstructed image;

[0098] S5.2: Add up the absolute values of all pixel differences in the image and then divide by the total number of pixels. Take the calculation result as the anomaly score.

[0099] S5.3: Compare the anomaly score with a set threshold. When the anomaly score is greater than the threshold, determine that the image of the bottle mouth to be detected is a defective image.

[0100] Specifically, assume that the trained autoencoder network already has the ability to reconstruct the masked patch well enough. By simple consideration, take the absolute value of the per-pixel difference between the query image and the reconstructed image as the anomaly score. Set the threshold according to the training situation. When the anomaly score is greater than this threshold, it is determined as an anomaly. The anomaly score is defined as:

[0101]

[0102] where and are the i-th query image and reconstructed image respectively, and |.| is the absolute value operation.

[0103] The anomaly score of each query image can be calculated through the definition of the anomaly score. The anomaly scores of all query images form an anomaly score set S. Limit the anomaly score to [0, 1] through feature scaling. The final anomaly score can be expressed as:

[0104]

[0105] where S max and S min represent the maximum and minimum values in the anomaly score set S respectively.

[0106] It should be noted that the surface defect detection task can be regarded as a binary classification problem. It can be divided into four cases according to the combination of the true class of the sample and the predicted class of the network, namely True Positive (TP), False Positive (FP), True Negative (TN), and False Negative (FN). The confusion matrix of the classification results is shown in the table.

[0107] Table 1. Confusion matrix of classification results

[0108]

[0109]

[0110] Among them, TP means the sample is positive and the prediction result is positive; FP means the sample is negative and the prediction result is positive; FN means the sample is positive and the prediction result is negative; TN means the sample is negative and the prediction result is negative.

[0111] AUC is very effective in evaluating the detection performance of binary models and has been widely used by researchers. Therefore, it can be used to evaluate the performance of the proposed defect detection method. The evaluation is completed using the AUC value, that is, the area under the Receiver Operating Characteristic (ROC) curve. As Figure 7 shown, it is a ROC curve graph. The horizontal axis of this curve is the True Positive Rate (TPR), and the vertical axis is the False Positive Rate (FPR). The definitions of the two are as follows:

[0112]

[0113]

[0114] According to the definition of the AUC value, the AUC value can be obtained by summing the areas of each part under the ROC curve. Assuming that the ROC curve is formed by connecting the points with coordinates {(x1,y1),(x2,y2),...,(x m ,y m )} in sequence, then the AUC value can be estimated as:

[0115]

[0116] The calculation of AUC requires two parts, one is the label and the other is the score. The larger the AUC value, the better the defect detection performance. Each query image has a pixel-level label. During the test process, the anomaly score of each obtained query image can be used as the score, and the AUC value of the bottle mouth defect reaches 98.5%.

[0117] The embodiment of the present invention changes the process of traditional manual screening of bottle mouth defective products. By training a defect detector, a more intelligent robotic arm can be used to complete the business of screening bottle mouth defective products, reducing labor costs and overcoming the problems of slow manual sorting speed and easy errors. By training a deep model for a large number of easily collected normal bottle mouth images, the process of requiring a large amount of manual data annotation is eliminated, reducing costs. The present invention can also timely sense its abnormal state and remind relevant staff to take a series of more precise and forward-looking measures to ensure the quality and efficiency of products during the production process.

[0118] Embodiments of the present invention are based on computer vision and use image data for training and testing. Based on the semi-supervised learning method and the Transformer structure, the model is trained with 50% masked images, and the complete defect-free image is reconstructed through the dual-stream masked images. The difference between the input image and the reconstructed image is analyzed to detect its abnormal state. This method has low cost, is simple and easy to implement, and achieves more efficient and accurate defect detection.

[0119] Embodiment 2

[0120] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following further describes the embodiments of the present invention in detail.

[0121] Refer to Figure 8 , this example uses the Bottle datasets dataset collected by MVTec Company in Germany under real industrial scenarios. Its training set contains 209 defect-free images, and the test set contains 20 defect-free images and 124 defective images. There are 3 types of defects, namely broken_large (images with large-degree damage), broken_small (images with small-degree damage), and contamination (contaminated images), and all have pixel-level labels.

[0122] Refer to Figure 3 , the training steps are as follows:

[0123] Step 1, training stage. First, use the data processing method to adjust the size of the defect-free images to 224×224, divide each defect-free image into 196 16×16 blocks, and use a 50% masking rate to randomly mask 98 of the 16×16 blocks. Then, after encoding the positions of the visible blocks (recording the position of each block in the original image), perform block embedding as the input to the encoder. The encoder uses the Vision Transformer structure to map the visible blocks to the latent space. The output of the encoder and the masked part are input to the decoder after position embedding. The decoder mainly reconstructs the masked part, and finally prompts the decoder to reconstruct the masked blocks to be infinitely close to the corresponding parts of the input image, and save the optimal reconstruction model.

[0124] Step 2, refer to Figure 6, during the testing phase, since a masking rate of 50% is adopted, in theory, when the number of training times is sufficient, the network can complete each 50% masking training process. Therefore, by using the optimal reconstruction model saved in the training phase with a two-stream mask (the two parts of the mask are complementary) as input, the complementary parts can be reconstructed respectively. Combining the reconstructions of the two parts together is the complete reconstruction of the original input image. For unseen defective images, more normal-looking images are reconstructed, so there will be significant differences. Thus, the anomaly score is obtained from the difference between the reconstructed image and the input image. By setting a threshold, when the anomaly score is greater than this threshold, it is determined as a defective image.

[0125] Step 3, defect detection. Assume that the trained autoencoder network already has the ability to reconstruct the masked patch excellently. By simple consideration, the absolute value of the per-pixel difference between the query image and the reconstructed image is used as the anomaly score. According to the training situation, set the threshold. When the anomaly score is greater than this threshold, it is determined as an anomaly.

[0126] The experimental results of the embodiments of the present invention are as shown in the attached drawings. Figure 9 As shown, in the test, a contaminated defective image is used. Through the method of the embodiments of the present invention, the defective image can be well reconstructed into a defect-free image. After taking the residual between the defective image and the reconstructed image, the obtained residual image is compared with the pixel-level label, and the defect can be clearly found visually.

[0127] The embodiments of the present invention consider the deficiencies of the high computational resource cost of training based on generative adversarial networks and the insufficient representation ability of the autoencoder model based on convolutional neural networks, and propose an autoencoder network based on Transformer. The encoder only maps the visible patches from the image space to the latent space, avoiding the process of the traditional generative adversarial network updating the search input vector to capture the normal sample manifold, reducing the computational resources and strengthening the model representation ability. At the same time, the decoder mainly focuses on the reconstruction of the masked part, further reducing the computational resources. The embodiments of the present invention transform the bottle mouth defect detection from the traditional manual screening mode to the big data behavior mode analysis. Based on semi-supervised learning, a semi-masking method is adopted, randomly masking 50% of the normal images, and introducing the Vision Transfomer (ViT) structure into the autoencoder. The encoder is used to extract the features of the visible part, and the decoder is used to reconstruct the masked part into a normal image. By combining the reconstructed images of the complementary masks together to form a complete reconstructed image, the difference between the input image and the reconstructed image is analyzed to detect its abnormal state. The method of the present invention has low cost, is simple and easy to implement, and has low computational complexity.

[0128] In several embodiments provided by the present invention, it should be understood that the devices and methods disclosed in the present invention can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed.

[0129] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module. The above integrated modules can be implemented in the form of hardware, or in the form of a combination of hardware and software functional modules.

[0130] Another embodiment of the present invention provides a storage medium, in which a computer program is stored, and the computer program is used to execute the steps of the bottle mouth defect detection method based on double-stream semi-mask reconstruction in the above embodiments. Another aspect of the present invention provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and when the processor calls the computer program in the memory, the steps of the bottle mouth defect detection method based on double-stream semi-mask reconstruction as described in the above embodiments are implemented. Specifically, the above integrated modules implemented in the form of software functional modules can be stored in a computer-readable storage medium. The above software functional modules are stored in a storage medium, including several instructions for causing an electronic device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute some steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0131] The above content is a further detailed description of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention belongs, without departing from the concept of the present invention, several simple deductions or substitutions can be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. A method for detecting bottle mouth defects based on dual-stream semi-mask reconstruction, characterized in that, Including: S1: Construct a bottle mouth defect detection system, which includes a random masking module, an autoencoder network, and an anomaly scoring module connected in sequence; S2: Input a training data set composed of defect-free bottle mouth images into the bottle mouth defect detection system to train the autoencoder network and obtain a trained bottle mouth defect detection system; S3: Use the random masking module to convert the bottle mouth image to be detected into two masked images with complementary masks; S4: Input the two masked images into the trained autoencoder network respectively to obtain reconstructed images; S5: Use the anomaly scoring module to compare the bottle mouth image to be detected with the reconstructed image to confirm whether there are bottle mouth defects in the bottle mouth image to be detected.

2. The method for detecting bottle mouth defects based on dual-stream semi-mask reconstruction according to claim 1, characterized in that, The random masking module is specifically used for: Adjust the input bottle mouth image to a predetermined size, divide the resized image into regular and non-overlapping image patches, and perform position encoding on each image patch; Randomly mask the multiple image patches with a preset masking rate, and embed the position encoding information of the visible image patches into the visible image patches.

3. The method for detecting bottle mouth defects based on dual-stream semi-mask reconstruction according to claim 1, characterized in that, The autoencoder network includes an encoder and a decoder connected. Among them, The encoder includes a plurality of first Transformer structural units connected in sequence, which are used to extract the feature information of the visible patches; The decoder includes a plurality of second Transformer structural units connected in sequence, which are used to perform position embedding on the output of the encoder and the masked part of the input bottle mouth image, and complete the reconstruction of the masked part of the input bottle mouth image.

4. The method for detecting bottle mouth defects based on dual-stream semi-mask reconstruction according to claim 3, characterized in that, The first Transformer structural unit and the second Transformer structural unit have the same structure, and each includes a first normalization layer, a fully connected layer, a multi-head attention sub-unit, a first fusion sub-unit, a second normalization layer, a multi-layer perceptron, and a second fusion sub-unit. Among them, The first normalization layer is used to normalize the input visible image patches; The fully connected layer is used to perform a linear transformation on the normalized visible image patches, and decompose them into three parts, namely query Q, key K, and value V; The multi-head attention sub-unit is used to process each decomposed part through multiple attention heads respectively, obtain the feature maps of each part and splice them to obtain a spliced feature map; The first fusion sub-unit is used to combine the spliced feature map with the input visible image patches; The second normalization layer is used to normalize the combined feature map; The multi-layer perceptron includes two fully connected layers and a non-linear activation function arranged between the two fully connected layers, which is used to improve the encoder's ability to fit non-linearity by using the fully connected layer and the non-linear activation function; The second fusion sub-unit is used to combine the output result of the first fusion sub-unit with the output result of the multi-layer perceptron and output a feature image.

5. The method for detecting bottle mouth defects based on dual-stream semi-mask reconstruction according to claim 1, characterized in that, The S2 includes: S2.1: Obtain multiple defect-free bottle mouth images, adjust them to a predetermined size, and divide each resized image into regular and non-overlapping image patches; S2.2: Randomly mask the multiple image patches using a preset masking rate, and embed the position encoding information of the visible image patches into the visible image patches; S2.3: Input the visible image patches embedded with position encoding information into the encoder to extract the feature information of the visible image patches. Subsequently, input the output of the encoder and the masked part of the current defect-free bottle mouth image into the decoder together to reconstruct the masked part of the current defect-free bottle mouth image, and use the loss function to update the parameters in the encoder and the masker; S2.4: When the value of the loss function tends to be stable, the training ends, and the trained encoder and decoder are obtained.

6. The method for detecting bottle mouth defects based on dual-stream semi-mask reconstruction according to claim 5, characterized in that, The loss function is: loss = ∑(loss' × mask) / ∑mask loss'=||I in -I rec ||2 where I rec is the reconstructed image output by the decoder, and I in is the image input to the encoder, and mask is the masked part of the input defect-free bottle mouth image.

7. The bottle mouth defect detection method based on dual-stream semi-mask reconstruction according to claim 1, characterized in that The S3 includes: Use the random masking module to convert the bottle mouth image to be detected into two masked images with complementary masks, and the masking rate of the two masked images is 50%.

8. The bottle mouth defect detection method based on dual-stream semi-mask reconstruction according to claim 4, characterized in that The S4 includes: S4.1: Input the two masked images into the trained autoencoder network respectively, reconstruct the masked part of each masked image in the two masked images, and output two masked reconstructed images; S4.2: Combine the two masked reconstructed images to obtain a complete reconstructed image of the bottle mouth image to be detected.

9. The bottle mouth defect detection method based on dual-stream semi-mask reconstruction according to any one of claims 1 to 8, characterized in that The S5 includes: S5.1: Calculate the absolute value of the pixel difference between the corresponding pixel points of the bottle mouth image to be detected and the reconstructed image; S5.2: Add up all the absolute values of the pixel differences in the image and divide by the total number of pixel points, and use the calculation result as the anomaly score; S5.3: Compare the anomaly score with a set threshold. When the anomaly score is greater than the threshold, it is determined that the bottle mouth image to be detected is a defective image.

Citation Information

Patent Citations

  • Industrial CT defect detection method based on deep learning

    CN111179229A

  • Image defect detection method and device, electronic equipment, storage medium and product

    CN112581463A