An aluminum foil sealing tightness detection method based on unsupervised learning

By employing unsupervised learning methods, combined with infrared image processing and multi-scale image reconstruction networks, the problems of low efficiency and poor accuracy in existing aluminum foil sealing detection are solved, achieving non-destructive and efficient sealing performance detection and defect localization.

CN116934725BActive Publication Date: 2025-11-28HOHAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310937867.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-28
Publication Date
2025-11-28
Estimated Expiration
2043-07-28

AI Technical Summary

Technical Problem

Existing methods for detecting the sealing performance of aluminum foil rely on manual sampling and are inefficient. Traditional machine vision methods require extensive prior knowledge and have poor generalization ability. Deep learning-based methods have high manual annotation costs, few defect samples, and a high degree of uncertainty regarding defects.

Method used

An unsupervised learning-based approach is adopted to acquire sealing images through an infrared camera, and adaptive ROI region extraction, grayscale processing, edge detection, and image enhancement are performed. Defect detection is carried out by combining a multi-scale image reconstruction network model, and multi-semantic features are extracted and defects are suppressed by using Transformer and CNN structures to achieve non-destructive detection.

Benefits of technology

It achieves efficient aluminum foil sealing performance testing without the need for defect samples and manual labeling, improving testing accuracy and efficiency. It enables fine-grained detection and precise positioning of sealing defects, and is suitable for real-time monitoring in bottling production lines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116934725B_ABST
    Figure CN116934725B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on the detection method of unattended learning of aluminum foil seal tightness. Including: S1, obtain aluminum foil seal infrared image;S2, to aluminum foil seal infrared image is carried out adaptive ROI region identification and extraction;S3, to the thermal imaging image of extraction is preprocessed and normal sample image is made into data set;S4, after pre-processing normal sample data set is input into multi-scale image reconstruction network model and is carried out model parameter training;S5, after pre-processing the image to be measured is input into the reconstruction model of training and is carried out image reconstruction, and output reconstruction image, and the residual error of the image to be measured and reconstruction image is calculated, and residual error map is obtained;S6, to residual error map is preprocessed, highlights defect part, whether infrared seal image exists defect is judged by abnormal score, if there is defect, the defect part is positioned.The application can be reconstructed and repaired to aluminum foil seal infrared image quickly and effectively, with higher detection precision.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of detection, and particularly relates to a detection method for the sealing property of an aluminum foil seal based on unsupervised learning. BACKGROUND

[0002] At present, bottle packaging, as one of the mainstream products in the packaging industry, is widely used in the food, medicine, cosmetic and other industries, and the sealing property of the bottle provided by the aluminum foil sealing technology is an important factor to ensure product quality. In the process of aluminum foil sealing, factors such as temperature, aluminum foil quality and cap tightness can affect the sealing property of the aluminum foil sealing, and thus affect the quality and effectiveness of the product. Therefore, it is necessary to detect the sealing property of the aluminum foil sealing.

[0003] The existing sealing detection method mainly relies on manual sampling investigation, such as water pressure method, air pressure method and the like, but the manual cost is high and the efficiency is low, and the sample to be detected is also easily damaged, which cannot meet mass production. At present, a method of judging whether the seal is complete by using machine vision has appeared on the market, which captures the temperature distribution map of the seal by an infrared camera to detect defects. However, the traditional machine vision method needs rich prior knowledge and has poor generalization, and needs to be adjusted by various parameters to adapt to different packaging materials.

[0004] In recent years, deep learning has been widely used in industrial detection due to its strong learning ability. Most of the networks are based on supervised learning, that is, a large number of normal samples and defect samples with artificial annotation are needed. However, as the precision of the packaging production line is continuously improved, the difficulty of obtaining the seal defect sample is suddenly increased. Therefore, the high cost of artificial annotation, the lack of defect samples and the large unknown nature of defects have become the main problems of the current aluminum foil seal defect detection based on deep learning. SUMMARY

[0005] In order to solve the above technical problems, the application provides a detection method for the sealing property of an aluminum foil seal based on unsupervised learning, which comprises the following steps:

[0006] Step S1, acquiring an infrared image of the aluminum foil seal;

[0007] Step S2, performing adaptive ROI region extraction on the infrared image of the aluminum foil seal;

[0008] Step S21, performing three-channel weighted average on the acquired infrared image to realize gray scale processing;

[0009] Step S22, performing smoothing and noise reduction processing on the gray scale image to remove small textures and irregular noise points, and the smoothing and noise reduction processing method is Gaussian filtering, and the expression is:

[0010]

[0011] where (x, y) represents the pixel coordinates in the image, G(x, y) represents the value of the Gaussian function calculated at this position, σ represents the standard of the Gaussian kernel, and e is the constant base number of the natural logarithmic function;

[0012] Step S23, the edge profile of the gray image after smoothing is extracted by using Canny edge detection, and a binary edge image is obtained;

[0013] Step S24, a contour extraction algorithm is used to traverse the pixel value of the binary image, extract all the completely closed edge profiles in the edge image, and use the maximum wheel contour, i.e. the circumscribed rectangle to mark the target region. According to the position and size of the circumscribed rectangle, the corresponding region is cropped from the original image, and the obtained corresponding region is the ROI region of the aluminum foil sealing infrared image.

[0014] Step S3, the extracted infrared image is preprocessed and the normal sample image is made into a data set;

[0015] Step S31, the ROI image is subjected to image enhancement. First, the contrast of the image is enhanced to improve the visibility of image details. Specifically, adaptive histogram equalization is used, and the formula is as follows:

[0016]

[0017]

[0018] where f(x, y) is the original image, (x, y) is the pixel point coordinate, h(i, j) is the window function, and W is the window size; T(x, y) refers to the pixel value with (x, y) as the center, f(x+i, y+j) refers to the pixel value with (x+i, y+j) as the center in the original image, and i and j refer to the coordinates of the window function;

[0019] Then, a Laplace filter is used to sharpen the image and enhance the edges and details of the image;

[0020] Step S32, the sample after image enhancement is subjected to interpolation upsampling operation to improve the resolution of the infrared image and provide more image information for subsequent deep learning model training. Specifically, the bicubic interpolation method is used, and the formula is as follows:

[0021]

[0022]

[0023] where (b x ,b y ) is the coordinate of the interpolation point, B(b x ,b y) is the bicubic interpolation result, (b xi yj ) is the 16 nearest neighbors of the interpolation point, W(x) is the weighting function, and a is-0.5 here.

[0024] Step S33, the image is randomly rotated to expand the training data set and increase the robustness of the model;

[0025] Step S34, the image pixel value is normalized to 0 to 1, and the normalized image pixel value is standardized to have a mean of 0 and a variance of 1 to speed up the model convergence, and the formula is:

[0026]

[0027]

[0028] where x norm is the normalized result, x std is the standardized result, x org is the pixel value of the original image, and μ and σ are the mean and standard deviation of the single channel pixel value.

[0029] Step S35, the preprocessed image data set is divided into training and test sets, and the normal sample data is used as the training set accounting for 75%, and the remaining 25% of the normal sample and the defect sample are used as the test set.

[0030] Step S4, the preprocessed normal sample data set is input into the multi-scale image reconstruction network model for model parameter training.

[0031] The multi-scale image reconstruction network model includes an image generation network and an image discrimination network, wherein the image generation network includes a multi-scale feature sampling module, a global context feature extraction module, an abnormal feature detection module, and an image generation module.

[0032] The multi-scale feature sampling module is composed of four convolution groups. The first layer convolution group is composed of a convolution layer with a convolution kernel of 7*7 and a maximum pooling layer, which aims to capture large-scale features and reduce the computational complexity of subsequent layers. The subsequent three convolution groups are composed of residual block structures, including a convolution layer with a 3*3 convolution kernel, batch normalization, an activation function, and a residual connection, which are used to better extract local features such as texture and edge, while avoiding gradient disappearance. The multi-scale feature sampling module outputs three different scale feature maps, which are output by the second, third, and fourth convolution groups, respectively. The multi-scale feature improves the generalization ability and robustness of the model and is beneficial to the reconstruction of image details by the image reconstruction network.

[0033] ​The global context feature extraction module fuses the Transformer structure and the convolution structure, and specifically includes a 3*3 convolution layer based on down-sampling and an improved lightweight Vision Transformer module, the improved lightweight Vision Transformer module is in sequence relative position coding, a local perception unit, a LayerNorm layer, a lightweight multi-head self-attention module, a LayerNorm layer and an improved MLP module;

[0034] Wherein, the local perception unit adopts a depth separable convolution, and the translational invariance of the convolution is introduced into the module, and a specific formula is as follows:

[0035] LPU(X)=DWConv(X)+X

[0036] In the formula, X is an input feature tensor, DWConv is a depth separable convolution layer, and the whole adopts a residual connection;

[0037] The lightweight multi-head self-attention module simplifies the generation of Key and Value through convolution operation on the basis of the original multi-self-attention, and the calculation formula of attention is as follows:

[0038]

[0039]

[0040]

[0041] In the formula, Q, K and V are respectively Query, Key and Value in the Transformer, Softmax is a normalized exponential function, K T is a transpose matrix of K, R is a real field, B is a bias matrix, and k is a multiple of K and V reduced along the spatial direction;

[0042] The improved MLP module adds a 3*3 convolution layer between the original fully connected layers, and through the residual connection, the gradient vanishing is avoided, and the ability of the Transformer module to extract local semantic information is enhanced;

[0043] The global context feature extraction module has three branches, the input feature is three different scale features output by the multi-scale feature sampling module, after feature extraction through the three branch structure networks, the features are fused and output.

[0044] The abnormal feature detection module is composed of K-means clustering and feature detection, and is mainly used for suppressing defect features in the image reconstruction process. The feature map output from the global context feature extraction module is dimensionally decomposed to contain N feature vectors P, P={p1, p2,..., pN}(P∈R C×1 , N=H*W), and is input into K-means clustering. K midpoints are selected as clustering centers, and the center vector is C, C={C1, C2,..., Ck}(C∈R C×1 ), and feature detection is to replace defect features in the feature vector that are too far away from the center vector with normal feature vectors;

[0045] The main part of the image generation module and the image discrimination network adopts the network structure of DCGAN. The image generation module is composed of five up-sampling modules. The normalization operation in the module adopts Instance Norm. The image generation module outputs a three-channel color reconstruction image, and the image size is consistent with the input image. The image discrimination network includes five feature down-sampling modules and one fully connected layer. A self-attention module is added between the fourth and fifth layers of the image generation module and between the first and second layers of the image discrimination network. The output of the image discrimination network is the probability that the input image is true, ranging from 0 to 1.

[0046] Step S41, random color Gaussian noise is added to the images in the training set, and a mask operation is performed;

[0047] Step S42, the abnormal feature detection module in the image reconstruction model is excluded, and the processed image is input into the model for parameter training. The adversarial loss is used as the training loss function of the image discrimination network, and the image reconstruction loss is used as the training loss function of the image generation network.

[0048] Step S43, the trained network weight is fixed, the abnormal feature detection module is added, and the K-means clustering is parameter trained: the feature vectors C of K clustering centers are randomly initialized, the distance between each input feature vector P and each clustering center is calculated, and the feature vector is assigned to the nearest clustering center. In each cluster, the mean of all samples in the cluster is calculated, and the mean is used as the new class center. The above steps are iterated to complete the training of K-means clustering.

[0049] Step S5, the preprocessed to-be-tested image is input into the trained reconstruction model for image reconstruction, and a reconstructed image is output. The to-be-tested image and the reconstructed image are used for residual calculation to obtain a residual image.

[0050] Step S51, add random color Gaussian noise to the pre-processed image to be tested, input the image to the image reconstruction model with fixed training weight for image reconstruction, wherein when passing through the abnormal feature detection module, the spatial distance d of the input feature vector P to all cluster centers C is calculated, if the spatial distance d exceeds the abnormal threshold T, the feature vector is regarded as a defect feature, and the nearest center feature vector is used to replace the defect feature vector, so as to achieve the effect of defect feature suppression, and the calculation formula of the spatial distance d is:

[0051]

[0052] In the formula, X and Y are the input feature vector P and the center feature vector C respectively, and d(X, Y) is the Euclidean distance between the two vectors;

[0053] The abnormal threshold T is obtained by training the feature vector of the positive sample, and the calculation formula is:

[0054]

[0055] In the formula, di is the spatial distance of the feature vector from the center feature vector during training of the positive sample, N is the number of feature vectors, and σ d is the standard deviation of the distance d i ;

[0056] Step S52, calculate the pixel residual error between the reconstructed image output by the model and the image to be tested to obtain a preliminary residual error image, and the residual error calculation formula is as follows:

[0057] L dif (i,j)=(L src (i,j)-L rec (i,j)) 2

[0058]

[0059] In the formula, L src (i,j) is the image to be detected, L rec (i,j) is the reconstructed image, and the calculated residual error image is further normalized to obtain the final result.

[0060] Step S6, pre-process the residual error image to highlight the defect part, judge whether the infrared sealing image has defects through the abnormal score, and if there are defects, position the defect part.

[0061] Step S61, perform three-channel weighted averaging on the residual error image to realize grayscale processing;

[0062] Step S62, adopt mean filtering to perform image denoising processing to eliminate pseudo-defects formed by noise points;

[0063] Step S63, the abnormal probability score of the processed gray scale image is calculated, and when the abnormal probability score is greater than the abnormal threshold, it indicates that there is an abnormal point in the residual image, and the formula is:

[0064]

[0065] In the formula, S map is a single-channel residual image matrix, S mapmax is the maximum value in the pixel matrix, S mapmin is the minimum value in the pixel matrix, and S is the maximum value of the normalized residual image matrix.

[0066] Step S64, the adaptive threshold method is used for binaryzation processing of the residual image, the threshold value of image segmentation is calculated by using the Otsu method, and the gray scale image is converted into a binary image, and the formula is as follows:

[0067]

[0068] In the formula, T OTSU is the threshold value obtained by the Otsu method, and t is the pixel value in the gray scale image.

[0069] Step S65, according to the obtained binary image, the proportion R of the pixel points with a pixel value of 1 in the whole image is calculated, if the abnormal probability score is greater than the abnormal threshold and the proportion of the defect pixels is greater than the proportion threshold, it is determined that the image has defects, otherwise, the image is normal.

[0070] Step S66, if the image has defects, the contour extraction algorithm is used to traverse the pixel value of the binary image, all the completely closed edge contours in the edge image are extracted, and the maximum bounding box (circumscribed rectangle) is used to mark the target area, so that the defect positioning is realized.

[0071] Compared with the prior art, the beneficial effects of the present application are that the aluminum foil sealing detection method based on unsupervised learning of the present application uses the infrared image of the sealing captured by the infrared camera to perform nondestructive defect detection on the aluminum foil sealing. The method adopts a neural network based on unsupervised learning, does not need defect samples and related manual labeling, greatly saves the preparation time and labor cost; the Transformer and CNN structure are combined to extract multi-semantics features from the global and local and suppress the defects, so that the effect of detail reconstruction is achieved, the fine-grained detection and accurate positioning of the defects are realized, the robustness is high, the accuracy and efficiency of the sealing defect detection are greatly improved, and the real-time monitoring of the sealing defects of the bottle production line is realized.

[0072] Other features and advantages of the present application will be described in detail in the specification, and some can be directly seen from the specification or specific embodiments.

[0073] In order to make the above invention purposes, advantages and characteristics more clear and intuitive, the following text has a good implementation case of the invention and related drawings are described in detail. BRIEF DESCRIPTION OF DRAWINGS

[0074] In order to more clearly illustrate the specific implementation cases of the present application or the existing technical solutions, the following will be related to the specific implementation of the present application. Embodiments are described with reference to the accompanying drawings.

[0075] Figure 1 For the specific embodiment process diagram of the aluminum foil sealing sealing detection method of the present application;

[0076] Figure 2 For the step diagram of the aluminum foil sealing sealing detection method of the present application;

[0077] Figure 3 For the multi-scale image reconstruction network model of the aluminum foil sealing sealing detection method of the present application;

[0078] Figure 4 For the residual module of the multi-scale image reconstruction network model of the present application;

[0079] Figure 5 For the lightweight Vision Transformer module of the multi-scale image reconstruction network model of the present application;

[0080] Figure 6 For the example diagram of a group of input images, reconstructed images, residual images and defect positioning images in the embodiment of the present application. DETAILED DESCRIPTION

[0081] In order to more clearly illustrate the purposes, technical solutions and advantages of each embodiment of the present application, the technical solutions of the present application will be described clearly and completely below with reference to the accompanying drawings. Obviously, the examples only represent some examples of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0082] Figure 2 is the step diagram of the aluminum foil sealing sealing detection method based on unsupervised learning provided by an embodiment of the present application, referring to Figure 1 The specific embodiment process diagram of the method is as follows:

[0083] Step S1, acquiring an aluminum foil sealing infrared image;

[0084] Step S2, performing adaptive ROI region extraction on the aluminum foil sealing infrared image;

[0085] Step S3, preprocessing the extracted infrared image and making normal sample images into a data set;

[0086] Step S4, input the preprocessed normal sample data set into the multi-scale image reconstruction network model for model parameter training;

[0087] Step S5, input the preprocessed to-be-detected image into the trained reconstruction model for image reconstruction, output the reconstructed image, and calculate the residual error between the to-be-detected image and the reconstructed image to obtain a residual error image;

[0088] Step S6, pre-process the residual error image to highlight the defect part, judge whether the infrared sealing image has defects through the abnormal score, and if there are defects, position the defect part.

[0089] In this embodiment, the sealing infrared image captured by the infrared camera is used for non-destructive defect detection of the aluminum foil seal. This method uses a neural network based on unsupervised learning, without the need for defect samples and related manual annotation, greatly saving preparation time and labor cost; combining Transformer and CNN structure, multi-semantic features are extracted from the global and local to suppress defects, achieving the effect of detail reconstruction, realizing fine-grained detection and accurate positioning of defects, having strong robustness, greatly improving the accuracy and efficiency of sealing defect detection, and realizing real-time monitoring of sealing defects on the bottle production line.

[0090] In the embodiment, the step S1 of obtaining the aluminum foil sealing infrared image can be achieved by installing an infrared camera 32mini with a resolution of 384*288 at a distance of 0.5m to 1m behind the electromagnetic aluminum foil sealing machine, cooperating with the photoelectric signal, and obtaining the infrared image of the bottle cap surface above the bottle. The infrared image is obtained by the sealing transient temperature conduction to the bottle cap.

[0091] The step S2 of adaptively extracting the ROI region of the aluminum foil sealing infrared image includes:

[0092] Step S21, performing three-channel weighted average on the obtained infrared image to realize gray-scale processing;

[0093] For example, the weighted average formula of the three channels can be:

[0094] Gray = 0.417B + 0.205G + 0.378R

[0095] In the formula, B, G, and R are the pixel values of the blue, green, and red channels, and the weights before the channels are calculated from the infrared image data set in this embodiment;

[0096] Step S22, performing smoothing and denoising processing on the gray-scale image to remove small textures and irregular noise points, and the smoothing and denoising processing method is Gaussian filtering, and its expression is:

[0097]

[0098] In the formula, (x, y) represents the pixel coordinates in the image, G(x, y) represents the value of the Gaussian function calculated at the position, sigma represents the standard of the Gaussian kernel, and e is the constant base number of the natural logarithmic function;

[0099] Step S23, the edge profile of the gray image after smoothing is extracted by using Canny edge detection, and a binary edge image is obtained;

[0100] Step S24, a contour extraction algorithm is used to traverse the pixel value of the binary image, extract all the completely closed edge profiles in the edge image, and use the maximum wheel outline, i.e. the circumscribed rectangle to mark the target region. According to the position and size of the circumscribed rectangle, the corresponding region is cropped from the original image, and the obtained corresponding region is the ROI region of the aluminum foil sealing infrared image.

[0101] The step S3 of preprocessing the extracted infrared image and making a normal sample image into a data set comprises:

[0102] Step S31, image enhancement is performed on the extracted ROI image. First, the contrast of the image is enhanced to improve the visibility of image details. Specifically, adaptive histogram equalization is used, and the formula is as follows:

[0103]

[0104]

[0105] In the formula, f(x, y) is the original image, (x, y) is the pixel point coordinate, h(i, j) is the window function, W is the window size; T(x, y) refers to the pixel value with (x, y) as the center, f(x+i, y+j) refers to the pixel value with (x+i, y+j) as the center in the original image, and i and j refer to the coordinates of the window function;

[0106] Then, a Laplace filter is used to sharpen the image and enhance the edges and details of the image;

[0107] Step S32, interpolation upsampling operation is performed on the sample after image enhancement to improve the resolution of the infrared image and provide more image information for subsequent deep learning model training. Specifically, a bicubic interpolation method is used, and the formula is as follows:

[0108]

[0109]

[0110] In the formula, (b x ,b y) is the coordinate of the interpolation point, B(b x ,b y ) is the bicubic interpolation result, (b xi ,b yj ) is the 16 nearest points of the interpolation point, and W(x) is the weighting function, where a is-0.5.

[0111] In the example, the size of the up-sampled image is 224x224;

[0112] In step S33, the image is randomly rotated to expand the training data set and increase the robustness of the model.

[0113] In the example, to ensure the integrity of the edges of the rotated image, the image is padded with background pixel values to a size of 256x256 before rotation, and after being randomly rotated by 10 degrees clockwise or counterclockwise, the center is cropped to a size of 224x224 to complete the data set expansion.

[0114] In step S34, the image pixel values are normalized to between 0 and 1, and the normalized image pixel values are standardized to have a mean of 0 and a variance of 1 to accelerate model convergence, and the formula is:

[0115]

[0116]

[0117] In the formula, x norm is the normalized result, x std is the standardized result, x org is the pixel value of the original image, and μ and σ are the mean and standard deviation of the single channel pixel value.

[0118] In step S35, the preprocessed image data set is divided into training and test sets, with normal sample data accounting for 75% as the training set, and the remaining 25% of normal samples and defect samples as the test set.

[0119] In this embodiment, the image sample data comes from a factory, a total of 612 images, of which 588 are normal sample images and 24 are artificially made defect images, with a resolution of 388x284. After image preprocessing, the size becomes 224x224, 460 normal samples are randomly selected as the training set, and the remaining 128 normal samples and 24 defect samples are used as the test set.

[0120] The multi-scale image reconstruction network model in step S4 is shown in Figure 3 , which includes an image generation network and an image discrimination network. The image generation network includes a multi-scale feature sampling module, a global context feature extraction module, an abnormal feature detection module, and an image generation module.

[0121] Specifically, the input image size is 224x224x3, and three different scale feature maps are output after the multi-scale feature sampling module, with sizes of 56x56x64, 28x28x128, and 14x14x256 respectively; the three feature maps are input into the global context feature extraction module for multi-branch feature extraction, the extracted feature map with a size of 7x7x512 is input into the abnormal feature detection module for defect feature suppression, and the output result is converted into a feature vector Z with a size of 1x1024, which is the normal feature of the input image; the Z is input into the image generation network to generate the final normal sample reconstruction result; the image discrimination network inputs the reconstruction result to determine whether the image is a normal image, which is used to enhance the detail reconstruction effect of the reconstruction network in the training process.

[0122] In the example, as shown in Figure 3 The multi-scale feature sampling module is composed of four convolution groups, the first layer convolution group is composed of a convolution layer with a convolution kernel of 7*7 and a maximum pooling layer, which aims to capture large-scale features and reduce the computational complexity of subsequent layers; the subsequent three convolution groups are composed of residual block structures, as shown in Figure 4 The residual block structure includes a convolution layer with a 3*3 convolution kernel, a convolution layer with a 1*1 convolution kernel, batch normalization, an activation function, and a residual connection, which is used to better extract local features such as texture and edge, while avoiding gradient disappearance; the multi-scale feature sampling module outputs three different scale feature maps, which are output by the second, third, and fourth convolution groups respectively, and the multi-scale feature improves the generalization ability and robustness of the model and is beneficial to the reconstruction of image details by the image reconstruction network;

[0123] The global context feature extraction module combines the Transformer structure and the convolution structure, and specifically includes a 3*3 convolution layer based on downsampling and an improved lightweight Vision Transformer module, the improved lightweight Vision Transformer module sequentially includes relative position encoding, a local perception unit, a LayerNorm layer, a lightweight multi-head self-attention module, a LayerNorm layer, and an improved MLP module, Figure 5 The lightweight Vision Transformer module used in the example is shown below:

[0124] The local perception unit adopts a depthwise separable convolution, which introduces the translational invariance of convolution into the module, and the specific formula is:

[0125] LPU(X)=DWConv(X)+X

[0126] In the formula, X is an input feature tensor, DWConv is a depthwise separable convolution layer, and the whole adopts a residual connection.

[0127] In this example, first, 3*3 convolution is used for channel-wise convolution, and then 1*1 convolution is used for point-wise convolution to realize the role of depth separable convolution;

[0128] The lightweight multi-head self-attention module simplifies the generation of Key and Value through convolution operation on the basis of the original multi-self-attention, greatly saving the calculation amount, and the calculation formula of attention is as follows:

[0129]

[0130]

[0131]

[0132] In the formula, Q, K and V are Query, Key and Value in the Transformer respectively, Softmax is a normalized exponential function, K T is the transpose matrix of K, R is a real field, B is a bias matrix, and k is the multiple of K and V reduced along the spatial direction;

[0133] Specifically, when Q, K and V are calculated, 1*1 convolution is used instead of the weight matrix of W q , W k and W v , to speed up the training and inference of the model, and nonlinear transformation is introduced to enhance the expression ability of the model;

[0134] The improved MLP module adds a 3*3 convolution layer between the original fully connected layers, and through residual connection, it avoids gradient disappearance while enhancing the ability of the Transformer module to extract local semantic information;

[0135] The global context feature extraction module has three branches, and the input features are three different scale features output by the multi-scale feature sampling module. After feature extraction by the three branch structure network, the features are fused and output.

[0136] The abnormal feature detection module is composed of K-means clustering and feature detection, and is mainly used to suppress the defect features in the image reconstruction process. The feature map output from the global context feature extraction module is decomposed along the dimension, containing N feature vectors P, P={p1, p2,..., pN}(P∈R C×1 , N=H*W), and K points are selected as the clustering centers in the K-means clustering, and the center vector is C, C={C1, C2,..., Ck}(C∈R C×1), the feature detection is to replace the defect features in the feature vector that are too far from the center vector with normal feature vectors.

[0137] The main part of the image generation module and the image discrimination network adopts the network structure of DCGAN. The image generation module is composed of five up-sampling modules. The normalization operation in the module adopts Instance Norm. The image generation module outputs a three-channel color reconstruction image. The image size is consistent with the input image. The image discrimination network includes five feature down-sampling modules and one fully connected layer. A self-attention module is added between the fourth and fifth layers of the image generation module and between the first and second layers of the image discrimination network. The output of the image discrimination network is the probability that the input image is true, ranging from 0 to 1. Specifically, the structure of the self-attention module can refer to the multi-head attention mechanism of the lightweight Vision Transformer module shown in Figure 5

[0138] The step of inputting the preprocessed normal sample data set into the multi-scale image reconstruction network model for model parameter training in the step S4 includes:

[0139] In step S41, random color Gaussian noise is added to the images in the training set, and a mask operation is performed.

[0140] In step S42, the abnormal feature detection module in the image reconstruction model is excluded, and the processed image is input into the model for parameter training. The adversarial loss is used as the training loss function of the image discrimination network, and the image reconstruction loss is used as the training loss function of the image generation network.

[0141] In this example, the training loss function of the image discrimination network can select the Wasserstein distance combined with gradient penalty as the adversarial loss function, where the formula of the Wasserstein distance is:

[0142]

[0143] In the formula, P r is the real sample distribution, P g is the sample distribution generated by the generator, Π(P r ,P g ) is the set of all possible joint distributions combined by P r and P g , inf is the lower bound of the expected value, ||x w -y w || represents the norm, and E represents the expected value.

[0144] According to the Wasserstein distance, the loss function of the discrimination network can be obtained, and the formula is: ​

[0145]

[0146] In the formula, The expected value of the discriminator output representing the real sample, The expected value of the discriminator output representing the sample generated by the generator, and λ is a constant coefficient of the gradient penalty term, is the gradient of the discriminator network, when The gradient is penalized when the distance is far from 1, and the farther the distance is from 1, the greater the penalty is, where λ is 0.09, and x d is the input value of the discriminator;

[0147] The training loss function of the image generation network is composed of the mean square error and the SSIM structural similarity coefficient. The formula of the mean square error is as follows:

[0148]

[0149] In the formula, n represents the number of samples, y m represents the true value, represents the predicted value;

[0150] The formula of the SSIM structural similarity is as follows:

[0151]

[0152] In the formula, u and v represent two input image samples, μ represents the mean of the pixel samples, σ u , σ v represents the tolerance of the pixel samples, σ uv represents the correlation coefficient of u and v, and c1 and c2 are generally constant terms;

[0153] The training loss function of the image generation network is as follows:

[0154]

[0155] In the formula, α is a weight coefficient, which is 0.3 in this example;

[0156] Step S43: Fix the trained network weight, add an abnormal feature detection module, and train the parameters of the K-means clustering: randomly initialize the feature vector C of K cluster centers, calculate the distance between each input feature vector P and each cluster center, and assign the feature vector to the nearest cluster center. Calculate the mean of all samples in each cluster, and take the mean as the new cluster center. Iterate the above steps to complete the training of the K-means clustering.

[0157] In this example, the model training environment and related parameters: under the version of python3.7, the deep learning framework Pytorch1.7.0, the graphics card is NVIDIA GeForce RTX 3060, the Adam optimizer with default parameters, the batchsize is 16, the learning rate is 0.0001, and the epoch is 100.

[0158] The step of inputting the pre-processed to-be-detected image into the trained reconstruction model for image reconstruction in the step S5 includes the steps of:

[0159] In step S51, random color Gaussian noise is added to the pre-processed to-be-detected image, and the image is input into the image reconstruction model with fixed training weights for image reconstruction. When passing through the abnormal feature detection module, the spatial distance d of the input feature vector P to all cluster centers C is calculated. If the spatial distance d exceeds the abnormal threshold T, the feature vector is regarded as a defect feature, and the nearest center feature vector is used to replace the defect feature vector, so as to achieve the effect of defect feature suppression. The calculation formula of the spatial distance d is:

[0160]

[0161] In the formula, X and Y are the input feature vector P and the center feature vector C respectively, and d(X, Y) is the Euclidean distance between the two vectors.

[0162] The abnormal threshold T is obtained by training the feature vector of the positive sample, and the calculation formula is:

[0163]

[0164] In the formula, di is the spatial distance of the feature vector from the center feature vector during training of the positive sample, N is the number of feature vectors, and σ d is the standard deviation of the distance d i .

[0165] In step S52, the reconstructed image output by the model is pixel residual calculated with the to-be-detected image to obtain a preliminary residual image. The residual calculation formula is as follows:

[0166] L dif (i,j)=(L src (i,j)-L rec (i,j)) 2

[0167]

[0168] In the formula, L src (i,j) is the to-be-detected image, and Lrec (i,j) is a reconstructed image, the calculated residual image is further normalized to obtain the final result.

[0169] In the embodiment, the step S6 of pre-processing the residual image highlights the defect part, and the step of judging whether the infrared sealing image has defects by the abnormal score includes:

[0170] In step S61, the residual image is subjected to three-channel weighted average to realize gray scale processing.

[0171] In step S62, mean filtering is used for image denoising processing to eliminate pseudo defects formed by noise points.

[0172] In step S63, the abnormal probability score of the processed gray scale image is calculated, and when the abnormal probability score is greater than the abnormal threshold, it indicates that there is an abnormal point in the residual image, and the formula is:

[0173]

[0174] In the formula, S map is a single-channel residual image matrix, S mapmax is the maximum value in the pixel matrix, S mapmin is the minimum value in the pixel matrix, and S is the maximum value of the normalized residual image matrix.

[0175] In step S64, the residual image is subjected to binaryzation processing by using an adaptive threshold method, the threshold value of image segmentation is calculated by using the Otsu method, and the gray scale image is converted into a binaryzation image, and the formula is as follows:

[0176]

[0177] In the formula, T OTSU is the threshold value obtained by the Otsu method, and t is the pixel value in the gray scale image.

[0178] In step S65, according to the obtained binaryzation image, the proportion R of the pixel points with a pixel value of 1 in the whole image is calculated, and if the abnormal probability score is greater than the abnormal threshold and the proportion of the defect pixels is greater than the proportion threshold, it is determined that the image has defects, otherwise, the image is normal.

[0179] In step S66, if the image has defects, a contour extraction algorithm is used to traverse the pixel value of the binaryzation image, extract all the completely closed edge contours in the edge image, and use the maximum boundary box (circumscribed rectangle) to mark the target area to realize defect positioning.

[0180] The complete defect detection process includes four images, namely, the image to be detected, the reconstructed image, the residual image, and the defect positioning image, and the specific embodiment is shown in Figure 6 .

[0181] In the embodiments provided by the present application, it is obvious that the disclosed method can also be implemented by other ways, and is not limited to the above described method. For example, the steps described in the detection method flowchart in the drawings can be performed simultaneously, or the order of the steps can be changed, which can be determined according to different functions. Meanwhile, each block in the flowchart can be implemented by a special hardware system for performing the related functions or actions, or by a computer program in cooperation with the hardware.

[0182] When the functions are implemented in the form of software function modules and sold or used as a separate product, the system can be stored in a readable storage medium. Based on this, the technical solutions of the present application or part of the technical solutions can be used as a software product, which contains a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) to implement all or part of the methods described in the embodiments of the present application. The aforementioned storage medium includes a U disk, a mobile hard disk, a RAM, a ROM, and an optical disk, etc. which can store programs.

[0183] Based on the above ideal embodiments according to the present application, through the above description, the relevant staff can make various changes and modifications without deviating from the scope of the technical idea of the present application. The technical scope of the present application is not limited to the content in the specification, and must be determined according to the scope of the claims.

Claims

1. A method for detecting the sealing property of an aluminum foil closure based on unsupervised learning, characterized by, The method comprises the following steps: Step S1, acquiring an infrared image of an aluminum foil seal; Step S2, performing adaptive ROI region extraction on the infrared image of the aluminum foil seal; Step S3, preprocessing the extracted infrared image and making a normal sample image into a data set; Step S4, inputting the preprocessed normal sample data set into a multi-scale image reconstruction network model for model parameter training; The multi-scale image reconstruction network model in step S4 comprises an image generation network and an image discrimination network, wherein the image generation network comprises a multi-scale feature sampling module, a global context feature extraction module, an abnormal feature detection module and an image generation module; The multi-scale feature sampling module is composed of four convolution groups, the first layer convolution group is composed of a convolution layer with a convolution kernel of 7*7 and a maximum pooling layer, the purpose is to capture large-scale features and reduce the computational complexity of subsequent layers; the subsequent three convolution groups are composed of residual block structures, including a convolution layer with a 3*3 convolution kernel, batch normalization, an activation function and a residual connection, which are used to better extract texture, edge local features, and at the same time avoid gradient disappearance; the multi-scale feature sampling module outputs three different scale feature maps, which are output by the second layer, the third layer and the fourth layer convolution group respectively, the multi-scale feature improves the generalization ability and robustness of the model and is beneficial to the reconstruction of image details by the image reconstruction network; The global context feature extraction module fuses the Transformer structure and the convolution structure, and specifically includes a 3*3 convolution layer based on down-sampling and an improved lightweight Vision Transformer module, the improved lightweight Vision Transformer module sequentially includes relative position coding, a local perception unit, a LayerNorm layer, a lightweight multi-head self-attention module, a LayerNorm layer and an improved MLP module; Wherein, the local perception unit adopts a depth separable convolution, which introduces the translational invariance of convolution into the module, and the specific formula is: LPU(X)=DWConv(X)+X In the formula, X is an input feature tensor, DWConv is a depth separable convolution layer, and the whole adopts a residual connection; The lightweight multi-head self-attention module simplifies the generation of Key and Value through convolution operation on the basis of the original multi-self-attention, and the calculation formula of attention is as follows: In the formula, Q, K, and V are Query, Key, and Value in the Transformer respectively, Softmax is a normalized exponential function, K T is the transpose matrix of K, R is a real field, B is a bias matrix, and k is the multiple of K and V reduced along the spatial direction; The improved MLP module adds a 3*3 convolution layer between the original fully connected layers, and through the residual connection, it avoids gradient disappearance while enhancing the ability of the Transformer module to extract local semantic information; The global context feature extraction module has three branches, the input features are three different scale features output by the multi-scale feature sampling module, and after feature extraction by the three branch structure networks, the features are fused and output. The abnormal feature detection module is composed of K-means clustering and feature detection, and is mainly used for suppressing defect features in the image reconstruction process. The feature map output from the global context feature extraction module is decomposed along the dimension, containing N feature vectors P, P={p1, p2,..., pN}(P∈R C×1 , N=H*W), and is input into K-means clustering. K midpoints are selected as clustering centers, and the center vector is C, C={C1, C2,..., Ck}(C∈R C×1 ), and feature detection is to replace defect features in the feature vector that are too far away from the center vector with normal feature vectors. The main part of the image generation module and the image discrimination network adopts the network structure of DCGAN, the image generation module is composed of five up-sampling modules, the normalization operation in the module adopts Instance Norm, the image generation module outputs a three-channel color reconstructed image, and the image size is consistent with the input image; the image discrimination network includes five feature down-sampling modules and a fully connected layer, a self-attention module is added between the fourth and fifth layers of the image generation module and between the first and second layers of the image discrimination network, and the output of the image discrimination network is the probability that the input image is true, ranging from 0 to 1; Step S5, input the preprocessed to-be-detected image into the trained reconstruction model for image reconstruction, output a reconstructed image, and perform residual error calculation on the to-be-detected image and the reconstructed image to obtain a residual error image; Step S6, pre-process the residual error image to highlight the defect part, and determine whether the infrared seal image has defects through an abnormal score, and if there are defects, the defect part is positioned.

2. The aluminum foil seal sealing performance detection method based on unsupervised learning according to claim 1, characterized in that, the step S2 of adaptively extracting the ROI region of the aluminum foil seal infrared image comprises: Step S21, performing three-channel weighted averaging on the obtained infrared image to realize gray-scale processing; Step S22, performing smoothing and noise reduction processing on the gray-scale image to remove small textures and irregular noise points, and the smoothing and noise reduction processing method is Gaussian filtering, and its expression is: In the formula, (x, y) represents the pixel coordinates in the image, G(x, y) represents the Gaussian function value calculated at the position, sigma represents the standard of the Gaussian kernel, and e is the constant base number of the natural logarithmic function; Step S23, the edge profile of the gray-scale image after smoothing processing is extracted by using Canny edge detection, and a binary edge image is obtained; Step S24, a contour extraction algorithm is used to traverse the pixel values of the binary image, extract all completely closed edge profiles in the edge image, and use the maximum wheel outline, that is, the circumscribed rectangle to mark the target region, and according to the position and size of the circumscribed rectangle, the corresponding region is cropped from the original image, and the obtained corresponding region is the ROI region of the aluminum foil seal infrared image.

3. The aluminum foil seal sealing performance detection method based on unsupervised learning according to claim 1, characterized in that, the step S3 of pre-processing the extracted infrared image and making normal sample images into a data set comprises: Step S31, image enhancement is performed on the extracted ROI image, first, the contrast of the image is enhanced to improve the visibility of image details, and specifically, adaptive histogram equalization is used, and the formula is as follows: In the formula, f(x, y) is the original image, (x, y) is the pixel point coordinate, h(i, j) is the window function, W is the window size; T(x, y) refers to the pixel value with (x, y) as the center, f(x+i, y+j) refers to the pixel value with (x+i, y+j) as the center in the original image, and i and j refer to the coordinates of the window function; Then, the Laplace filter is used to sharpen the image, enhance the edge and details of the image. Step S32, the sample after image enhancement is interpolated and up-sampled to improve the resolution of the infrared image, and more image information is provided for subsequent deep learning model training. Specifically, the bicubic interpolation method is used, and the formula is as follows: where (b x ,b y ) is the interpolation point coordinate, B(b x ,b y ) is the bicubic interpolation result, (b xi ,b yj ) is the 16 nearest points of the interpolation point, and W(x) is the weighting function, where a is -0.

5. Step S33, the image is randomly rotated to expand the training data set and increase the robustness of the model; Step S34, the image pixel value is normalized to 0 to 1, and the normalized image pixel value is standardized to have a mean of 0 and a variance of 1, so as to accelerate the model convergence, and the formula is: where x norm is the normalized result, x std is the standardized result, x org is the pixel value of the original image, and μ and σ are the mean and standard deviation of the single channel pixel value. Step S35, the preprocessed image data set is divided into training and test sets, and the normal sample data is used as the training set, accounting for 75%, and the remaining 25% of normal samples and defect samples are used as the test set.

4. The aluminum foil seal tightness detection method based on unsupervised learning according to claim 1, wherein The step S4 of inputting the preprocessed normal sample data set into the multi-scale image reconstruction network model for model parameter training comprises: Step S41, random color Gaussian noise is added to the image in the training set, and a mask operation is performed; Step S42, the abnormal feature detection module in the image reconstruction model is excluded, the processed image is input into the model for parameter training, the adversarial loss is used as the training loss function of the image discrimination network, and the image reconstruction loss is used as the training loss function of the image generation network; Step S43, the trained network weight is fixed, the abnormal feature detection module is added, and the K-means clustering is trained: the feature vector C of the K cluster centers is randomly initialized, the distance between each input feature vector P and each cluster center is calculated, and the feature vector is assigned to the nearest cluster center. For each cluster, calculate the mean of all samples in the cluster, and take the mean as the new cluster center. Iterate the above steps to complete the training of K-means clustering.

5. The aluminum foil seal tightness detection method based on unsupervised learning according to claim 1, wherein The step S5 of inputting the preprocessed test image into the trained reconstruction model for image reconstruction, outputting the reconstructed image, and calculating the residual error between the test image and the reconstructed image to obtain the residual error image comprises: Step S51, random color Gaussian noise is added to the preprocessed test image, and the image is input into the image reconstruction model with fixed training weight for image reconstruction. When passing through the abnormal feature detection module, the spatial distance d between the input feature vector P and all cluster centers C is calculated. If the spatial distance d exceeds the abnormal threshold T, the feature vector is regarded as a defect feature, and the nearest center feature vector is used to replace the defect feature vector to achieve the effect of defect feature suppression. The formula for calculating the spatial distance d is: In the formula, X and Y are the input feature vector P and the center feature vector C, respectively, and d(X, Y) is the Euclidean distance between the two vectors. The abnormal threshold T is obtained by training the feature vector of the positive sample, and the formula is: where di is the spatial distance of the feature vector from the center feature vector during positive sample training, N is the number of feature vectors, and σ d is the standard deviation of the distances d i . Step S52, pixel residual error calculation is performed between the reconstructed image output by the model and the to-be-tested image to obtain a preliminary residual error image. The formula for residual error calculation is as follows: L dif (i,j) = (L src (i,j) - L rec (i,j)) 2 wherein L src (i,j) is the image to be detected, L rec (i,j) is the reconstructed image, the computed residual map is further normalized to obtain the final result.

6. The method for detecting the sealing performance of an aluminum foil seal based on unsupervised learning according to claim 1, characterized in that, The step S6 of pre-processing the residual error image to highlight the defect part and determining whether the infrared seal image has defects by using the abnormal score, if there is a defect, the step of locating the defect part includes: Step S61, three-channel weighted average is performed on the residual error image to realize grayscale processing; Step S62, mean filter is used for image denoising processing to eliminate pseudo-defects caused by noise points; Step S63, the abnormal probability score of the processed grayscale image is calculated. When the abnormal probability score is greater than the abnormal threshold, it indicates that there is an abnormal point in the residual error image. The formula is as follows: In the formula, S map is a single-channel residual map matrix, S mapmax is the maximum value in the pixel matrix, S mapmin is the minimum value in the pixel matrix, S is the maximum value of the normalized residual map matrix; Step S64, the residual error image is binarized by using the adaptive threshold method. Otsu method is used to calculate the image segmentation threshold to convert the grayscale image into a binary image. The formula is as follows: In the formula, T OTSU is the threshold value obtained by the Otsu method, and t is the pixel value in the gray image. Step S65, according to the obtained binary image, the proportion R of the pixel points with pixel value 1 in the whole image is calculated. If the abnormal probability score is greater than the abnormal threshold and the proportion of defect pixels is greater than the proportion threshold, it is determined that the image has defects. Otherwise, the image is normal. Step S66, if the image has defects, the contour extraction algorithm is used to traverse the pixel value of the binary image, extract all the completely closed edge contours in the edge image, and use the maximum wheel width boundary box to mark the target area to realize defect positioning.