A Semi-Supervised Anomaly Detection Method Based on Memory Enhancement and Pseudo-Labeling
Feature extraction and reconstruction are carried out through the QC-Net model, combined with pseudo-labels and historical information, and the problems of low accuracy and poor stability of existing anomaly detection methods under complex image data are solved, achieving high-precision and low-cost industrial anomaly detection.
Patent Information
- Application Number
- CN202411969124.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2044-12-30
AI Technical Summary
The existing anomaly detection methods have low accuracy and poor stability under complex image data, and high data labeling costs, making it difficult to meet the high accuracy and stability requirements of industrial production.
The semi-supervised anomaly detection method based on memory enhancement and pseudo-tagging is adopted, and feature extraction and reconstruction are performed through the QC-Net model, and abnormal detection is performed by combining pseudo-tagging and historical information, including the Encoder-Decoder module, VAE-GAN encoder, pseudo-tagging machine and memory module. A small amount of annotated data is used to generate pseudo-tagging and fuse a variety of information for abnormal judgment.
It improves the accuracy of complex image anomaly detection, enhances the stability of the model under complex data, reduces the cost of data labeling, and improves the practicality and scalability of the detection method.
Smart Images

Figure CN119831972B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine vision and anomaly detection, and particularly to a semi-supervised anomaly detection method based on memory enhancement and pseudo-labeling. Background Art
[0002] With the development of the intelligentization of industrial production, the demand for the automation of product quality inspection is increasing day by day. When inspecting products, image-based anomaly detection has become a research and application hotspot due to its intuitive and efficient characteristics. By analyzing product images, it is possible to quickly and accurately identify whether there are anomalies in the products, which helps to improve production efficiency and product quality.
[0003] However, in the face of complex image data, existing anomaly detection methods are difficult to comprehensively and accurately capture the anomaly features in the images due to the lack of effective feature extraction and fusion mechanisms, resulting in low accuracy of anomaly detection and inability to meet the requirements of high-precision anomaly detection in actual production. In the surface defect detection of industrial products, complex textures, lighting and other factors will interfere with the identification of anomalies, resulting in many misjudgments and missed detections in the detection results.
[0004] Traditional supervised learning anomaly detection methods require a large amount of labeled data to train the model. However, in actual applications, obtaining a large amount of labeled data is often costly, time-consuming and laborious. Moreover, for some rare anomaly situations, due to the scarcity of labeled data and lack of sufficient samples, the application effect of supervised learning methods in anomaly detection is limited. Existing semi-supervised anomaly detection methods, in the face of complex image data, due to the lack of effective feature extraction and fusion mechanisms, as well as the reasonable utilization of historical anomaly information, result in insufficient accuracy and stability of anomaly detection, and are unable to meet the requirements of high-precision anomaly detection in actual production. When processing data in different batches or in a complex and changeable environment, the stability of its detection results is poor and the fluctuations are large, unable to provide reliable and consistent detection results, and it is difficult to be competent for some application scenarios with high stability requirements. Summary of the Invention
[0005] In view of the above analysis, embodiments of the present invention aim to provide a semi-supervised anomaly detection method based on memory enhancement and pseudo-labeling to solve the problems of low accuracy of existing anomaly detection, poor stability of the model, and high cost of data annotation.
[0006] On the one hand, embodiments of the present invention provide a semi-supervised anomaly detection method based on memory enhancement and pseudo-labeling, including the following steps:
[0007] Obtain the image data of the product to be detected, input the image data of the product to be detected into the pre-constructed QC-Net model, and obtain the anomaly detection result; wherein, the QC-Net model includes an input module, an Encoder-Decoder module, a VAE-GAN encoder, a pseudo-labeler, a memory module, and an anomaly detection module;
[0008] The input module is used to receive the image data of the product to be detected and send the image data of the product to be detected to the Encoder-Decoder module, the VAE-GAN encoder, the pseudo-labeler, the anomaly detection module, and the memory module respectively;
[0009] The Encoder-Decoder module extracts features from the image data of the product to be detected and sends them to the pseudo-labeler; and generates image reconstruction data based on the extracted features and sends the image reconstruction data to the anomaly detection module;
[0010] The VAE-GAN encoder is used to generate potential anomaly information based on the image data of the product to be detected, generate anomaly reconstruction data based on the potential anomaly information, and send the potential anomaly information and the anomaly reconstruction data to the anomaly detection module;
[0011] The pseudo-labeler is used to generate an abnormal label based on the features extracted by the Encoder-Decoder module and send it to the anomaly detection module;
[0012] The memory module is used to store historical abnormal image data and its corresponding anomaly detection results;
[0013] The anomaly detection module obtains the anomaly detection result based on the image data of the product to be detected, the image reconstruction data, the potential anomaly information, the anomaly reconstruction data, the pseudo-label, the historical abnormal image data stored in the memory module and its corresponding anomaly detection result, and transmits the anomaly detection result to the memory module.
[0014] As a further improvement of this application, the Encoder-Decoder module includes an Encoder unit and a Decoder unit;
[0015] The Encoder unit includes a first convolutional layer, a first pooling layer, a second convolutional layer, and a second pooling layer that are connected end to end in sequence;
[0016] The first convolutional layer extracts the features of the image data, and the first pooling layer performs a pooling operation on the features extracted by the first convolutional layer; the second convolutional layer performs a convolutional operation on the output of the first pooling layer to extract the features of the image data; the second pooling layer performs a dimensionality reduction process on the output of the second convolutional layer to obtain the final features of the image data. The output of the second pooling layer is used as the output of the Encoder unit, which is connected to the inputs of the Decoder unit and the pseudo-labeler.
[0017] The Decoder unit includes a transposed convolutional layer, an upsampling layer, and a Decoder convolutional layer that are connected end to end in sequence.
[0018] The transposed convolutional layer receives the output of the second pooling layer and performs a transposed convolutional operation on the final features of the image data; the upsampling layer upsamples the output of the transposed convolutional layer to increase the feature space dimension; the Decoder convolutional layer performs a convolutional operation on the output of the upsampling layer to generate the image reconstruction data. The output of the Decoder convolutional layer is used as the output of the Decoder unit, which is connected to the input of the anomaly detection module.
[0019] As a further improvement of this application, the VAE-GAN encoder includes an encoder and a decoder connected end to end. The encoder is used to generate potential anomaly information based on the image data; the decoder is used to generate reconstructed image data based on the potential anomaly information.
[0020] As a further improvement of this application, the encoder includes a first convolutional layer, a first activation function layer, a first pooling layer, a second convolutional layer, a second activation function layer, a second pooling layer, a Flatten layer, a first fully connected layer, a third activation function layer, and a second fully connected layer that are connected end to end in sequence.
[0021] Among them, the first convolutional layer extracts the features of the image data, and the first activation function layer performs a non-linear process on the output of the first convolutional layer; the first pooling layer performs a pooling operation on the output of the first encoder activation function layer to reduce the dimension of the features after non-linear processing; the second convolutional layer performs a convolutional operation on the output of the first pooling layer to extract the features of the image data; the second activation function layer performs a non-linear process on the output of the second convolutional layer; the second pooling layer performs a pooling operation on the output of the second activation function layer; the Flatten layer converts the output of the second pooling layer into a one-dimensional feature vector; the first fully connected layer compresses the features of the one-dimensional feature vector; the third activation function layer performs a non-linear process on the output of the first fully connected layer; the second fully connected layer performs a feature mapping on the output of the third activation function layer to generate the potential anomaly information of the image data.
[0022] As a further improvement of this application, the decoder includes a first transposed convolutional layer, a first activation function layer, an upsampling layer, a second transposed convolutional layer, a second activation function layer, and a decoder convolutional layer that are connected end to end in sequence.
[0023] Among them, the first transposed convolution layer performs a transposed convolution operation on the potential abnormal information received from the encoder; the first activation function layer performs a non-linear process on the output of the first transposed convolution layer; the upsampling layer performs upsampling on the output of the first activation function layer; the second transposed convolution layer performs a transposed convolution operation on the output of the upsampling layer; the second activation function layer performs a non-linear process on the output of the second transposed convolution layer; the decoder convolution layer performs a convolution operation on the output of the second activation function layer to obtain reconstructed image data.
[0024] As a further improvement of the present application, the pseudo-labeler includes an input layer, a fully connected layer, an activation function layer, and an output layer that are connected end to end in sequence;
[0025] The input layer receives the final features of the image data output by the Encoder unit and transmits them to the fully connected layer;
[0026] The fully connected layer performs a spatial mapping on the final features of the image data to generate the number of feature categories;
[0027] The activation function layer generates the distribution probability of each feature category based on the number of feature categories through the Softmax function;
[0028] The output layer generates an abnormal label based on the distribution probability of each feature category and transmits it to the anomaly detection module.
[0029] As a further improvement of the present application, the anomaly detection module includes:
[0030] A reconstruction deviation calculation unit for calculating a reconstruction deviation based on the image reconstruction data and the image data of the product to be detected;
[0031] An abnormal deviation calculation unit for calculating an abnormal deviation based on the abnormal reconstruction data and the image data of the product to be detected;
[0032] A pseudo-label deviation calculation unit for calculating a pseudo-label deviation based on the abnormal label and the true label of the pre-annotated image data of the product to be detected;
[0033] A regularization deviation calculation unit for calculating a regularization deviation based on the potential abnormal information;
[0034] An anomaly localization unit for performing feature fusion on the reconstruction deviation, abnormal deviation, pseudo-label deviation, and regularization deviation to obtain a comprehensive deviation feature; retrieving the historical abnormal image with the highest similarity to the comprehensive deviation feature in the memory module, and determining the position of the abnormal area in the image of the product to be detected based on the abnormal area of the historical abnormal image;
[0035] An anomaly classification unit for obtaining an anomaly category based on the abnormal label output by the pseudo-labeler.
[0036] As a further improvement of the present application, a reconstruction deviation is calculated based on the image reconstruction data and the image data of the product to be detected, as shown in calculation formula (1);
[0037]
[0038] where n is the number of image data of the product to be detected, y i is the image data of the i-th product to be detected, is the i-th image reconstruction data.
[0039] The abnormal deviation is calculated based on the abnormal reconstruction data and the image data of the product to be detected, as shown in calculation formula (2);
[0040]
[0041] where n is the number of image data of the product to be detected, a i is the image data of the i-th product to be detected, is the i-th abnormal reconstruction data.
[0042] As a further improvement of the present application, a pseudo-label deviation is calculated based on the pseudo-label and the true label of the pre-annotated image data of the product to be detected, as shown in calculation formula (3);
[0043]
[0044] where n is the number of image data of the product to be detected, y ic is the encoding of the true label of the i-th image data, is the pseudo-label encoding of the i-th image data.
[0045] As a further improvement of the present application, a regularization deviation is calculated based on the potential abnormal information, as shown in calculation formula (4);
[0046] L latent = D KL (q(z|xi)||p(z)) (4)
[0047] where D KL is the KL divergence, q(z|xi) is the potential abnormal information, and p(z) is the prior distribution of the abnormal information.
[0048] Compared with the prior art, the present invention can at least achieve one of the following beneficial effects:
[0049] 1. The present invention constructs a QC-Net model, including an Encoder-Decoder module, a VAE-GAN encoder, and a pseudo-labeler. Each module extracts, analyzes, and reconstructs image data features from different perspectives, can capture abnormal features more comprehensively and accurately, calculates deviations based on multiple pieces of information and performs feature fusion, comprehensively judges abnormal situations, improves the accuracy of complex image anomaly detection, and solves the problem of low accuracy of existing methods in complex image anomaly detection.
[0050] 2. The memory module of the present invention can store historical abnormal image data and their corresponding anomaly detection results. When detecting new product images, the anomaly localization unit retrieves the historical abnormal image with the highest similarity to the comprehensive deviation feature in the memory module, and determines the location of the abnormal area based on historical experience. The model can make full use of historical information, reduce the volatility of detection results, and enhance stability under complex data, solving the problem of insufficient stability of existing models.
[0051] 3. The present invention adopts a semi-supervised learning method. The pseudo-labeler generates pseudo-labels based on a small amount of labeled data. When calculating the deviation of pseudo-labels, the relationship between pseudo-labels and other features is fully considered, enabling the model to more accurately utilize the information in unlabeled data, effectively improving the accuracy of semi-supervised anomaly detection, reducing the need for large-scale labeled data, lowering the cost of data labeling, and improving the practicability and scalability of the detection method.
[0052] In the present invention, the above technical solutions can also be combined with each other to achieve more preferred combination schemes. Other features and advantages of the present invention will be described in the subsequent specification, and some advantages will be obvious from the specification or understood by implementing the present invention. The objectives and other advantages of the present invention can be realized and obtained through the content specifically pointed out in the specification and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] The drawings are only for the purpose of showing specific embodiments and are not considered as limiting the present invention. Throughout the drawings, the same reference signs denote the same components;
[0054] Figure 1 It is a schematic structural diagram of the QC-Net model provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0055] The following will specifically describe the preferred embodiments of the present invention with reference to the drawings. The drawings form a part of this application and are used together with the embodiments of the present invention to explain the principles of the present invention, rather than to limit the scope of the present invention.
[0056] A specific embodiment of the present invention discloses a semi-supervised anomaly detection method based on memory enhancement and pseudo-labeling, including:
[0057] Obtain the image data of the product to be detected, and input the image data of the product to be detected into a pre-constructed QC-Net model to obtain an anomaly detection result; wherein, the QC-Net model includes an input module, an Encoder-Decoder module, a VAE-GAN encoder, a pseudo-labeler, a memory module, and an anomaly detection module;
[0058] The input module is used to receive the image data of the product to be detected, and send the image data of the product to be detected to the Encoder-Decoder module, the VAE-GAN encoder, the pseudo-labeler, the anomaly detection module, and the memory module respectively;
[0059] The Encoder-Decoder module extracts features from the image data of the product to be detected and sends them to the pseudo-labeler; and generates image reconstruction data based on the extracted features, and sends the image reconstruction data to the anomaly detection module;
[0060] The VAE-GAN encoder is used to generate potential anomaly information based on the image data of the product to be detected, generate anomaly reconstruction data based on the potential anomaly information, and send the potential anomaly information and the anomaly reconstruction data to the anomaly detection module;
[0061] The pseudo-labeler is used to generate an abnormal label based on the features extracted by the Encoder-Decoder module and send it to the anomaly detection module;
[0062] The memory module is used to store historical abnormal image data and its corresponding anomaly detection results;
[0063] The anomaly detection module obtains an anomaly detection result based on the image data of the product to be detected, the image reconstruction data, the potential anomaly information, the anomaly reconstruction data, the pseudo-label, the historical abnormal image data stored in the memory module, and its corresponding anomaly detection result, and transmits the anomaly detection result to the memory module.
[0064] As Figure 1 shown. The image data of the product to be detected is detected by a pre-constructed QC-Net model. The QC-Net model includes an input module, an Encoder-Decoder module, a VAE-GAN encoder, a pseudo-labeler, a memory module, and an anomaly detection module.
[0065] The input module is the input end of the QC-Net model and receives the image data of the product to be detected. The image data of the product to be detected refers to a two-dimensional or three-dimensional image including information such as the appearance and structure of the product to be detected, and can be a grayscale image, a color image, or other forms of image data. The input module sends the received image data to the Encoder-Decoder module, the VAE-GAN encoder, the anomaly detection module, and the memory module respectively.
[0066] The Encoder-Decoder module extracts features from the image data of the product to be detected and sends them to the pseudo-labeler; and generates image reconstruction data based on the extracted features and sends the image reconstruction data to the anomaly detection module. The extracted features include features that can represent image information such as edges, textures, shapes, etc. Generating image reconstruction data means restoring the extracted features into an image, and gradually restoring the low-dimensional features to a spatial size similar to the original image through a combination of transposed convolutional layers, upsampling layers, and convolutional layers, and the difference between the original image and the reconstructed image can be judged to discover anomalies.
[0067] Specifically, the Encoder-Decoder module includes an Encoder unit and a Decoder unit;
[0068] The Encoder unit includes a first convolutional layer, a first pooling layer, a second convolutional layer, and a second pooling layer connected end to end in sequence;
[0069] The first convolutional layer extracts the features of the image data, and the first pooling layer performs a pooling operation on the features extracted by the first convolutional layer; the second convolutional layer performs a convolutional operation on the output of the first pooling layer to extract the features of the image data; the second pooling layer performs a dimensionality reduction process on the output of the second convolutional layer to obtain the final features of the image data, and the output of the second pooling layer serves as the output of the Encoder unit, connecting the input of the Decoder unit and the pseudo-labeler;
[0070] The Decoder unit includes a transposed convolutional layer, an upsampling layer, and a Decoder convolutional layer connected end to end in sequence;
[0071] The transposed convolutional layer receives the output of the second pooling layer and performs a transposed convolutional operation on the final features of the image data; the upsampling layer performs upsampling on the output of the transposed convolutional layer to increase the feature space dimension; the Decoder convolutional layer performs a convolutional operation on the output of the upsampling layer to generate image reconstruction data, and the output of the Decoder convolutional layer serves as the output of the Decoder unit, connecting the input of the anomaly detection module.
[0072] The VAE-GAN encoder is used to generate potential anomaly information based on the image data of the product to be detected, generate anomaly reconstruction data based on the potential anomaly information, and send the potential anomaly information and the anomaly reconstruction data to the anomaly detection module.
[0073] Specifically, the VAE-GAN encoder includes an encoder and a decoder connected end to end. The encoder is used to generate potential anomaly information based on the image data; the decoder is used to generate reconstructed image data based on the potential anomaly information. The potential anomaly information is a potential representation obtained by encoding the image data, which can contain the anomaly features existing in the image. The potential anomaly information needs to be restored to observable image information through the decoder. The anomaly reconstruction data is the image obtained by decoding the potential anomaly information through the decoder part of the VAE-GAN encoder. Compared with the original image, the decoded image will show obvious differences in the anomaly area, and this difference can be used as the basis for anomaly detection.
[0074] Generating anomaly reconstruction data based on the potential anomaly information provides another perspective for detecting anomalies, and sending the potential anomaly information and the anomaly reconstruction data to the anomaly detection module.
[0075] The encoder includes a first convolutional layer, a first activation function layer, a first pooling layer, a second convolutional layer, a second activation function layer, a second pooling layer, a Flatten layer, a first fully connected layer, a third activation function layer, and a second fully connected layer connected end to end in sequence; among them, the first convolutional layer extracts the features of the image data, and the first activation function layer performs nonlinear processing on the output of the first convolutional layer; the first pooling layer performs a pooling operation on the output of the first encoder activation function layer to reduce the feature dimension after nonlinear processing; the second convolutional layer performs a convolutional operation on the output of the first pooling layer to extract the features of the image data; the second activation function layer performs nonlinear processing on the output of the second convolutional layer; the second pooling layer performs a pooling operation on the output of the second activation function layer; the Flatten layer converts the output of the second pooling layer into a one-dimensional feature vector; the first fully connected layer compresses the features of the one-dimensional feature vector; the third activation function layer performs nonlinear processing on the output of the first fully connected layer; the second fully connected layer performs feature mapping on the output of the third activation function layer to generate the potential anomaly information of the image data.
[0076] The decoder includes a first transposed convolutional layer, a first activation function layer, an upsampling layer, a second transposed convolutional layer, a second activation function layer, and a decoder convolutional layer connected end to end in sequence;
[0077] Among them, the first transposed convolutional layer performs a transposed convolution operation on the latent anomaly information received from the encoder; the first activation function layer performs a non-linear processing on the output of the first transposed convolutional layer; the upsampling layer upsamples the output of the first activation function layer; the second transposed convolutional layer performs a transposed convolution operation on the output of the upsampling layer; the second activation function layer performs a non-linear processing on the output of the second transposed convolutional layer; the decoder convolutional layer performs a convolution operation on the output of the second activation function layer to obtain anomaly reconstruction image data.
[0078] The pseudo-labeler is used to generate anomaly labels based on the features extracted by the Encoder-Decoder module and send them to the anomaly detection module. Pseudo-labels are a type of predicted labels, which are obtained through the processing of the input layer, fully connected layer, activation function layer, and output layer, and can provide additional supervision information in the case of limited labeled data. For product images, the pseudo-labeler can generate anomaly labels based on the distribution probabilities of each feature category.
[0079] Specifically, the pseudo-labeler includes an input layer, a fully connected layer, an activation function layer, and an output layer that are connected end to end in sequence;
[0080] The input layer receives the final features of the image data output by the Encoder unit and transmits them to the fully connected layer;
[0081] The fully connected layer performs a spatial mapping on the final features of the image data to generate the number of feature categories;
[0082] The activation function layer generates the distribution probabilities of each feature category based on the number of feature categories through the Softmax function;
[0083] The output layer generates anomaly labels based on the distribution probabilities of each feature category and transmits them to the anomaly detection module.
[0084] Furthermore, the anomaly detection module obtains the anomaly detection result based on the image data of the product to be detected, the image reconstruction data, the latent anomaly information, the anomaly reconstruction data, the anomaly labels, the historical anomaly image data stored in the memory module and their corresponding anomaly detection results, and transmits the anomaly detection result to the memory module. The anomaly detection result includes the anomaly area and the anomaly type.
[0085] The anomaly detection module includes:
[0086] The reconstruction deviation calculation unit is used to calculate the reconstruction deviation based on the image reconstruction data and the image data of the product to be detected. The reconstruction deviation calculation unit calculates the reconstruction deviation by comparing the difference between the image reconstruction data and the original image data.
[0087] Calculating the reconstruction deviation based on the image reconstruction data and the image data of the product to be detected is shown in calculation formula (1);
[0088]
[0089] Among them, n is the number of image data of the product to be detected, and y i is the image data of the i-th product to be detected, is the i-th image reconstruction data.
[0090] An abnormal deviation calculation unit for calculating an abnormal deviation based on the abnormal reconstruction data and the image data of the product to be detected. The abnormal deviation calculation unit calculates the abnormal deviation by comparing the difference between the abnormal reconstruction data and the image data of the product to be detected.
[0091] Calculating the abnormal deviation based on the abnormal reconstruction data and the image data of the product to be detected is shown in calculation formula (2);
[0092]
[0093] Among them, n is the number of image data of the product to be detected, and a i is the image data of the i-th product to be detected, is the i-th abnormal reconstruction data.
[0094] A pseudo-label deviation calculation unit for calculating a pseudo-label deviation based on the abnormal label and the true label of the pre-annotated image data of the product to be detected. The pseudo-label deviation calculation unit calculates the pseudo-label deviation by comparing the difference between the abnormal label generated by the pseudo-labeler and the pre-annotated true label.
[0095] Calculating the pseudo-label deviation based on the abnormal label and the true label of the pre-annotated image data of the product to be detected is shown in calculation formula (3);
[0096]
[0097] Among them, n is the number of image data of the product to be detected, and y ic is the encoding of the true label of the i-th image data, is the pseudo-label encoding of the i-th image data.
[0098] A regularization deviation calculation unit for calculating a regularization deviation based on the latent abnormal information. The regularization deviation measures the difference between the distribution of the latent abnormal information and the prior distribution through divergence. By calculating the divergence, it can be known whether the distribution of the latent abnormal information generated by the current model conforms to the expected prior distribution. Divergence is used as a constraint condition for the model. By minimizing the divergence, the model can be trained to make the distribution of the latent abnormal information closer to the prior distribution and improve the prediction accuracy of the model.
[0099] Calculating the regularization deviation based on the latent abnormal information is shown in calculation formula (4);
[0100] L latent = D KL (q(z|xi)||p(z)) (4)
[0101] where D KL is the KL divergence, q(z|xi) is the potential anomaly information, and p(z) is the prior distribution of the anomaly information.
[0102] An anomaly localization unit is used to perform feature fusion on the reconstruction deviation, anomaly deviation, pseudo-label deviation, and regularization deviation to obtain a comprehensive deviation feature; retrieve the historical anomaly image with the highest similarity to the comprehensive deviation feature in the memory module, and determine the position of the anomaly area in the image of the product to be detected based on the anomaly area of the historical anomaly image.
[0103] The anomaly localization unit performs feature fusion on the reconstruction deviation, anomaly deviation, pseudo-label deviation, and regularization deviation to obtain a comprehensive deviation feature. The comprehensive deviation feature is shown in the calculation formula (5).
[0104] L total = ω1.L recon + ω2.L anomaly + ω3.L latent + ω4.L pseudo (5);
[0105] where L total is the comprehensive deviation feature; L recon is the reconstruction deviation; L anomaly is the anomaly deviation; L latent is the pseudo-label deviation; L pseudo is the regularization deviation. ω1, ω2, ω3, and ω4 are the weights of the reconstruction deviation, anomaly deviation, pseudo-label deviation, and regularization deviation, which are preset in advance through experiments or experience based on different anomaly detection tasks. If the reconstructed image is significantly different from the original image, the weight of the reconstruction deviation can be increased; for tasks that rely more on the accuracy of the pseudo-label, the weight of the pseudo-label deviation can be increased.
[0106] When retrieving the historical anomaly image with the highest similarity to the comprehensive deviation feature in the memory module, the Euclidean distance or hash value can be used as the similarity metric to retrieve the historical anomaly image with the highest similarity to the comprehensive deviation feature in the memory module. The anomaly classification unit maps the anomaly label output by the pseudo-labeler to a specific anomaly category.
[0107] Furthermore, the memory module is used to store the historical anomaly image data and its corresponding anomaly detection results. The historical anomaly image data is the image data that has been determined to have anomalies in the previously completed detected image data, and its corresponding anomaly detection results include the anomaly area and anomaly type.
[0108] The above various embodiments of the present invention have the following beneficial effects: By constructing a QC-Net model, including an Encoder-Decoder module, a VAE-GAN encoder, and a pseudo-labeler, each module extracts, analyzes, and reconstructs image data features from different perspectives, can capture abnormal features more comprehensively and accurately, calculates deviations based on various information and performs feature fusion, comprehensively judges abnormal situations, improves the accuracy of complex image anomaly detection, and solves the problem of low accuracy of existing methods in complex image anomaly detection; The memory module can store historical abnormal image data and their corresponding anomaly detection results. When detecting new product images, the anomaly localization unit retrieves the historical abnormal image with the highest similarity to the comprehensive deviation feature in the memory module, determines the location of the abnormal area based on historical experience, the model can make full use of historical information, reduce the volatility of detection results, and enhance stability under complex data, solving the problem of insufficient stability of existing models; Adopting a semi-supervised learning method, the pseudo-labeler generates pseudo-labels based on a small amount of labeled data. When calculating the deviation of the pseudo-labels, the relationship between the pseudo-labels and other features is fully considered, enabling the model to more accurately utilize the information in the unlabeled data, effectively improving the accuracy of semi-supervised anomaly detection, reducing the need for large-scale labeled data, lowering the cost of data annotation, and improving the practicability and scalability of the detection method.
[0109] Those skilled in the art can understand that all or part of the processes of implementing the above method embodiments can be completed by instructing relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium is a magnetic disk, an optical disk, a read-only memory, or a random access memory, etc.
[0110] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention.
Claims
1. A semi-supervised anomaly detection method based on memory enhancement and pseudo-labeling, comprising the following steps: Obtain the image data of the product to be detected, and input the image data of the product to be detected into a pre-constructed QC-Net model to obtain an anomaly detection result; wherein, the QC-Net model includes an input module, an Encoder-Decoder module, a VAE-GAN encoder, a pseudo-labeler, a memory module, and an anomaly detection module; The input module is used to receive the image data of the product to be detected, and send the image data of the product to be detected to the Encoder-Decoder module, the VAE-GAN encoder, the pseudo-labeler, the anomaly detection module, and the memory module respectively; The Encoder-Decoder module extracts features from the image data of the product to be detected and sends them to the pseudo-labeler; and generates image reconstruction data based on the extracted features, and sends the image reconstruction data to the anomaly detection module; The VAE-GAN encoder is used to generate potential anomaly information based on the image data of the product to be detected, generate anomaly reconstruction data based on the potential anomaly information, and send the potential anomaly information and the anomaly reconstruction data to the anomaly detection module; The pseudo-labeler is used to generate anomaly labels based on the features extracted by the Encoder-Decoder module and send them to the anomaly detection module; The memory module is used to store historical anomaly image data and their corresponding anomaly detection results; The anomaly detection module obtains an anomaly detection result based on the image data of the product to be detected, the image reconstruction data, the potential anomaly information, the anomaly reconstruction data, the anomaly label, the historical anomaly image data stored in the memory module, and their corresponding anomaly detection results, and transmits the anomaly detection result to the memory module.
2. The method according to claim 1, characterized in that, The Encoder-Decoder module includes an Encoder unit and a Decoder unit; The Encoder unit includes a first convolutional layer, a first pooling layer, a second convolutional layer, and a second pooling layer connected end to end in sequence; The first convolutional layer extracts the features of the image data, and the first pooling layer performs a pooling operation on the features extracted by the first convolutional layer; the second convolutional layer performs a convolutional operation on the output of the first pooling layer to extract the features of the image data; the second pooling layer performs dimensionality reduction processing on the output of the second convolutional layer to obtain the final features of the image data, and the output of the second pooling layer is used as the output of the Encoder unit, connecting the input of the Decoder unit and the pseudo-labeler; The Decoder unit includes a transposed convolutional layer, an upsampling layer, and a Decoder convolutional layer connected end to end in sequence; The transposed convolutional layer receives the output of the second pooling layer and performs a transposed convolutional operation on the final features of the image data; the upsampling layer performs upsampling on the output of the transposed convolutional layer to increase the feature space dimension; The Decoder convolutional layer performs a convolutional operation on the output of the upsampling layer to generate image reconstruction data, and the output of the Decoder convolutional layer is used as the output of the Decoder unit, connecting the input of the anomaly detection module.
3. The method according to claim 1, wherein The VAE-GAN encoder includes an encoder and a decoder connected end to end. The encoder is used to generate potential anomaly information based on image data; The decoder is used to generate reconstructed image data based on the potential anomaly information.
4. The method according to claim 3, characterized in that, The encoder includes a first convolutional layer, a first activation function layer, a first pooling layer, a second convolutional layer, a second activation function layer, a second pooling layer, a Flatten layer, a first fully connected layer, a third activation function layer, and a second fully connected layer connected end to end in sequence; Among them, the first convolutional layer extracts the features of the image data, and the first activation function layer performs nonlinear processing on the output of the first convolutional layer; the first pooling layer performs a pooling operation on the output of the first encoder activation function layer to reduce the feature dimension after nonlinear processing; the second convolutional layer performs a convolutional operation on the output of the first pooling layer to extract the features of the image data; the second activation function layer performs nonlinear processing on the output of the second convolutional layer; the second pooling layer performs a pooling operation on the output of the second activation function layer; the Flatten layer converts the output of the second pooling layer into a one-dimensional feature vector; the first fully connected layer compresses the features of the one-dimensional feature vector; the third activation function layer performs nonlinear processing on the output of the first fully connected layer; the second fully connected layer performs feature mapping on the output of the third activation function layer to generate the potential anomaly information of the image data.
5. The method according to claim 4, characterized in that, The decoder includes a first transposed convolutional layer, a first activation function layer, an upsampling layer, a second transposed convolutional layer, a second activation function layer, and a decoder convolutional layer connected end to end in sequence; Among them, the first transposed convolutional layer performs a transposed convolutional operation on the potential anomaly information received from the encoder; the first activation function layer performs nonlinear processing on the output of the first transposed convolutional layer; the upsampling layer performs upsampling on the output of the first activation function layer; the second transposed convolutional layer performs a transposed convolutional operation on the output of the upsampling layer; the second activation function layer performs nonlinear processing on the output of the second transposed convolutional layer; the decoder convolutional layer performs a convolutional operation on the output of the second activation function layer to obtain the reconstructed image data.
6. The method according to claim 5, wherein The pseudo-labeler includes an input layer, a fully connected layer, an activation function layer, and an output layer connected end to end in sequence; The input layer receives the final features of the image data output by the Encoder unit and transmits them to the fully connected layer; The fully connected layer performs spatial mapping on the final features of the image data to generate the number of feature categories; The activation function layer generates the distribution probability of each feature category based on the number of feature categories through the Softmax function; The output layer generates an anomaly label based on the distribution probability of each feature category and transmits it to the anomaly detection module.
7. The method according to claim 6, wherein The anomaly detection module includes: A reconstruction deviation calculation unit for calculating the reconstruction deviation based on the image reconstruction data and the image data of the product to be detected; An anomaly deviation calculation unit for calculating the anomaly deviation based on the anomaly reconstruction data and the image data of the product to be detected; A pseudo-label deviation calculation unit for calculating the pseudo-label deviation based on the anomaly label and the true label of the pre-labeled image data of the product to be detected; A regularization deviation calculation unit for calculating the regularization deviation based on the potential anomaly information; Anomaly localization unit, which is used to perform feature fusion on reconstruction deviation, anomaly deviation, pseudo-label deviation, and regularization deviation to obtain comprehensive deviation features; retrieve the historical anomaly image with the highest similarity to the comprehensive deviation features in the memory module, and determine the position of the anomaly area in the image of the product to be detected based on the anomaly area of the historical anomaly image; Anomaly classification unit, which is used to obtain the anomaly category based on the anomaly label output by the pseudo-labeler.
8. The method according to claim 7, wherein Calculate the reconstruction deviation based on the image reconstruction data and the image data of the product to be detected as shown in calculation formula (1); where n is the number of image data of the products to be detected, and y i is the image data of the i-th product to be detected, and is the i-th image reconstruction data; Calculate the anomaly deviation based on the anomaly reconstruction data and the image data of the product to be detected as shown in calculation formula (2); where n is the number of image data of the products to be detected, and a i is the image data of the i-th product to be detected, and is the i-th abnormal reconstruction data.
9. The method according to claim 1, wherein Calculate the pseudo-label deviation based on the anomaly label and the true label of the pre-annotated image data of the product to be detected as shown in calculation formula (3); where n is the number of image data of the product to be detected, and y ic is the encoding of the true label of the i-th image data, and is the encoding of the pseudo-label of the i-th image data.
10. The method according to claim 7, wherein Calculate the regularization deviation based on the potential anomaly information as shown in calculation formula (4); L latent = D KL (q(z|xi)||p(z)) (4) Among them, D KL is the KL divergence, q(z|xi) is the potential anomaly information, and p(z) is the prior distribution of the anomaly information.
Citation Information
Patent Citations
Unified anomaly detection method based on multi-source uncertainty mining
CN118505600A
Patch feature learning method for anomaly detection, and system therefor
WO2024186178A1