Zero-sample industrial anomaly detection method based on knowledge distillation

Through the knowledge distillation of the DINO self-supervised model and the lightweight feature matching module, combining multi-scale feature fusion and text attention module, the problem of insufficient generalization ability of existing methods in unknown categories of industrial products is solved, and efficient and accurate multi-scale anomaly detection is achieved.

CN120298397AActive Publication Date: 2025-07-11NANJING UNIV OF INFORMATION SCI & TECH

Patent Information

Application Number
CN202510765559.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-07-11
Estimated Expiration
2045-06-10

AI Technical Summary

Technical Problem

The existing zero-sample anomaly detection method based on text image large models lacks generalization capabilities on industrial products of unknown categories, is difficult to adapt to anomaly detection in multi-scale and complex contexts, and lacks cross-modal interaction and deep feature modeling, resulting in insufficient detection stability and accuracy.

Method used

The DINO self-supervised model is used as the teacher model, combining the lightweight feature matching module and the text attention module, and the fusion of knowledge distillation and multi-scale feature is achieved to achieve accurate detection of abnormal areas of industrial parts.

Benefits of technology

It improves the model's abnormal detection capabilities in industrial products of unknown categories, reduces calculation costs, enhances the adaptability and robustness to abnormalities of different scales, and meets the real-time and accuracy requirements of industrial detection tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298397A_ABST
    Figure CN120298397A_ABST
Patent Text Reader

Abstract

The invention discloses a zero-sample industrial anomaly detection method based on knowledge distillation, and belongs to the technical field of image detection. The method comprises a training stage and a testing stage. In the training stage, normal industrial images and abnormal industrial images in any type of industrial products are selected; constructing a knowledge distillation network architecture which comprises a teacher model, a student model, an image encoder and a text encoder; training the teacher model by using the training set data, and transferring knowledge in the trained teacher model to the student model; and in the test stage, constructing an industrial anomaly detection model, and inputting the to-be-detected picture in the test set into the industrial anomaly detection model to obtain a detection result. According to the method, efficient zero-sample anomaly detection can be realized in industrial detection tasks with unknown categories and various anomaly forms, and a solution which is efficient, high in generalization and excellent in robustness is provided for industrial visual detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image anomaly detection method, specifically a zero-shot industrial anomaly detection method based on knowledge distillation, belonging to the field of image detection technology. Background Art

[0002] Image zero-shot anomaly detection (ZSAD) aims to accurately identify anomaly regions in unknown industrial products without relying on any pre-defined training data. With the increasing requirements for data privacy protection, the increasing cost of obtaining labeled data, and the continuous expansion of industrial product types, the application value of ZSAD in the field of industrial vision detection has become increasingly prominent and has become a current research hotspot. ZSAD methods need to have strong generalization ability to cope with the significant changes in product appearance and the diversity of anomaly targets in different industrial detection tasks. Therefore, achieving efficient anomaly detection on industrial products of unknown categories has become the core challenge in current research.

[0003] Currently, anomaly detection methods based on large text-image models mainly rely on constructing two types of text prompts, normal or abnormal, and identify anomaly regions by calculating the cosine similarity between the image and the text in the joint space. Some existing methods adopt multi-scale window or image block strategies to extract dense visual features and match normal and abnormal regions in combination with text prompts. However, these methods still have certain limitations in actual industrial applications.

[0004] First of all, many existing methods assume that the category of the product to be detected is known and construct corresponding text descriptions based on this category, such as "picture of a normal metal part" or "undamaged glass sheet". However, in actual industrial scenarios, especially when it comes to data privacy protection, new product detection, or a changing production environment, the category of the detection object is often unknown, making this way that relies on prior knowledge difficult to generalize. In addition, the vocabulary selection of text prompts will also significantly affect the detection ability of the model. For example, in experiments, simply replacing "screw" with "bolt" or "fastener" may lead to significant fluctuations in the average precision index, affecting the stability and accuracy of anomaly detection.

[0005] Secondly, existing methods often independently map images and texts during the feature modeling process, lacking cross-modal deep interaction and multi-scale feature extraction, resulting in the model relying too much on specific text prompts and insufficient generalization ability of the model in the face of different scales, complex backgrounds, or diverse anomaly categories. Especially in industrial detection tasks, anomalies may appear in the form of fine cracks, surface defects, deformations, etc., and the global matching strategy of existing methods is difficult to accurately capture these detailed features.

[0006] In summary, the current anomaly detection methods based on text-image large models still have the following problems: (1) Relying on predefined text prompts, it is difficult to adapt to industrial products of unknown categories; (2) Lack of cross-modal interaction and deep feature modeling, resulting in insufficient ability to distinguish anomaly regions; (3) Poor adaptability to multi-scale anomaly targets, and it is difficult to stably detect anomaly regions of different scales in actual industrial scenarios. Summary of the Invention

[0007] Object of the Invention: Aiming at the above problems, the object of the present invention is to provide a zero-shot industrial anomaly detection method based on knowledge distillation. The DINO self-supervised model is selected as the teacher model, and its high-quality local feature extraction ability is fully utilized, and combined with the vision-language alignment ability of the text-image large model, to achieve precise detection of anomaly regions of industrial parts.

[0008] Technical Solution: The zero-shot industrial anomaly detection method based on knowledge distillation of the present invention includes a training stage and a testing stage; The training stage includes the following steps: Select normal industrial images and abnormal industrial images in any category of industrial products, and after preprocessing, construct a training set; Construct a knowledge distillation network architecture, including a teacher model, a student model, an image encoder, and a text encoder. The teacher model uses a pre-trained DINO model, and the student model includes a lightweight feature matching module and a text attention module; Use the training set data to train the teacher model, and transfer the knowledge in the trained teacher model to the student model; The testing stage includes the following steps: Construct an industrial anomaly detection model, input the to-be-detected pictures in the test set into the industrial anomaly detection model, and obtain the detection result; the industrial anomaly detection model includes an image encoder, a student model, and a text encoder.

[0009] Furthermore, the step of using the training set data to train the teacher model and transferring the knowledge in the trained teacher model to the student model includes: Input the to-be-detected images in the training set into the image encoder to obtain four-layer to-be-detected image features , where is the first-layer image feature, is the second-layer image feature, is the third-layer image feature, is the fourth-layer image feature; Process the four-layer to-be-detected image features to obtain the corresponding correlation matrix , and the formula is: , wherein, is the transpose matrix of the matrix; Input the correlation matrix into the lightweight feature matching module for feature transformation to obtain the feature map , and the formula is: , wherein, represents being composed of two layers of normalized convolutional layers; Input the same image to be detected into the pre-trained DINO model to obtain four layers of image features to be detected , where is the first layer of image features, is the second layer of image features, is the third layer of image features, is the fourth layer of image features; Calculate the correlation matrix after taking the dot product of the four layers of features output by the DINO model , and the formula is: , wherein, is the transpose matrix of the matrix; Adopt binary cross-entropy loss BCE to calculate the difference between the correlation matrix after conversion of the image encoder output and the correlation matrix after conversion of the DINO model output. The formula is: , wherein, represents binary cross-entropy loss; Use to train the lightweight feature matching module, and select the model with the minimum loss as the final lightweight feature matching module; Process the fourth layer of image features and input them into the multi-scale feature mapping module with 3 different parallel branches to obtain the feature map , so that the dimension of the feature map is the same as that of the text encoder; Construct the normal state and abnormal state of the image to be detected into a text prompt template in a fixed format, and align the feature map to the text embedding space; Generate the abnormal coarse heat map of the image to be detected.

[0010] Furthermore, process the fourth layer of image features After processing, it is input into a multi-scale feature mapping module with 3 different parallel branches to obtain a feature map The steps include: Remove the class label representing the global feature from the fourth-layer image feature Keep only the patch-level label representing the local feature, and then resize it to a preset size to obtain a feature map ; Input the feature map Into 3 different parallel branches respectively to obtain three different feature maps 、 And , and the formulas are respectively: , , , In the formula, Represents a standard Convolution, Represents a standard Convolution; Represents a dilated Convolution, and R represents the dimensional space of the tensor; Then concatenate the original feature With the multi-scale features extracted from each branch on the channel dimension to obtain a feature map , and the formula is: , In the formula, Represents the concatenation operation on the channel dimension, Use a 1×1 convolution To reduce the dimension of the feature map To obtain a feature map , and the formula is: , Finally, apply a residual connection to add the feature map To the feature map To obtain a feature map , and the formula is: .

[0011] Furthermore, the steps to make the dimension of the feature map The same as that of the text encoder include: Flatten the feature map Back to one dimension to obtain a feature map , and the formula is: , In the formula, Indicates a flattening operation; Then, a 1×1 convolutional layer is used as a linear projection layer to perform feature transformation and obtain a feature map , and the formula is: .

[0012] Furthermore, the normal and abnormal states of the image to be detected are constructed into a text prompt template in a fixed format, and the steps of aligning the feature map to the text embedding space include: Construct a text prompt template H in a fixed format, which is expressed as: , where, is the fixed text part, represents the state of the object, including good and damaged, corresponding to normal and abnormal; represents r learnable class tokens; Generate normal and abnormal texts through the text prompt template H, which are respectively expressed as: , , Map the normal and abnormal texts to the embedding space to obtain feature maps , where C represents the embedding dimension of the text encoder, represents the text length; Reduce the dimension of the feature map using global average pooling to extract global features and obtain a feature map , and the formula is: , where, represents the global average pooling operation; Align is expressed as , and then update the class token by adding residuals, which is expressed as: , where, represents the token randomly generated during initialization; The updated class token is re-injected into the normal and abnormal texts to form visually guided texts, which are respectively expressed as: , , where, Indicates the normal text after visual guidance, Indicates the abnormal text after visual guidance; Map and to the embedding space through the text encoder to obtain the final normal text features and abnormal text features , which are respectively expressed as: , , In the formula, Indicates the linear transformation operation, Indicates the normalization operation.

[0013] Furthermore, the steps to generate the abnormal coarse heat map of the image to be detected include: For the four-layer image features to be detected , remove the class label and only retain the patch label to obtain the feature map , which is expressed as: , In the formula, Indicates removing the class label of the image features to be detected; Then perform a linear projection on the four-layer feature map to obtain the projected feature map , which is expressed as: , For each layer of the projected feature map , calculate the cosine similarity between each patch and the normal text feature and the abnormal text feature and , which is expressed as: , , For each layer of image features, calculate the difference between the normal and abnormal matching scores to obtain the abnormal score matrix for each layer , which is expressed as: , Next, generate multi-layer heat maps and perform a merging operation: First, reshape the abnormal score matrix for each layer into a two-dimensional matrix of , which is expressed as: , In the formula, Indicates the matrix size reshaping operation; Then, the four - layer anomaly score matrix is fused by taking the mean, which is expressed as: , Finally, in order to map the patch - level anomaly score to the original image space, bilinear interpolation is required, which is expressed as: , In the formula, represents the upsampling operation, represents the original size of the image, specifies that the method adopted for upsampling is bilinear interpolation; The heatmap is normalized, and the values of the heatmap are restricted within the range, which is expressed as: , In the formula, min represents taking the minimum value of the matrix, and max represents taking the maximum value of the matrix; Set the threshold T 0 to generate the anomaly segmentation mask: , In the formula, represents the anomaly segmentation mask at position , represents the anomaly score at the (x, y) position of the normalized anomaly heatmap, T 0 is the threshold for anomaly detection; The anomaly segmentation mask is superimposed on the original image to obtain the final anomaly detection result: , In the formula, represents the original input image, and the symbol represents element - by - element multiplication, which is used to highlight the anomaly area; Finally, calculate the prediction loss. Select the ground - truth mask as the supervision signal of the multi - scale feature mapping module, which represents the true anomaly situation of each pixel in the image, E represents the height of the image, W represents the width of the image, represents the anomaly area, represents the normal area; For the L - th image in a batch, the total loss of the anomaly mask is: , In the formula, represents the Focal Loss, represents ; Sum the losses of all images in a batch to obtain the final total loss function as follows: , In the formula, batchsize represents the number of samples used for model training or inference in one iteration; Use the loss function to train the multi-scale feature mapping module, and select the model with the minimum loss as the final multi-scale feature mapping module.

[0014] Beneficial effects: Compared with the prior art, the significant advantages of the present invention are as follows: 1. The present invention mainly aims at the industrial anomaly detection task and proposes a zero-shot detection method based on knowledge distillation to improve the anomaly detection ability of the model on industrial products of unknown categories. This method depends on normal and abnormal images in the training stage, but does not require a dataset of specific detection targets. Instead, it uses other datasets, such as training on the VisA dataset and testing on the MVTec dataset, so that the model has stronger generalization ability and can adapt to the anomaly detection needs of different industrial products; 2. This method makes full use of the high-quality local features generated by the DINO model and optimizes the lightweight module through knowledge distillation, enabling the model to have stronger anomaly recognition ability. In the inference stage, it is not necessary to load the complete DINO pre-trained model, which reduces the computational cost while improving the detection efficiency; 3. The present invention introduces a multi-scale feature fusion module RMF to enhance the model's adaptability to anomalies of different scales, and updates the fine-grained text embedding through a vision-guided text attention module VTAM to make the representation of the anomaly region more accurate; 4. In the present invention, a strategy of combining rough anomaly detection and fine anomaly detection is adopted. First, calculate the similarity between the image and the text through the features obtained by the image encoder to generate a rough anomaly detection result, and then use the lightweight feature matching module and VTAM to further optimize the anomaly region perception ability. Finally, fuse the multi-level detection results to improve the robustness and generalization ability of anomaly detection; 5. Since the training data of this method is not limited to specific industrial products, but improves the adaptability of the model through cross-dataset training, the present invention can achieve efficient and accurate anomaly detection in an industrial background with unknown categories and diverse anomaly forms, significantly reducing the dependence on abnormal samples of target products and meeting the dual requirements of real-time performance and accuracy for industrial detection tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 is a framework diagram of the method of the present invention in the training stage; Figure 2It is the flowchart of the method described in the present invention in the training stage; Figure 3 It is the schematic diagram of the multi-scale feature mapping module; Figure 4 It is the schematic diagram of the text attention module; Figure 5 It is the framework diagram of the method described in the present invention in the testing stage; Figure 6 It is the flowchart of the method described in the present invention in the testing stage. Detailed implementation manner

[0016] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments.

[0017] The zero-shot industrial anomaly detection method based on knowledge distillation described in this embodiment includes a training stage and a testing stage; The training stage includes the following steps: Select normal industrial images and abnormal industrial images from any category of industrial products, and after preprocessing, construct a training set; Construct a knowledge distillation network architecture, including a teacher model, a student model, an image encoder, and a text encoder, where the teacher model uses a pre-trained DINO model, and the student model includes a lightweight feature matching module and a text attention module; Use the training set data to train the teacher model, and transfer the knowledge in the trained teacher model to the student model; The testing stage includes the following steps: Construct an industrial anomaly detection model, input the to-be-detected pictures in the test set into the industrial anomaly detection model to obtain detection results; where the industrial anomaly detection model includes an image encoder, a student model, and a text encoder.

[0018] Combined with Figures 1 to 2 As shown, in the training stage, select normal industrial images and abnormal industrial images from any category of industrial products, and after preprocessing, construct a training set.

[0019] In this example, the publicly available MVTec AD color image set is selected as the data set. This data set covers 15 categories of industrial products and is divided into a training set and a test set, both of which contain normal and abnormal images. The zero-shot evaluation is carried out on the VISA data set.

[0020] 16 training images can be selected from the training set at a time, and the size of each picture is changed to pixels, 518 is the length and height, 3 is the number of channels, this size is the input requirement of the image encoder, and the normalization parameters are the mean and the standard deviation and perform random rotation 、horizontal flipping, and color jitter such as brightness ±0.2, contrast ±0.2, and saturation ±0.2 to enhance data diversity, and finally obtain the preprocessed input image.

[0021] Furthermore, the steps of training the teacher model with the training set data and transferring the knowledge in the trained teacher model to the student model include: Input the images to be detected in the training set into the image encoder, and the image encoder adopts a transformer structure to obtain four layers of image features to be detected , where is the first layer of image features, is the second layer of image features, is the third layer of image features, is the fourth layer of image features; Each layer of image features to be detected consists of class tokens of global information and patch-level tokens of local information, and the feature size is , where 1370 represents the number of patches, 16 is the number of input images, and 1024 represents the embedding dimension of each patch feature; Process the four layers of image features to be detected. First, remove the global information in each feature and only retain the local information; then, reshape the feature into , and finally calculate the corresponding correlation matrix for each feature , and the formula is: , In the formula, is the transpose matrix of the Input the correlation matrix into the lightweight feature matching module for feature transformation to obtain the feature map , in order to learn to map the features extracted by the image encoder to the feature space in the style of the DINO model, and the formula is: , In the formula, represents being composed of two layers of normalized convolutional layers, including activation, keeping the number of channels unchanged to enhance the local feature expression ability and at the same time match the output style of the DINO model; Reshape the feature into and add the class tokens of global information to make its final output the correlation matrix after dot product with the four layers of features of DINO respectively Keep consistent; Input the same image to be detected into the pre-trained DINO model to obtain four layers of image features to be detected , where is the first layer of image features, is the second layer of image features, is the third layer of image features, is the fourth layer of image features; Calculate the correlation matrix of the dot products of the four layers of features output by the DINO model , and the formula is: , In the formula, is the transpose matrix of the matrix; Use the binary cross-entropy loss BCE to calculate the difference between the correlation matrix after conversion of the output of the image encoder and the correlation matrix after conversion of the output of the DINO model. The formula is: , In the formula, represents the binary cross-entropy loss; Use to train the lightweight feature matching module, and select the model with the smallest loss as the final lightweight feature matching module; Input the fourth layer of image features , whose shape is into the multi-scale feature mapping module with 3 different parallel branches after processing to obtain the feature map , and make the dimension of the feature map the same as that of the text encoder; Construct the normal state and abnormal state of the image to be detected into a text prompt template in a fixed format, and align the feature map to the text embedding space; Generate the abnormal coarse heat map of the image to be detected.

[0022] In the example, the image encoder uses the cross-modal model CLIP trained by contrastive learning and is used as the basic model in the zero-shot anomaly segmentation task. In the training stage of this method, use the publicly available pre-trained image encoder, freeze all its parameters, and extract features from the image to be detected.

[0023] The DINO model is also a pre-trained model. Through self-distillation and contrastive learning, it achieves efficient image representation learning. DINO is trained on unlabeled data, capable of capturing subtle differences in images and generating high-quality local features. During the training phase of this method, the pre-trained model of DINO freezes all parameters and is used to extract high-quality features of industrial images. These features are then distilled into a lightweight feature matching module to enhance the learning ability of the lightweight module for the DINO model. It should be noted that in the testing phase, the DINO model is no longer used, and instead, the lightweight feature matching module trained in the training phase is used to replace its function.

[0024] The lightweight feature matching module consists of two 3 ×3 convolutional blocks, which are used to learn to represent the image features output by the image encoder as high-quality patch-level features close to the output of the DINO model. It replaces the DINO model in the inference phase, reducing the hardware cost and time cost of inference.

[0025] Combined Figure 3 As shown, further, the steps of processing the fourth-layer image features and inputting them into the multi-scale feature mapping module with 3 different parallel branches to obtain the feature map include: Removing the class token representing the global feature from the fourth-layer image features and only retaining the patch-level token representing the local feature, and then resizing it to a preset size to obtain the feature map ; Inputting the feature map into 3 different parallel branches respectively to obtain three different feature maps 、 and , and the formulas are respectively: , , , wherein, represents a standard convolution, represents a standard convolution; represents a dilated convolution, and R represents the dimensional space of the tensor; Then, the original feature is concatenated with the multi-scale features extracted from each branch in the channel dimension to obtain the feature map , and the formula is: , wherein, Represents the concatenation operation in the channel dimension, uses a 1×1 convolution on the feature map to reduce the dimension and obtain the feature map , and the formula is: , Finally, apply a residual connection to add the feature map and the feature map to obtain the feature map , and the formula is: .

[0026] Furthermore, the steps to make the dimension of the feature map the same as that of the text encoder include: To embed the image features into the semantic space of the text encoder, flatten the feature map to one dimension to obtain the feature map , and the formula is: , wherein, represents the flattening operation; Then, perform feature transformation through a 1×1 convolutional layer as a linear projection layer to obtain the feature map , and the formula is: , where 768 is the output dimension of the text encoder, enabling the image features to be seamlessly aligned to the text embedding space for subsequent multimodal information fusion.

[0027] Furthermore, the steps to construct the normal and abnormal states of the image to be detected into a fixed-format text prompt template and align the feature map to the text embedding space include: Construct a fixed-format text prompt template H, which is expressed as: , wherein, is the fixed text part, represents the state of the object, including good and damaged, corresponding to normal and abnormal; represents r learnable class tokens for learning information about specific product categories, and is fixed at 4; Generate normal and abnormal texts through the text prompt template H, which are respectively expressed as: , , Map the normal text and abnormal text to the embedding space to obtain feature maps respectively , where C represents the embedding dimension of the text encoder, represents the text length, including the fixed template, status words, and learnable class tokens; For the feature map Perform dimensionality reduction using global average pooling to extract global features and obtain the feature map , and the formula is: , In the formula, represents the global average pooling operation; For is represented as , and then update the class token by adding residuals , which is represented as: , In the formula, represents the token randomly generated during initialization; The updated class token is re-injected into the normal text and abnormal text to form the visually guided text, which are respectively represented as: , , In the formula, represents the visually guided normal text, represents the visually guided abnormal text; For and Map them to the embedding space through the text encoder to obtain the final normal text feature and abnormal text feature , which are respectively represented as: , , In the formula, represents the linear transformation operation, represents the normalization operation.

[0028] Furthermore, the steps to generate the abnormal coarse heat map of the image to be detected include: For the four-layer features of the image to be detected , remove the class tokens and only retain the patch tokens to obtain the feature map , which is represented as: , In the formula, Indicates removing the class label of the image features to be detected; Then perform a linear projection on the four-layer feature map to obtain a projected feature map which is expressed as: , For each layer of the projected feature map , calculate the cosine similarity between each patch and the normal text feature and the abnormal text feature , which is expressed as: and , which is expressed as: , , For each layer of image features, calculate the difference between the normal and abnormal matching scores to obtain the abnormal score matrix for each layer , which is expressed as: , When is small, it indicates that the patch is closer to the abnormal feature, suggesting that this area may be an abnormal area; otherwise, it may be a normal area; Next, generate multi-layer heatmaps and perform a merging operation: First, reshape the abnormal score matrix for each layer into a two-dimensional matrix of , which is expressed as: , wherein, represents the matrix size reshaping operation; Then, perform a mean fusion on the four-layer abnormal score matrices, which is expressed as: , Finally, to map the patch-level abnormal scores to the original image space, bilinear interpolation is required, which is expressed as: , wherein, represents the upsampling operation, represents the original size of the image, specifies that the method used for upsampling is bilinear interpolation; Normalize the heatmap and limit the values of the heatmap within the range of , which is expressed as: , wherein, min represents taking the minimum value of the matrix, and max represents taking the maximum value of the matrix; Set the thresholdT Generate an abnormal segmentation mask with 0: , wherein, represents the abnormal segmentation mask at position , represents the abnormal score of the normalized abnormal heat map at the (x, y) position, which is used to measure the possibility of belonging to the abnormal area at the (x, y) position, T 0 is the threshold for anomaly detection; Overlay the abnormal segmentation mask on the original image to obtain the final anomaly detection result: , wherein, represents the original input image, and the symbol represents element-wise multiplication, which is used to highlight the abnormal area; Finally, calculate the prediction loss, and select the ground truth mask as the supervision signal of the multi-scale feature mapping module, which represents the true abnormal situation of each pixel in the image, E represents the height of the image, W represents the width of the image, represents the abnormal area, represents the normal area; For the L-th image in a batch, the total loss of the abnormal mask is: , wherein, represents the Focal Loss, represents ; Sum the losses of all images in a batch to obtain the final total loss function: , wherein, batchsize represents the number of samples used for model training or inference in one iteration; Use the loss function to train the multi-scale feature mapping module, and select the model with the minimum loss as the final multi-scale feature mapping module.

[0029] Among them, the threshold for anomaly detection can be adjusted according to the change of the dataset. The specific method is as follows: First, use the function precision_recall_curve provided by the scikit-learn library to calculate the prediction rate (Precision) and recall rate (Recall) at different thresholds, and then calculate the threshold , the specific formula is: , The formula for Focal Loss is: , In the formula, is 's focusing parameter, set to 2.0; represents the number of all pixels in the mask.

[0030] 's formula is: , In the formula, is the smoothing term, set to to prevent the denominator from being zero.

[0031] In summary, use to train the lightweight feature matching module, and use to train the multi-scale feature mapping module. Train for 10 rounds, set the batchsize to 16, set the learning rate to 0.0001, and select the model with the smallest loss as the final lightweight feature matching module and multi-scale feature mapping module.

[0032] The test phase includes the following steps: Construct an industrial anomaly detection model, input the images to be detected in the test set into the industrial anomaly detection model, and obtain the detection results; among them, the industrial anomaly detection model includes an image encoder, a student model, and a text encoder.

[0033] Combined with Figures 5 to 6 as shown, the specific implementation process of the test phase is as follows: S1, Read in the images for testing, and uniformly adjust their sizes to pixels, as the input images in the test phase .

[0034] S2, Send the preprocessed images into the pre-trained image encoder and the student model, that is, the lightweight feature matching module, to extract four layers of features and the correlation matrix .

[0035] S3, Send into the multi-scale feature mapping module to obtain that can be embedded into text features.

[0036] S4, Send the unified text template into the text encoder, and after a series of operations, obtain the fused normal and abnormal text embeddings .

[0037] S5, Combine the text embedding with the image features to obtain a rough anomaly heat map .

[0038] S6, Use the four-layer features extracted by the image encoder and the correlation matrix , and obtain the image fusion features through element-wise sum operation , and send them together with the concatenated text embedding into the VTAM vision-guided text attention module

[0039] S7, In the text attention module, combine Figure 4 , first denote the concatenated text features as project them into the Query space through a 1D convolutional layer to obtain the query features , and the formula is , then reshape the dimension to dimensions, where represents the 1D convolutional layer is the number of attention heads is the dimension of the projected features is the dimension of each head ; B represents the batch size then use a 1D convolutional layer to project the image fusion features into the Key and Value spaces respectively, and the formula is , , In the formula represents the key projected from the image features represents the 1D convolutional weight matrix that maps the image features to the Key space represents the bias term of the Key space mapping represents the value projected from the image features represents the 1D convolutional weight matrix that maps the image features to the Value space represents the bias term of the Value space mapping reshape to form the multi-head dimension .

[0040] Then calculate the similarity between Query and Key through dot product, and the formula is , In the formula is a learnable scaling factor, which is set to 1 at the start of training; This is used as the scaling dimension to prevent the dot product result from being too large.

[0041] Then use to normalize the attention weights. The normalized attention matrix reflects the matching degree between the text embedding and the image token.

[0042] Then use the attention weights to perform a weighted sum on Value. The formula is: , Then reshape it into the original feature dimension .

[0043] Then perform L2 normalization on the updated features to keep the feature lengths consistent. The normalized features serve as the finally optimized text features.

[0044] Finally, after multi-modal interaction, the optimized image features will be calculated for cosine similarity with the concatenated and optimized text embedding to generate a similarity matrix. After removing the class tokens, the patch-level token information similarity is obtained and reshaped into an anomaly heatmap . The heatmap is upsampled to the original image size through bilinear interpolation to obtain a fine prediction result, which is combined with the rough predicted heatmap in a 1:1 ratio and normalized to obtain the final prediction result.

[0045] S8, finally obtaining the predicted anomaly heatmap . By setting a threshold it is converted into a binary mask for identifying the anomaly region. Among them, the black region is the normal region, and vice versa. If there is no white region in the whole picture, then the picture has no anomaly.

[0046] This method is trained on the MvTecAD dataset and the VisA dataset, and tested on the other dataset. Two metrics, the area under the pixel-level curve (AUROC) and the pixel-level region overlap rate (PRO), are calculated respectively. Among them, the AUROC of the model trained on the VisA dataset and tested on the MvTecAD dataset is 92.2, and the PRO is 88.1; the AUROC of the model trained on the MvTecAD dataset and tested on the VisA dataset is 96.0, and the PRO is 91.3. The results show that the performance of this method is better than the zero-shot anomaly detection method that only relies on text prompts.

Claims

1. A zero-shot industrial anomaly detection method based on knowledge distillation, characterized in that It includes a training phase and a testing phase; The training phase includes the following steps: Select normal industrial images and abnormal industrial images from any category of industrial products. After preprocessing, construct a training set; Construct a knowledge distillation network architecture, including a teacher model, a student model, an image encoder, and a text encoder. Among them, the teacher model uses a pre-trained DINO model, and the student model includes a lightweight feature matching module and a text attention module; Use the training set data to train the teacher model, and transfer the knowledge in the trained teacher model to the student model; The testing phase includes the following steps: Construct an industrial anomaly detection model, input the images to be detected in the test set into the industrial anomaly detection model to obtain the detection results; among them, the industrial anomaly detection model includes an image encoder, a student model, and a text encoder.

2. The zero-shot industrial anomaly detection method based on knowledge distillation according to claim 1, wherein The steps of using the training set data to train the teacher model and transferring the knowledge in the trained teacher model to the student model include: Input the image to be detected in the training set into the image encoder to obtain four layers of image features to be detected , where is the first-layer image feature, is the second-layer image feature, is the third-layer image feature, is the fourth-layer image feature; Process the features of the four-layer image to be detected to obtain the corresponding correlation matrix , and the formula is: , In the formula, is the transpose matrix of the matrix; Input the correlation matrix into the lightweight feature matching module for feature transformation to obtain a feature map , and the formula is: , In the formula, represents being composed of two normalized convolutional layers; Input the same image to be detected into the pre-trained DINO model to obtain four layers of image features to be detected , where is the first-layer image feature, is the second-layer image feature, is the third-layer image feature, is the fourth-layer image feature; Calculate the correlation matrix after taking the dot product of the four layers of features output by the DINO model , and the formula is: , In the formula, is the transpose matrix of the matrix; Use binary cross-entropy loss BCE to calculate the difference between the correlation matrix after the output conversion of the image encoder and the correlation matrix after the output conversion of the DINO model. The formula is: , In the formula, represents the binary cross-entropy loss; Utilize Train a lightweight feature matching module and select the model with the minimum loss as the final lightweight feature matching module; Process the fourth-layer image features and input them into a multi-scale feature mapping module with 3 different parallel branches to obtain a feature map , such that the dimension of the feature map is the same as that of the text encoder; Construct a text prompt template in a fixed format for the normal and abnormal states of the image to be detected, and align the feature map to the text embedding space; Generate a rough anomaly heat map of the image to be detected.

3. The zero-shot industrial anomaly detection method based on knowledge distillation according to claim 2, characterized in that, Process the fourth-layer image features and input them into a multi-scale feature mapping module with 3 different parallel branches to obtain feature maps The steps are as follows: Remove the class label representing the global feature from the fourth-layer image feature and only retain the patch-level label representing the local feature, and then resize it to a preset size to obtain a feature map ; Input the feature map into three different parallel branches respectively to obtain three different feature maps , and , and the formulas are respectively: , , , In the formula, denotes standard convolution, denotes standard convolution; denotes dilated convolution, and R represents the dimensional space of the tensor; Then, the original features are concatenated in the channel dimension with the multi-scale features extracted by each branch to obtain a feature map , and the formula is: , In the formula, represents the concatenation operation in the channel dimension. Use a 1×1 convolution on the feature map to reduce the dimension and obtain the feature map , and the formula is: , Finally, apply a residual connection to add the feature map and the feature map to obtain the feature map , and the formula is: 。 4. The zero-shot industrial anomaly detection method based on knowledge distillation according to claim 3, characterized in that Steps to make the dimensionality of the feature map the same as that of the text encoder include: Flatten the feature map to one dimension again to obtain the feature map , and the formula is: , In the formula, represents a flattening operation; Then, a 1×1 convolutional layer is used as a linear projection layer to perform feature transformation and obtain a feature map , and the formula is: 。 5. The zero-shot industrial anomaly detection method based on knowledge distillation according to claim 4, wherein, Construct the normal state and abnormal state of the image to be detected into a text prompt template in a fixed format, and align the feature map The steps of aligning to the text embedding space include: Construct a text prompt template H in a fixed format, expressed as: , In the formula, is the fixed text part, represents the state of the object, including good and damaged, corresponding to normal and abnormal; represents r learnable class labels; Generate normal through the text prompt template H and abnormal text , which are respectively expressed as: , , Map the normal text and abnormal text to the embedding space to obtain the feature maps respectively , where C represents the embedding dimension of the text encoder represents the text length The feature map is dimensionally reduced using global average pooling to extract global features, resulting in the feature map , and the formula is: , In the formula, represents the global average pooling operation; Represent as and then update the category label by residual summation which is expressed as: ​ , In the formula, represents a flag randomly generated during initialization; The updated class labels are reinjected into the normal text and abnormal text to form visually guided text, respectively expressed as: , , In the formula, represents the normal text after visual guidance, represents the abnormal text after visual guidance; Map and to the embedding space through the text encoder to obtain the final normal text features and abnormal text features , which are respectively represented as: , , In the formula, represents a linear transformation operation, represents a normalization operation.

6. The zero-shot industrial anomaly detection method based on knowledge distillation according to claim 5, characterized in that The steps of generating a rough anomaly heat map of the image to be detected include: For the four-layer image features to be detected , remove the class labels and only retain the patch labels to obtain the feature map , which is expressed as: , In the formula, represents the class label for removing the features of the image to be detected; Next, perform a linear projection on the four-layer feature map to obtain a projected feature map , which is expressed as: , For each layer of the projected feature map calculate the cosine similarity between each patch and the normal text feature and the abnormal text feature which is expressed as and is expressed as: , , For each layer of image features, calculate the difference between the normal and abnormal matching scores to obtain the abnormal score matrix for each layer , denoted as: , Next, generate multi-layer heat maps and perform a merging operation: First, the anomaly score matrix of each layer is reshaped into a two-dimensional matrix, denoted as: , In the formula, represents the matrix size reshaping operation; Then, perform mean fusion on the four-layer anomaly score matrices, expressed as: , Finally, in order to map the patch-level anomaly scores to the original image space, bilinear interpolation is required, expressed as: , In the formula, represents the upsampling operation, represents the original size of the image, specifies that the method used for upsampling is bilinear interpolation; Normalize the heat map and limit the values of the heat map within the range, expressed as: , In the formula, min represents taking the minimum value of the matrix, and max represents taking the maximum value of the matrix; Set a threshold T equal to 0 to generate an abnormal segmentation mask: , In the formula, represents the abnormal segmentation mask at the position , represents the abnormal score of the normalized abnormal heat map at the (x, y) position, T and 0 is the threshold for anomaly detection; Overlay the anomaly segmentation mask on the original image to obtain the final anomaly detection result: , In the formula, represents the original input image, and the symbol represents element-wise multiplication, which is used to highlight the abnormal area; Finally, calculate the prediction loss and select the ground truth mask as the supervision signal for the multi-scale feature mapping module, representing the ground truth anomalies of each pixel in the image, E represents the height of the image, W represents the width of the image, represents the abnormal region, represents the normal region; For the L-th image in a batch, the total loss of the anomaly mask is: , In the formula, represents Focal Loss, represents ; Sum the losses of all images in a batch to obtain the final total loss function as: , In the formula, batchsize represents the number of samples used for model training or inference in one iteration; Using a loss function Train the multi-scale feature mapping module and select the model with the minimum loss as the final multi-scale feature mapping module.

Citation Information

Patent Citations

  • Industrial anomaly detection method and system based on multi-scale feature guidance and fusion

    CN117710757A

  • Knowledge distillation-based anomaly detection convolution encoder and anomaly detection method

    CN119205687A

  • Knowledge distillation-based unsupervised industrial anomaly detection method and system

    CN119991555A

Cited By

  • AI-based bottle cap 360-degree full-view-angle defect high-speed detection method and system

    CN120451176A

  • Visual self-supervised learning-based high-altitude operation personnel operation specification detection method

    CN121354034A

  • Abnormality detection method and system for compensation device

    CN121599976A

  • Industrial anomaly detection method based on correlation percentage and adaptive screening strategy

    CN121686067A

  • Aircraft skin surface anomaly detection method and system independent of defect sample and storage medium

    CN122265288A