Industrial image anomaly detection method based on noise suppression modal fusion alignment
Through the noise suppression modal fusion alignment method, the problems of insufficient text semantic positioning and noise interference in industrial image anomaly detection are solved, and more accurate abnormal area recognition and generalization capabilities are achieved.
Patent Information
- Application Number
- CN202510767310.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-10
AI Technical Summary
The existing industrial image anomaly detection methods lack spatial positioning capabilities in text semantic guidance modules, the multimodal interaction mechanism has information bottlenecks, and are susceptible to noise interference, resulting in inaccurate detection results.
The noise suppression modal fusion alignment method is adopted, position information is introduced through the text prompt template, noise is suppressed using a dual-channel image encoder, and information interaction and alignment between text and image modes is promoted in combination with a multimodal interaction module.
The model's spatial perception ability of abnormal areas is enhanced, false detection and missed detection is reduced, and the accuracy and robustness of detection are improved.
Smart Images

Figure CN120298399A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an industrial image anomaly detection method, specifically an industrial image anomaly detection method for noise suppression modal fusion alignment, belonging to the field of image detection technology. Background Art
[0002] Image anomaly detection is a key task in computer vision, aiming to identify abnormal regions or objects in images. In the industrial field, it is mainly used for product quality control and defect detection, such as detecting whether electronic components are damaged. Its accuracy is crucial: high-precision detection can prevent defective products from entering the market, reduce losses and maintain the reputation of enterprises; at the same time, it reduces manual intervention and improves production efficiency and competitiveness. With the digital transformation of industry, data security and privacy protection have become increasingly important. Zero-shot learning methods use pre-trained models and a small amount of labeled data to reduce the labeling cost and the risk of data leakage, promoting the intelligentization, high efficiency and data security of industrial production.
[0003] In recent years, multi-modal pre-training technologies based on contrastive learning have achieved important breakthroughs in the field of computer vision. Such methods construct a joint semantic space of language and vision, enabling the model to directly understand image content through text semantics and demonstrating significant zero-shot reasoning capabilities. Their core innovation lies in breaking through the limitations of traditional single-modal learning and mapping image features and text descriptions to a unified representation space through a cross-modal alignment mechanism. This technical route effectively alleviates the dependence on a large amount of labeled data and demonstrates excellent domain adaptation capabilities in tasks such as image classification and object recognition, especially suitable for the pain points of scarce abnormal samples and high labeling costs in industrial quality inspection scenarios. At the same time, the characteristics of multi-modal interaction provide a new paradigm for constructing privacy-friendly detection systems by replacing the processing of raw data with semantic-level feature extraction.
[0004] Currently, although the anomaly detection methods based on this paradigm have made progress in text semantic guidance mechanisms and visual feature modeling, they still face multiple technical challenges: First, the text semantic guidance module generally lacks spatial positioning capabilities, and the existing prompt template design is difficult to effectively integrate position prior knowledge. Since the original framework focuses on global semantic matching, the local region features and text descriptions fail to achieve fine-grained alignment. Second, there is an information bottleneck in the cross-modal interaction mechanism. The text encoding process is easily interfered by redundant vocabulary, and the attention allocation mechanism fails to fully focus on key semantic units. Third, complex background interference in industrial scenarios (such as metal reflection, surface texture, etc.) poses a severe challenge to visual feature extraction. The current methods are not perfect in noise suppression and feature decoupling, and are prone to interference from pseudo anomalies in the detection results.
[0005] In summary, the existing anomaly detection algorithms have the following defects: (1)General text prompts lack specific descriptions of defect locations and cannot fully utilize text semantic information.
[0006] (2)There are deficiencies in the information interaction between the text modality and the image modality, and the potential of multimodal data cannot be fully utilized. The semantic alignment between text embeddings and image embeddings is insufficient, restricting the model's ability to understand multimodal semantics, thus affecting the detection accuracy and generalization performance.
[0007] (3)Both the image and text modalities are affected by noise interference. Background noise in industrial images, such as light changes and texture interference, can cause the model to over-focus, resulting in false detections and missed detections; while the text encoder is prone to assigning too high weights to irrelevant semantics, distracting the model's attention and affecting the accuracy of anomaly detection. Summary of the Invention
[0008] Object of the Invention: Aiming at the above problems, the object of the present invention is to provide an industrial image anomaly detection method with noise suppression, modality fusion, and alignment. By introducing location information through a text prompt template, it fully utilizes text semantic information; a noise suppression attention mechanism is used in the text encoder and the image encoder, and a dual-channel image encoder is introduced to suppress irrelevant noise and pay more attention to key semantic information and target object features.
[0009] Technical Solution: An industrial image anomaly detection method with noise suppression, modality fusion, and alignment of the present invention includes the following steps: Obtain normal industrial image samples and abnormal industrial image samples with labels, construct a dataset, and divide it into a training set and a test set according to a ratio; Construct an industrial image anomaly detection model; wherein, the industrial image anomaly detection model includes a text encoder, a dual-channel image encoder, an MLP module, and a multimodal interaction module; Use the training set to train the industrial image anomaly detection model; Use the trained industrial image anomaly detection model to perform anomaly detection on the image samples in the test set.
[0010] Furthermore, the step of using the training set to train the industrial image anomaly detection model includes: Construct a structured normal text prompt template and a structured abnormal text prompt template; Input the normal text prompt template and five groups of abnormal text prompt templates with location information into the text encoder for encoding. The normal text prompt template is encoded through a noise suppression attention layer to output a normal text embedding vector , and the five groups of abnormal text prompt templates with location information are respectively encoded into a group of abnormal text embedding vectors through a noise suppression attention layer. Mean pooling operation is performed on the five groups of abnormal text embedding vectors to generate a fused abnormal text embedding vector Concatenate the normal text embedding vector and the fused abnormal text embedding vector to obtain a text prompt embedding vector , , , where C is the embedding dimension and R represents the real matrix space; Adjust the images to be detected in the training set to a fixed resolution and then input them into a two-channel image encoder to obtain global image embedding vectors, four-layer patch-level embedding vectors, and four-layer denoised patch-level embedding vectors; Project both the four-layer patch-level embedding vectors and the four-layer denoised patch-level embedding vectors through their respective MLP modules; Use a multimodal interaction module to interact and align the projected four-layer patch-level embedding vectors and four-layer denoised patch-level embedding vectors with the text prompt embedding vector respectively; Calculate the cosine similarities between the four-layer patch-level embedding vectors, four-layer denoised patch-level embedding vectors, and global image embedding vectors after interaction and alignment with the text prompt embedding vector respectively to generate anomaly maps, calculate the loss between the obtained anomaly maps and the true anomaly maps, and train the detection ability of the industrial image anomaly detection model through loss backpropagation.
[0011] Furthermore, the text encoder consists of 12 noise suppression attention layers.
[0012] Furthermore, in the noise suppression attention layer, multiply the input term by a weight matrix to generate a query matrix and a key matrix , divide the generated matrices into two parts, denoted as , where N represents the sequence length; Calculate the first attention weight matrix and the second attention weight matrix , where scale is a scaling factor, is an activation function; Perform a difference operation through the first attention weight matrix and the second attention weight matrix to generate a denoised attention output, expressed as: , In the formula, is a learnable weight coefficient, is a hyperparameter, is the value matrix generated by multiplying the text template by the weight matrix .
[0013] Furthermore, the step of adjusting the images to be detected in the training set to a fixed resolution and then inputting them into the two-channel image encoder includes: Input the image to be detected into the first-channel CLIP image encoder. Extract image features through the QKV attention mechanism from the 1st layer to the 5th layer, and extract image features through the VV attention mechanism from the 6th layer to the 24th layer, and output the global image embedding vector. and the patch-level embedding vectors of the 6th, 12th, 18th, and 24th layers , where i = 1, 2, 3, 4 correspond to the [6th, 12th, 18th, 24th] layers respectively; Input the image to be detected into the second-channel CLIP image encoder, and output the denoised patch-level embedding vectors of the 6th, 12th, 18th, and 24th layers through the noise suppression attention layer. .
[0014] Furthermore, the steps of projecting the four-layer patch-level embedding vectors and the four-layer denoised patch-level embedding vectors through their respective MLP modules include: The patch-level embedding vectors and the denoised patch set embedding vectors are respectively fed into their respective multi-layer perceptrons for mapping operations to align the embedding vectors, and the mapped patch-level embedding vectors are respectively denoted as ; where the multi-layer perceptron includes an input layer, an intermediate layer, and an output layer.
[0015] Furthermore, the steps of interacting and aligning the projected four-layer patch-level embedding vectors and the four-layer denoised patch-level embedding vectors with the text prompt embedding vectors through the multi-modal interaction module include: The text prompt embedding vector is projected into the Key space through a learnable weight matrix to obtain the matrix , and then the text prompt embedding vector is projected into the Value space through a learnable weight matrix to obtain the matrix ; The preliminarily aligned patch-level embedding vectors are respectively projected into the Query space through their respective learnable weight matrices to obtain ; Calculate the attention weights respectively, and the formula is: , , where ; Scale and normalize the attention weights , and the formula is: , where, , is the dimension of; Multiply the attention weights by the matrix V and pass through a linear projection layer to obtain the mixed four-layer patch-level embedding vectors. The formula is: , , , , Add the corresponding elements of the patch-level embedding vectors and the mixed patch-level embedding vectors and normalize to obtain the finally aligned patch-level embedding vectors .
[0016] Furthermore, calculate the cosine similarities between the four-layer patch-level embedding vectors, four-layer denoised patch-level embedding vectors, and global image embedding vectors after interactive alignment and the text prompt embedding vector respectively. The steps to generate the anomaly map include: Calculate the cosine similarity between the patch-level embedding vectors and the text prompt embedding vector respectively. The calculation formula is: , where represents the number of patches in each layer. In each layer, the image is segmented into 1370 patches, and each patch is converted into an embedding vector; The similarity matrix obtained after calculating the four-layer patch-level embedding vectors and the text prompt embedding vector is: ; The first patch records the global image information. When calculating the local part, ignore the similarity of the first patch and use the similarities of 1369 patches for upsampling. The obtained similarity map is: , where , represents the similarity matrix after removing the first patch, denotes upsampling; Calculate the similarity between the denoised patch-level embedding vectors and the text prompt embedding vector and perform upsampling. The obtained similarity map is denoted as ; Calculate the cosine similarity between the global image embedding vector and the text prompt embedding vector to obtain the similarity matrix of the global image and the text prompt as , the calculation formula is: , For similar images and , apply the softmax function to each pixel point to convert it into the probability distributions of normal and abnormal, select the category with the highest probability as the prediction label, and obtain eight abnormal mask images, denoted as . Perform weighted fusion on the eight abnormal mask images to generate the final abnormal image.
[0017] Furthermore, the loss of the industrial image anomaly detection model includes pixel-level loss and image-level loss , and the formula is: .
[0018] Furthermore, the normal text prompt template is [a][photo][of][good][v1][v2][cls], and the abnormal text prompt template is [a][photo][of][damage][v3][v4][cls][located][at][the][position], where [cls] is the name of this type of item, [v1], [v2], [v3], [v4] represent learnable text prompt words, the abnormal prompt introduces position information, and [position] represents the position prompt word, which is set to [top], [center], [left], [right] or [bottom].
[0019] Beneficial effects: Compared with the prior art, the remarkable advantages of the present invention are: The present invention uses a noise suppression attention mechanism in the text encoder to make the model pay more attention to the key semantics in the prompt words and reduce the interference of irrelevant semantics; by introducing position information for abnormal prompts, the amount of information brought by the text modality is increased, which helps the model better understand the specific area where the anomaly occurs and enhances the model's spatial perception ability; the dual-channel image encoder with a noise suppression attention channel can better retain the detailed information of the original image on the basis of effectively eliminating environmental noise, effectively reducing the false detection of irrelevant areas of the image and the missed detection of relevant areas; through the multi-modal interaction module, the information interaction between the text and image modalities is promoted, the information interaction and fusion between the text prompt embedding and the image patch embedding are promoted, and the image patch embedding is mapped to the same semantic space as the text embedding, so as to better complete the modality alignment; this alignment not only enhances the model's ability to understand multi-modal semantics, but also improves the model's interpretability, enabling it to more accurately locate and identify abnormal areas; at the same time, the modality alignment also strengthens the model's generalization ability for unknown anomaly types, further improving the accuracy and robustness of detection. Description of the Drawings
[0020] Figure 1 It is a flowchart of an industrial image anomaly detection method for noise suppression modal fusion alignment; Figure 2 It is an architecture diagram for training an industrial image anomaly detection model; Figure 3 It is an architecture diagram for testing an industrial image anomaly detection model; Figure 4 It is an architecture diagram of the MMI module; Figure 5 It is a flowchart of the training stage; Figure 6 It is a flowchart of the testing stage. Specific implementation manner
[0021] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments.
[0022] Combined with Figure 1 As shown, an industrial image anomaly detection method for noise suppression modal fusion alignment described in this embodiment includes the following steps: Step 1, obtain normal industrial image samples and abnormal industrial image samples with labels, construct a data set, and divide it into a training set and a testing set according to a ratio.
[0023] In this example, the zero-shot learning paradigm is followed, where the test set images do not belong to the categories that appear in the auxiliary anomaly detection data set. The data sets MVTec AD and VisA are selected as the sample sets. Both data sets are color image sets, containing normal samples and abnormal samples with labels. When the test set selects the data set MVTec AD, the data set VisA is used as the auxiliary anomaly detection data set. When the test set is the data set VisA, the data set MVTec AD is used as the auxiliary anomaly detection data set.
[0024] Step 2, construct an industrial image anomaly detection model; wherein, the industrial image anomaly detection model includes a text encoder, a dual-channel image encoder, an MLP module, and a multi-modal interaction module.
[0025] Among them, the text encoder is used to convert the input text prompt into a semantic embedding vector, which can more accurately capture the key semantic information in the text and at the same time suppress the interference of irrelevant semantics. In this example, the text encoder consists of 12 layers of noise suppression attention layers (NSA layer). The NSA layer is obtained by adding a noise attention mechanism to the Transformer layer, that is, each improved Transformer layer contains a noise suppression attention mechanism and a feed-forward neural network.
[0026] A dual-channel image encoder is used to extract the global features and local detail features of the input image, and remove irrelevant noise through a noise suppression mechanism, retaining the image information useful for anomaly detection. The dual-channel image encoder includes a first-channel CLIP image encoder and a second-channel CLIP image encoder. The first-channel CLIP image encoder consists of 24 vision Transformer layers ( Figure 2 denoted as Vit layer in
[0027] ), and the second-channel CLIP image encoder consists of 24 NSA layers.
[0028] An MLP (Multi-Layer Perceptron) module is used to map the patch-level embeddings output by the image encoder to the same semantic space as the text embeddings, enabling effective alignment and interaction with the text embeddings by adjusting the dimension and semantic expression of the image features.
[0029] Step 3: Use the training set to train the industrial image anomaly detection model.
[0030] Combined with Figures 2 to 3 as shown, this example mainly includes two main stages: the training stage and the testing stage. The training stage is used to train the weights of the text prompt template, MLP module, MMI, and the noise suppression attention layer. The testing stage is used to perform anomaly detection and anomaly segmentation on unknown images to be detected.
[0031] In the training stage, first use the text prompt template to generate normal and abnormal prompts, introduce location information into the abnormal prompts to enhance the text modality, and generate accurate text embedding vectors through the text encoder. Then, adjust the input image to a fixed resolution, and obtain the global image embedding, multi-layer patch-level embeddings, and multi-layer denoised patch-level embeddings through the dual-channel image encoder. Both the patch-level embeddings and the denoised patch embeddings are operated through their respective MLP modules. Then, interact and align the patch-level embeddings with the text embeddings through the multi-modal interaction module. Finally, calculate the cosine similarity between the interacted patch-level embeddings, denoised patch-level embeddings, and text embeddings to generate an anomaly map, and calculate the loss with the true anomaly map to perform backpropagation to train the detection ability of the model.
[0032] Combined with Figure 5 , further, the steps of using the training set to train the industrial image anomaly detection model include: Construct a structured normal text prompt template and a structured abnormal text prompt template; Input the normal text prompt template and five groups of abnormal text prompt templates with position information into the text encoder for encoding. The normal text prompt template is encoded through the noise suppression attention layer to output the normal text embedding vector , and the five groups of abnormal text prompt templates with position information are respectively encoded into a group of abnormal text embedding vectors through the noise suppression attention layer. Perform mean pooling operation on the five groups of abnormal text embedding vectors to generate a fused abnormal text embedding vector ; Concatenate the normal text embedding vector and the fused abnormal text embedding vector to obtain the text prompt embedding vector , , , C is the embedding dimension, and R represents the real matrix space; Adjust the images to be detected in the training set to the preset resolution and then input them into the dual-channel image encoder to obtain the global image embedding vector, four-layer patch-level embedding vectors, and four-layer denoised patch-level embedding vectors; Perform projection operations on the four-layer patch-level embedding vectors and the four-layer denoised patch-level embedding vectors through their respective MLP modules; Through the multi-modal interaction module, interact and align the projected four-layer patch-level embedding vectors and four-layer denoised patch-level embedding vectors with the text prompt embedding vector respectively; Calculate the cosine similarities between the four-layer patch-level embedding vectors, the four-layer denoised patch-level embedding vectors, and the global image embedding vector after interaction and alignment with the text prompt embedding vector respectively to generate abnormal maps. Calculate the loss between the obtained abnormal maps and the real abnormal maps, and train the detection ability of the industrial image anomaly detection model through loss backpropagation.
[0033] Furthermore, the normal text prompt template is [a][photo][of][good][v1][v2][cls], and the abnormal text prompt template is [a][photo][of][damage][v3][v4][cls][located][at][the][position], where [cls] is the name of this type of item, [v1], [v2], [v3], [v4] represent learnable text prompt words, the abnormal prompt introduces position information, and [position] represents the position prompt word, which is set to [top], [center], [left], [right], [bottom]. [good] represents normal text, and [damage] represents abnormal text.
[0034] To make the model pay more attention to key semantics and suppress noise information, in this example, the attention mechanism in the text encoder is improved, and a noise suppression attention mechanism is used to generate more accurate text embedding vectors.
[0035] Furthermore, the text encoder consists of 12 NSA layers, and the NSA layer is obtained by adding a noise attention mechanism to the Transformer layer.
[0036] Furthermore, in the noise suppression attention layer, the input item is multiplied by a weight matrix , generating a query matrix and a key matrix . Both generated matrices are divided into two parts, denoted as , where N is the sequence length, such as 77, and C is the feature dimension, such as 768; Calculate the first attention weight matrix and the second attention weight matrix , where scale is the scaling coefficient, is the activation function; Through the difference operation between the first attention weight matrix and the second attention weight matrix, a denoising attention output is generated, expressed as: , In the formula, is a learnable weight coefficient; is a hyperparameter, initialized to 0.75, is the value matrix generated by multiplying the text template by the weight matrix , .
[0037] Furthermore, the steps of adjusting the image to be detected in the training set to a preset resolution and inputting it into the dual-channel image encoder include: Input the image to be detected into the first-channel CLIP image encoder. Extract image features through the QKV attention mechanism from the 1st layer to the 5th layer, and extract image features through the VV attention mechanism from the 6th layer to the 24th layer, and output the global image embedding vector and the patch-level embedding vectors of the 6th, 12th, 18th, and 24th layers , where i = 1, 2, 3, 4 correspond to the [6th, 12th, 18th, 24th] layers respectively, N is the number of patches with a value of 1370, and C is the embedding dimension with a value of 1024.
[0038] Input the image to be detected into the second-channel CLIP image encoder, and output the denoised patch-level embedding vectors of the 6th, 12th, 18th, and 24th layers through the noise suppression attention layer .
[0039] In one example, the image to be detected can be adjusted to a fixed resolution , and the real anomaly map is a black-and-white map, where black represents the normal area and white represents the abnormal area. The real anomaly map is adjusted to size.
[0040] Furthermore, the step of projecting both the four-layer patch-level embedding vectors and the four-layer denoised patch-level embedding vectors through their respective MLP modules includes: The patch-level embedding vectors and the denoised patch set embedding vectors are respectively fed into their respective multi-layer perceptrons for mapping operations to align the embedding vectors, and the mapped patch-level embedding vectors are respectively denoted as , with the corresponding number of patches being 1370 and the embedding dimension being 768.
[0041] Among them, the multi-layer perceptron includes an input layer, an intermediate layer, and an output layer.
[0042] Combined with Figure 4 shown, furthermore, the step of interacting and aligning the projected four-layer patch-level embedding vectors and the four-layer denoised patch-level embedding vectors with the text prompt embedding vectors through the multi-modal interaction module includes: The text prompt embedding vectors including normal text embedding and abnormal text embedding are projected into the Key space through a learnable weight matrix to obtain the matrix , and then the text prompt embedding vectors are projected into the Value space through a learnable weight matrix to obtain the matrix , and at this time the feature dimension is 768; The preliminarily aligned patch-level embedding vectors are respectively projected into the Query space through their respective learnable weight matrices to obtain ; Calculate the attention weights respectively, and the formula is: , , where ; Scale and normalize the attention weights , and the formula is: , where, , the dimension remains unchanged, is the dimension of the attention weights Multiply with matrix V and obtain the mixed four-layer patch-level embedding vectors through a linear projection layer. The formula is: , , , , Add the corresponding elements of the patch-level embedding vector and the mixed patch-level embedding vector and normalize them to obtain the finally aligned patch-level embedding vector .
[0043] Furthermore, calculate the cosine similarities between the four-layer patch-level embedding vectors, four-layer denoised patch-level embedding vectors, and global image embedding vector after interactive alignment and the text prompt embedding vector respectively. The steps to generate the anomaly map include: Calculate the cosine similarities between the patch-level embedding vector and the text prompt embedding vector respectively. The calculation formula is: , where represents the number of patches in each layer. In each layer, the image is segmented into 1370 patches, and each patch is converted into an embedding vector; The similarity matrix obtained after calculating the four-layer patch-level embedding vectors and the text prompt embedding vector is: , corresponding to the similarity matrices of the 6th, 12th, 18th, and 24th layers respectively. Each matrix is formed by calculating the cosine similarities between all patches in that layer and the text prompt embedding vector; The first patch records the global image information. When calculating the local part, ignore the similarity of the first patch, and use the similarities of 1369 patches for upsampling to obtain The size of the similarity map is: , where , represents the similarity matrix after removing the first patch, represents upsampling; Calculate the similarity between the denoised patch-level embedding vector and the text prompt embedding vector and perform upsampling to obtain The size of the similarity map is denoted as ; Calculate the similarity between the global image embedding vector and the text prompt embedding vector The cosine similarity is calculated to obtain the similarity matrix between the global image and the text prompt as : , and the calculation formula is: , For each and in the similarity graph pixel, the softmax function is applied to transform it into the probability distributions of normal and abnormal, and the category with the highest probability is selected as the prediction label, obtaining eight abnormal mask graphs, denoted as , with the size of . The eight abnormal mask graphs are weighted and fused to generate the final abnormal graph.
[0044] Furthermore, the loss of the industrial image anomaly detection model includes the pixel-level loss and the image-level loss , and the formula is: .
[0045] The obtained predicted abnormal graph is compared with the true abnormal graph of the picture , and the calculations are performed through the Focal loss and the Dice loss. The calculation formulas are respectively: , , where represents the predicted abnormal graph, is the probability that the pixel point belongs to its true category, is the balancing factor, is the modulation factor; represents the probability that the pixel point in the predicted abnormal graph belongs to the abnormal category; represents the true label corresponding to the pixel point.
[0046] Then the pixel-level loss is calculated as: , The image-level loss uses binary cross-entropy loss, and the calculation formula is: , In the formula, represents the true label of the Z sample, represents the probability that the Z sample belongs to the abnormal type.
[0047] Then the image-level loss is calculated as: 。
[0048] In one example, the total loss is used to train the learnable weights in the text prompt template, MMI module, MLP, and noise suppression attention. A total of 30 rounds of training are performed, and the industrial image anomaly detection model with the smallest loss is selected as the final model.
[0049] Step 4: Use the trained industrial image anomaly detection model to perform anomaly detection on the image samples in the test set.
[0050] In the test phase, load the trained text prompt and send it into the improved text encoder to obtain normal and abnormal text embeddings. After adjusting the test image to a fixed resolution, send it into the dual-channel image encoder to obtain the global image embedding, multi-layer patch-level embedding, and denoised patch-level embedding. Align and interact the patch-level embedding with the text embedding through the same operations as in the training phase, and calculate the similarity to generate the anomaly segmentation map. The category judgment result of the image is obtained by calculating the similarity between the global image embedding and the text embedding. The flowchart is as Figure 6 shown. The difference in the test phase is that the parameters of the industrial image anomaly detection model have been determined, and there is no need to calculate the loss function anymore. Only the anomaly segmentation map needs to be calculated through the generated similarity map. The specific implementation process is illustrated by the following example.
[0051] In one example, load the trained text prompt. The normal prompt template is [a] [photo] [of][good] [v1] [v2] [cls], and the 5 groups of abnormal prompt templates are [a] [photo] [of] [damage] [v3] [v4][cls] [located] [at] [the] [top], [a] [photo] [of] [damage] [v3] [v4] [cls][located] [at] [the] [bottom], [a] [photo] [of] [damage] [v3] [v4] [cls][located] [at] [the] [center], [a] [photo] [of] [damage] [v3] [v4] [cls][located] [at] [the] [left], [a] [photo] [of] [damage] [v3] [v4] [cls][located] [at] [the] [right].
[0052] Send the trained text prompt with class names embedded into the trained text encoder to obtain a normal text embedding and an abnormal text embedding . Resize the test image to size and feed it into the trained image encoder to obtain a global image embedding , and four-layer patch-level embeddings , , , and four-layer denoised patch-level embeddings , , , . Perform MLP mapping on the patch-level embeddings , , , , , , , to obtain , , , , , , , . Calculate the cosine similarity between , , , , , , , and the normal text embedding to obtain a similarity map with the prompt text embedding: , , , , , , , : . Perform a softmax operation on each similarity map to obtain the probabilities of describing abnormalities and normality. The final abnormal segmentation map is obtained through the following calculation formula: , where 、 represent the probabilities of the similarity map predicting abnormality and normality respectively; 、 represent the probabilities of the similarity map predicting abnormality and normality respectively.
[0053] For the abnormal segmentation image Perform binarization processing. If < 0.5, it is set to 0 to represent normal pixel points, otherwise it is 1 to represent abnormal pixel points. Embed the global image with the normal text embedding and the abnormal text embedding to calculate the cosine similarity to obtain the normal and abnormal scores. Perform softmax calculation on the two scores and finally take the maximum probability between the normal probability and the abnormal probability to determine whether the picture is an abnormal picture.
Claims
1. An industrial image anomaly detection method for noise suppression mode fusion alignment, characterized in that The steps are as follows: Obtain normal industrial image samples and abnormal industrial image samples with labels, construct a dataset, and divide it into a training set and a test set according to a ratio; Construct an industrial image anomaly detection model; wherein, the industrial image anomaly detection model includes a text encoder, a dual-channel image encoder, an MLP module, and a multimodal interaction module; Use the training set to train the industrial image anomaly detection model; Use the trained industrial image anomaly detection model to perform anomaly detection on the image samples in the test set.
2. The industrial image anomaly detection method for noise suppression modal fusion alignment according to claim 1, wherein The steps of using the training set to train the industrial image anomaly detection model include: Construct a structured normal text prompt template and a structured abnormal text prompt template; Input the normal text prompt template and five groups of abnormal text prompt templates with position information into the text encoder for encoding. The normal text prompt template is encoded through the noise suppression attention layer and then outputs the normal text embedding vector , and the five groups of abnormal text prompt templates with position information are respectively encoded into a group of abnormal text embedding vectors through the noise suppression attention layer. Perform mean pooling operation on the five groups of abnormal text embedding vectors to generate the fused abnormal text embedding vector ; Concatenate the normal text embedding vector and the fused abnormal text embedding vector to obtain the text prompt embedding vector , , , where C is the embedding dimension and R represents the real matrix space; Adjust the images to be detected in the training set to a fixed resolution and input them into the dual-channel image encoder to obtain a global image embedding vector, four-layer patch-level embedding vectors, and four-layer denoised patch-level embedding vectors; Perform projection operations on the four-layer patch-level embedding vectors and the four-layer denoised patch-level embedding vectors through their respective MLP modules; Through the multimodal interaction module, interact and align the projected four-layer patch-level embedding vectors and four-layer denoised patch-level embedding vectors with the text prompt embedding vectors respectively; Calculate the cosine similarities between the four-layer patch-level embedding vectors, the four-layer denoised patch-level embedding vectors, and the global image embedding vector after interaction and alignment with the text prompt embedding vector respectively to generate an anomaly map, calculate the loss between the obtained anomaly map and the true anomaly map, and train the detection ability of the industrial image anomaly detection model through loss backpropagation.
3. The industrial image anomaly detection method for noise suppression mode fusion alignment according to claim 2, characterized in that The text encoder consists of 12 layers of noise suppression attention layers.
4. An industrial image anomaly detection method for noise suppression modal fusion alignment according to claim 3, characterized in that, In the noise suppression attention layer, multiply the input terms by a weight matrix , generating a query matrix and a key matrix . Divide the generated matrices into two parts, denoted as , where N represents the sequence length; Calculate the first attention weight matrix and the second attention weight matrix , where scale is the scaling factor, is the activation function; Generate a denoised attention output through the difference operation between the first attention weight matrix and the second attention weight matrix, expressed as: , wherein, is a learnable weight coefficient, is a hyperparameter, is the value matrix generated by multiplying the text template by the weight matrix 5. A method for detecting industrial image anomalies by fusing and aligning noise suppression modes according to claim 4, characterized in that, The steps of adjusting the images to be detected in the training set to a fixed resolution and inputting them into the dual-channel image encoder include: Input the image to be detected into the first-channel CLIP image encoder. Extract image features through the QKV attention mechanism from the 1st layer to the 5th layer, and extract image features through the VV attention mechanism from the 6th layer to the 24th layer, and output the global image embedding vector and the patch-level embedding vectors of the 6th, 12th, 18th, and 24th layers , where i = 1, 2, 3, 4 correspond to the [6th, 12th, 18th, 24th] layers respectively; Input the image to be detected into the second-channel CLIP image encoder, and output the denoised patch-level embedding vectors of the 6th, 12th, 18th, and 24th layers through the noise suppression attention layer .
6. The industrial image anomaly detection method for noise suppression modal fusion alignment according to claim 5, characterized in that, The steps of performing projection operations on the four-layer patch-level embedding vectors and the four-layer denoised patch-level embedding vectors through their respective MLP modules include: Patch-level embedding vectors and denoised patch set embedding vectors are respectively fed into their respective multi-layer perceptrons for mapping operations to align the embedding vectors, and the mapped patch-level embedding vectors are respectively denoted as ; The multi-layer perceptron includes an input layer, an intermediate layer, and an output layer.
7. The industrial image anomaly detection method for noise suppression mode fusion alignment according to claim 6, wherein The steps of interacting and aligning the projected four-layer patch-level embedding vectors and four-layer denoised patch-level embedding vectors with the text prompt embedding vectors respectively through the multimodal interaction module include: Embed the text prompt into a vector Through a learnable weight matrix Project it into the Key space to obtain a matrix , and then embed the text prompt into a vector Through a learnable weight matrix Project it into the Value space to obtain a matrix ; The preliminarily aligned patch-level embedding vectors are respectively projected into the Query space through their respective learnable weight matrices to obtain ; Calculate the attention weights respectively, and the formula is: , , wherein ; Scale and normalize the attention weights using the formula: , wherein, , is the dimension of; Multiply the attention weights by the matrix V and pass through a linear projection layer to obtain the mixed four-layer patch-level embedding vectors. The formula is as follows: , , , , Add the patch-level embedding vectors to the mixed patch-level embedding vectors element-wise and normalize them to obtain the finally aligned patch-level embedding vectors .
8. A method for abnormal detection of industrial images with noise suppression modal fusion alignment according to claim 7, characterized in that The steps of calculating the cosine similarities between the four-layer patch-level embedding vectors, the four-layer denoised patch-level embedding vectors, and the global image embedding vector after interaction and alignment with the text prompt embedding vector respectively to generate an anomaly map include: Patch-level embedding vectors are respectively calculated with the cosine similarity to the text prompt embedding vectors The calculation formula is as follows: , Among them, represents the number of patches per layer. At each layer, the image is segmented into 1370 patches, and each patch is converted into an embedding vector; The similarity matrix obtained after calculating the four-layer patch-level embedding vectors and the text prompt embedding vectors is as follows: ; The first patch records the global image information, ignore the similarity of the first patch when calculating the local part, upsample using the similarities of 1369 patches, and the obtained similarity map is: , Among them, , represents the similarity matrix after removing the first patch, denotes upsampling; Denoise the patch-level embedding vector Calculate the similarity with the text prompt embedding vector and perform upsampling. The resulting similarity map is denoted as ; Embed the global image into a vector and the text prompt embedding vector to calculate the cosine similarity, and obtain the similarity matrix between the global image and the text prompt as , and the calculation formula is: , For similar images and Apply the softmax function to each pixel point, convert it into the probability distributions of normal and abnormal, select the category with the highest probability as the prediction label, and obtain eight abnormal mask images, denoted as , perform weighted fusion on the eight abnormal mask images to generate the final abnormal image.
9. A method for abnormal detection of industrial images with noise suppression mode fusion alignment according to any one of claims 1 to 8, characterized in that The loss of the industrial image anomaly detection model includes pixel-level loss and image-level loss , and the formula is: 。 10. A method for abnormal detection of industrial images with noise suppression modal fusion alignment according to any one of claims 2 to 8, characterized in that The normal text prompt template is [a][photo][of][good][v1][v2][cls], and the abnormal text prompt template is [a][photo][of][damage][v3][v4][cls][located][at][the][position], where [cls] is the name of this type of item, [v1], [v2], [v3], [v4] represent learnable text prompt words, the abnormal prompt introduces position information, and [position] represents the position prompt word, set to [top], [center], [left], [right] or [bottom].
Citation Information
Patent Citations
Low-light image enhancement method and device fusing high and low frequency feature information
CN116152120A
Depth representation learning and fusion method based on multi-modal trajectory
CN116956224A
Virtual fitting model training method, virtual fitting method and electronic equipment
CN119723281A
Small-sample industrial anomaly detection method based on cross-modal adaptive interaction
CN119989247A
Children MPP auxiliary diagnosis system based on multi-modal time series data modeling
CN120015296A
Cited By
Abnormal sample generation system and anomaly detection device for quality control before manufacturing
CN121167312A
Vertical industry abnormal behavior supervision method based on artificial intelligence
CN121388956A