An Industrial Anomaly Image Segmentation Method and System Based on Large Models
Generate text prompts through multimodal models and filter out abnormal areas by local perceptrons, and combine them with CLIP encoder to optimize confidence scores, solving the problem of large-scale vision-language models adapting to multiple anomalies categories in industrial anomaly detection, achieving efficient industrial anomaly image segmentation.
Patent Information
- Application Number
- CN202411762779.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-03
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-12-03
AI Technical Summary
The prior art is difficult to adapt to multiple industrial anomaly categories without the need for normal image data, and large vision-language models lack industrial domain knowledge, resulting in inefficient detection of detection in industrial anomaly detection.
A multimodal model is used to generate text prompts, combined with a mask generator and a self-local perceptron, filter out anomaly candidate areas through a self-attention mechanism, and optimize the confidence scores using CLIP text and image encoder to achieve efficient segmentation of industrial anomaly images.
The detection quality and efficiency of industrial anomaly images are improved, and the zero-sample visual perception ability of large models is used to adapt to multiple anomaly categories without the need for normal image data.
Smart Images

Figure CN119693390B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to an industrial anomaly image segmentation method and system based on a large model. Background Art
[0002] Industrial Anomaly Detection (IAD) is a technology applied in the manufacturing and production fields, mainly used for defect detection of industrial components. The goal is to automatically detect potential anomaly types and their locations based on industrial-grade images. By early identification of anomalies, IAD can reduce losses, improve production efficiency, and ensure product quality. Previous technologies usually adopted deep convolutional neural networks as image feature extractors, mainly focused on designing dedicated models for each anomaly category, and relied on a large amount of normal image data as the training basis.
[0003] In current technologies, collecting a large number of training images for each anomaly category is not only extremely challenging but also costly. In addition, applying different models separately for each type of anomaly will make the actual industrial scenario too complex. These limitations have hindered the widespread application of existing industrial anomaly detection (IAD) methods. Given the characteristics of insufficient data and diverse anomaly types in industrial scenarios, a single model that can adapt to multiple anomaly categories without referring to normal images has become a practical need.
[0004] In recent years, many studies based on Large Vision-Language Models (LVLMs) have demonstrated powerful zero-shot visual perception capabilities. Even in the case of scarce or completely lacking normal images, these models can quickly adapt to multiple categories of tasks. Through carefully designed prompts, they can retrieve the prior knowledge stored in the model, thereby enhancing the detection ability. In addition, these models can be fully applied to downstream tasks through simple natural language descriptions and prompts, such as zero-shot classification using carefully constructed prompts.
[0005] Large Vision-Language Models (LVLMs) are only pre-trained on visual and language datasets from the network, lacking an understanding of industrial domain knowledge and sensitivity to local anomaly details in images. The huge domain difference between the pre-training dataset and the industrial anomaly dataset results in the mask generator generating a large number of irrelevant mask candidate regions. These limitations restrict their potential in industrial anomaly detection (IAD) tasks. Summary of the Invention
[0006] Based on the technical problems existing in the background art, the present invention proposes an industrial anomaly image segmentation method and system based on a large model, realizing task migration of industrial anomaly datasets on the large model, and improving the detection quality and efficiency of anomaly images.
[0007] An industrial anomaly image segmentation method based on a large model proposed by the present invention includes the following steps:
[0008] Step 1, obtain industrial anomaly images;
[0009] Step 2, obtain text prompts corresponding to the industrial anomaly images based on a multimodal model;
[0010] Step 3, input the industrial anomaly images and the text prompts into a mask generator to generate a set of anomaly candidate regions, and input the set of anomaly candidate regions into an anomaly candidate region generator to generate multiple segmentation mask images and corresponding confidence scores;
[0011] Step 4, divide each segmentation mask image into multiple sub-regions, map each sub-region into a one-dimensional vector through linear mapping, and combine it with a position embedding to generate a query Q. At the same time, perform feature extraction on the divided image through a mask encoder to generate a key K and a value V. Calculate the sub-weights of each sub-region based on the Q-K-V self-attention mechanism and take the average. If the average value of the sub-weights is greater than the set weight upper limit, it is used as an anomaly weight, indicating that there is an anomaly in the sub-region of the corresponding anomaly candidate region. Initially screen out the anomaly candidate regions irrelevant to the anomaly, thereby adjusting the mask encoder, and generate a set of anomaly segmentation mask candidates corresponding to all segmentation mask images based on the adjusted mask encoder;
[0012] Step 5, input the text prompts, as well as the pre-set overall prompt and detail prompt, into a CLIP text encoder, and add the output text overall feature and text detail feature to obtain a text feature;
[0013] Step 6, input the industrial anomaly images and the set of anomaly segmentation mask candidates into a CLIP image encoder respectively, and add the output image feature and mask feature to obtain a visual feature;
[0014] Step 7, calculate the similarity between the text feature and the visual feature, re-correct the confidence scores corresponding to each anomaly candidate region in the set of anomaly segmentation mask candidates based on the similarity calculation result, arrange the corrected confidence scores in descending order, and take the top K anomaly candidate regions corresponding to the confidence scores and aggregate them to obtain the final anomaly map.
[0015] Further, in Step 2, the multi-modal model adopts the Blip+ChatGPT approach. The industrial anomaly image generates a text description containing nouns related to the anomaly object based on Blip. ChatGPT performs comparative reasoning based on the received text description and the pre-provided query statement to generate the target noun T that conforms to the anomaly concept. o 。
[0016] Further, in Step 3, the mask generator adopts the Grounding-Dino model pre-trained on ImageNet to generate a set of anomaly candidate regions with bounding boxes and their corresponding confidence scores.
[0017] The anomaly region generator is an open-world vision segmentation base model constructed using SAM to generate a pixel-level segmentation mask image.
[0018] Further, in Step 4, in calculating and averaging the sub-weights of each sub-region based on the Q-K-V self-attention mechanism, specifically:
[0019] Configure a set number of fixed keys K for each query Q to focus on a small group of attention points around the query point.
[0020] Remove the top-level MLP of the deformable attention mechanism and apply a projection layer for dimension expansion to calculate the sub-weights of each sub-region and perform an averaging operation on the sub-weights of all sub-regions.
[0021] Further, divide each segmentation mask image into 9 sub-regions. The weight calculation process of the segmentation mask image is as follows:
[0022] Q 1-9 =Proj.[Pos;p1;...;p9];
[0023] K=V=Proj.(γ clip ·(M c ·Q));
[0024]
[0025] W′=Avg(W 1-9 );
[0026] where Q 1-9 is the query Q for sub-regions 1 to 9, Pos represents the position embedding, M c is the segmentation mask image, W n is the sub-weight of the nth sub-region, n ∈ [1, 9], i represents the ith segmentation mask image, I is the total number of segmentation mask images, i.e., the industrial anomaly image, Attn. is the self-attention mechanism, Q n is the query Q for the nth sub-region, Ki K is the key for the i-th segmentation mask image, Avg is the average operation, W′ is the weight of the segmentation mask image, i.e., the average of sub-weights, p1;...; p9 respectively represent 9 sub-regions, γ clip represents the CLIP text encoder.
[0027] An industrial anomaly image segmentation system based on a large model, including an acquisition module, a text acquisition module, a mask image generation module, an anomaly segmentation mask candidate set generation module, a text feature generation module, a visual feature generation module, and an anomaly map generation module;
[0028] The acquisition module is used to acquire industrial anomaly images;
[0029] The text acquisition module is used to obtain the text prompt corresponding to the industrial anomaly image based on the multimodal model;
[0030] The mask image generation module is used to input the industrial anomaly image and the text prompt into the mask generator to generate a set of anomaly candidate regions, and input the set of anomaly candidate regions into the anomaly candidate region generator to generate multiple segmentation mask images and corresponding confidence scores;
[0031] The anomaly segmentation mask candidate set generation module is used to divide each segmentation mask image into multiple sub-regions, map each sub-region into a one-dimensional vector through linear mapping, and combine it with the position embedding to generate the query Q. At the same time, the divided image is subjected to feature extraction through the mask encoder to generate the key K and the value V. Based on the Q-K-V self-attention mechanism, the sub-weights of each sub-region are calculated and averaged. Taking the average of the sub-weights greater than the set weight upper limit as the anomaly weight indicates that there is an anomaly in the sub-region of the corresponding anomaly candidate region, and the anomaly candidate regions irrelevant to the anomaly are initially screened out, thereby adjusting the mask encoder, and generating the anomaly segmentation mask candidate set corresponding to all segmentation mask images based on the adjusted mask encoder;
[0032] The text feature generation module is used to input the text prompt, the overall prompt and the detail prompt set in advance into the CLIP text encoder, and add the output text overall feature and text detail feature to obtain the text feature;
[0033] The visual feature generation module is used to input the industrial anomaly image and the anomaly segmentation mask candidate set into the CLIP image encoder respectively, and add the output image feature and mask feature to obtain the visual feature;
[0034] The anomaly map generation module is used to calculate the similarity between text features and visual features, re-correct the confidence scores corresponding to each anomaly candidate region in the anomaly segmentation mask candidate set based on the similarity calculation results, arrange the corrected confidence scores in descending order, and take the top K anomaly candidate regions corresponding to the confidence scores and aggregate them to obtain the final anomaly map.
[0035] A computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the industrial anomaly image segmentation method described above is implemented.
[0036] A computer-readable storage medium stores a number of classification programs, and the number of classification programs is used to be called by a processor and execute the industrial anomaly image segmentation method described above.
[0037] The advantages of an industrial anomaly image segmentation method and system based on a large model provided by the present invention are as follows: By using an existing pre-trained multi-modal large model, the anomaly segmentation task is completed. Previously, due to the domain gap between the pre-trained language-vision dataset and the industrial anomaly dataset, it was impossible to directly transfer tasks on the large model. In this embodiment, multi-modal prompts are used to successfully help the large model release its powerful zero-shot visual perception ability, thereby improving the detection quality and efficiency of anomaly images. Description of the Drawings
[0038] Figure 1 It is a schematic structural diagram of the present invention. Detailed Embodiments
[0039] Next, the technical solutions of the present invention will be described in detail through specific embodiments. Many specific details are set forth in the following description in order to fully understand the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0040] In this embodiment, industrial anomaly detection is defined as: identifying and monitoring abnormal situations or abnormal behaviors in a system, device, or process through technical means so as to take early measures to avoid possible problems. Its goal is to identify patterns or events that do not conform to normal behavior from a large amount of data, thereby improving production efficiency and product quality.
[0041] As Figure 1 shown, an industrial anomaly image segmentation method based on a large model proposed by the present invention mainly includes four parts: a mask generator, an anomaly descriptor, a self-local perception device, and an anomaly visual-text aligner. The specific steps are as follows:
[0042] Step 1: Obtain industrial anomaly images;
[0043] Step 2: Obtain the text prompt corresponding to the industrial anomaly image based on the multimodal model;
[0044] In this embodiment, the generation of the text prompt is implemented based on an anomaly descriptor. Specifically, the "Blip + ChatGPT" method is adopted as an alternative to the large multimodal model. An important advantage of this method is that it does not require retraining, but achieves open-world object detection through pre-trained weights. Specifically, Blip is responsible for receiving the anomaly image and completing the understanding and generation tasks. It analyzes the image based on the pre-trained weights and generates a text description containing nouns related to the anomaly object. Subsequently, ChatGPT takes over the task and performs comparative reasoning by receiving the generated description and the provided query statement. It matches and generates the target noun T that best conforms to the concept of "anomaly" from the text description o , which is the text prompt.
[0045] Step 3: Input the industrial anomaly image and the text prompt into a mask generator to generate a set of anomaly candidate regions, and input the set of anomaly candidate regions into an anomaly candidate region generator to generate multiple segmentation mask images and corresponding confidence scores;
[0046] Based on the text-guided open-set object detection framework, the architecture of the mask generator is established for visual localization. This architecture adopts the Grounding-Dino model pre-trained on ImageNet. Through the text encoder and the visual encoder, this model extracts the features of the anomaly text prompt and the query object respectively. These features are utilized by the cross-modal decoder to generate a set of anomaly candidate regions with bounding boxes and their corresponding confidence scores.
[0047] Subsequently, SAM (Segment Anything Model) is used to construct an open-world visual segmentation base model as the anomaly region generator. This model is also trained based on a large-scale image segmentation dataset. Specifically, it takes the bounding boxes of the set of anomaly candidate regions as input and passes them to the prompt-conditioned mask decoder, thereby generating pixel-level anomaly segmentation masks. Under the conditions of the input image I and the text prompt P, the segmentation mask M c and the corresponding confidence score S can be obtained:
[0048] M c , S = Generator(I, P);
[0049] where Generator is the mask generator.
[0050] Step 4: Divide each segmented mask image into multiple sub-regions, map each sub-region into a one-dimensional vector through linear mapping, and combine it with positional embeddings to generate query Q. At the same time, perform feature extraction on the divided image through a mask encoder to generate key K and value V. Calculate the sub-weights of each sub-region based on the Q-K-V self-attention mechanism and take the average. If the average value of the sub-weights is greater than the set weight upper limit, it is regarded as an abnormal weight, indicating that there is an abnormality in the sub-region of the corresponding abnormal candidate region. Initially screen out the abnormal candidate regions that have nothing to do with the abnormality, thereby adjusting the mask encoder, and generating an abnormal segmentation mask candidate set corresponding to all segmented mask images based on the adjusted mask encoder;
[0051] This embodiment proposes a self-local perceptron based on an improved self-attention mechanism. Compared with the original Q-K-V self-attention mechanism in Transformer, in this embodiment, the segmented mask image is divided into 9 sub-regions (patches). Each sub-region is mapped into a one-dimensional vector through linear mapping and combined with positional embeddings to generate the final query Q (query). At the same time, the divided segmented mask image is subjected to feature extraction through a mask encoder to generate key K (keys) and value V (values). Referring to the Deformable Attention mechanism, in this embodiment, only a small number of fixed keys K (keys) are assigned to each query, focusing on a small group of attention points around the query point. Remove the top MLP layer (multi-layer perceptron) of the Deformable Attention mechanism, and finally apply a projection layer (projection layer) to flatten the dimensions to calculate the sub-weights and perform an averaging operation. A higher average value indicates that there is an abnormal part in the sub-region of the abnormal candidate region, thereby helping the abnormal detection model to initially screen out the candidate regions that have nothing to do with the abnormality. Among them, the calculation process of the abnormal weight W′ of a certain segmented mask image can be expressed as:
[0052] Q 1-9 = Proj.[Pos; p1;...; p9];
[0053] K = V = Proj.(γ clip ·(M c ·Q));
[0054]
[0055] W′ = Avg(W 1-9 );
[0056] Among them, Q 1-9 is the query Q of sub-regions 1 to 9, Pos represents positional embeddings, M c is the segmented mask image, Wn is the sub - weight of the nth sub - region, where n ∈ [1, 9], i represents the i - th segmentation mask image, I is the total number of segmentation mask images, that is, industrial anomaly images, Attn. is the self - attention mechanism, and Q n is the query Q of the nth sub - region, and K i is the key K of the i - th segmentation mask image, Avg is the average operation, W′ is the weight of the segmentation mask image, that is, the average value of sub - weights, and p1;...; p9 respectively represent 9 sub - regions. represents the CLIP text encoder.
[0057] Step Five: Input the text prompt, the pre - set overall prompt and detail prompt into the CLIP text encoder together, and add the output text overall feature and text detail feature to obtain the text feature.
[0058] Step Six: Input the industrial anomaly image and the anomaly segmentation mask candidate set into the CLIP image encoder respectively, and add the output image feature and mask feature to obtain the visual feature.
[0059] Step Seven: Calculate the similarity between the text feature and the visual feature, re - correct the confidence score corresponding to each anomaly candidate region in the anomaly segmentation mask candidate set based on the similarity calculation result, sort the corrected confidence scores in descending order, and take the top K anomaly candidate regions corresponding to the confidence scores and aggregate them to obtain the final anomaly map.
[0060] Steps Five to Seven are implemented based on the anomaly visual - text aligner. Specifically:
[0061] To reduce language ambiguity and fully utilize the potential of the domain - expert - knowledge text features to achieve better detection effects, this embodiment proposes a text - feature - based anomaly localization module. First, the anomaly prompt consists of two parts, namely the detail prompt and the overall prompt. Among them, the detail prompt is also used to supplement the initialization prompt in the anomaly descriptor. The detail prompt is designed based on the expert knowledge of the anomaly patterns of similar products to supplement more specific anomaly details. This prompting method reformulates the task of finding anomaly regions as locating objects with specific anomaly - state expressions, which makes the extracted text features better match the pre - trained model and makes more use of the base model than identifying "anomaly" in the object context. By combining the above two prompts, it is possible to refine a simple "anomaly" prompt into a more specific prompt that describes the anomaly state in more detail. Use the pre - trained CLIP text encoder to extract the text detail feature DF t :
[0062] DF t = θ clip (dp, T O )
[0063] Among them, dp is the detail hint, T O is the text hint, θ clip is the pre-trained CLIP text encoder.
[0064] The overall hint is the overall description of each image in the industrial anomaly dataset. Since the industrial anomaly images in the industrial anomaly dataset are all normalized, they only differ in the abnormal state or quantity. In this embodiment, the number of industrial product anomalies, the maximum number of abnormal regions, and the size of the abnormal area are uniformly described to make up for the deficiency caused by the lack of understanding of specific attributes of the basic model. Such a design helps to focus on the target object itself while understanding the overall expression of the target. Similarly, the overall text feature OF is extracted using the pre-trained CLIP text encoder t :
[0065] OF t = θ clop (op);
[0066] Among them, θ clip is the pre-trained CLIP text encoder, and op is the overall hint.
[0067] Then, the text detail feature DF t and the overall text feature OF t are combined to obtain the text feature T f , as follows:
[0068] T f = DF t + OF t .
[0069] After obtaining the context feature of the text of the domain expert knowledge (i.e., the text feature T f ), the previously obtained candidate set of anomaly segmentation masks is used for further optimization. Since different masks segment the features of different regions, in fact, different regions contain different feature information. However, the original visual feature of CLIP is designed to generate a single feature vector to describe the entire image, which is not very suitable for pixel-level dense prediction. To solve this problem, in this embodiment, the visual encoder in CLIP is modified to extract the features containing the information of the occluded region and the surrounding region. Specifically, given the input industrial anomaly image and the candidate set of anomaly segmentation masks, the visual feature consists of the overall image feature and a new image obtained from the anomaly segmentation mask image to obtain only the region around the mask proposal, and they are respectively passed to the corresponding Clip image encoder:
[0070]
[0071] Among them, DFv is the mask feature, OF v is the image feature, is the Clip image encoder, dv crop is the mask encoder, I is the industrial anomaly image, ov is the image feature, ⊙ is the Hadamard product, m is the number of times to extract detailed features from the industrial anomaly image, that is, there are m masks and m times of image features need to be extracted, and ov is the extracted image feature.
[0072] Then the "whole + detail" visual context features can be extracted, that is, the visual feature V f , as follows:
[0073] V f = DF v + OF v .
[0074] Since this embodiment is based on CLIP, where visual features and text features are embedded in a common space. So far, each segmentation mask in the anomaly segmentation mask candidate set has a fixed set of visual - language feature combinations. By calculating the similarity between the text feature T f and the visual feature V f , the score of each anomaly segmentation mask proposal is obtained. Given the input of the image I and the reference expression T, this embodiment can calculate the similarity C m between the visual features and text features of all mask proposals:
[0075] C m = sim(V f , T f );
[0076] where sim is the cosine similarity calculation.
[0077] Generally, the number of abnormal regions in the object to be inspected is limited. Therefore, this embodiment screens the top K candidate regions with the highest confidence scores and uses their average value for the final abnormal region detection.
[0078] That is, the score corresponding to a single region in the anomaly segmentation mask candidate set is represented as S m , and it is recalibrated using the cosine similarity:
[0079] s = C m · S m
[0080] M′ C , S′ C = Top k (M c , s)
[0081] M′ C , S ′ C are the calibrated segmentation masks and the corresponding confidence scores respectively. Then, these K candidate regions are aggregated to estimate the final anomaly map A:
[0082]
[0083] where s is the corrected confidence score, and Top k (M c , s) are the anomaly candidate regions corresponding to the top K confidence scores, M′ C , S C ′ are the corrected anomaly segmentation mask and the corresponding confidence score, and α is a balance coefficient, preferably set to 0.1, and the specific value can be adjusted by itself in the experiment.
[0084] When performing anomaly detection on industrial anomaly images above, a normal reference image is used as a reference benchmark, so as to facilitate the identification of anomalies from industrial anomaly images.
[0085] Through more comprehensive industrial domain expert knowledge and more accurate prompt engineering, the finally optimized framework is obtained, which can generate more reliable predictions.
[0086] According to Steps 1 to 7, this embodiment utilizes an existing pre-trained multi-modal large model to complete the anomaly segmentation task without any further training. Previously, due to the domain gap between the pre-trained language-vision dataset and the industrial anomaly dataset, it was impossible to directly perform task migration on the large model. In this embodiment, multi-modal prompts are used to successfully help the large model release its powerful zero-shot visual perception ability. By integrating multiple data types from different data sources, the large model can generalize from the learned patterns to new data and tasks, which enables them to achieve good results in other applications and provides strong technical support for various complex multi-modal processing. It provides valuable inspiration for the non-adaptive design in the multi-modal field.
[0087] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered within the protection scope of the present invention.
Claims
1. An industrial abnormal image segmentation method based on a large model, characterized in that, It includes the following steps: Step 1, obtain industrial anomaly images; Step 2, obtain the text prompt corresponding to the industrial anomaly image based on the multi-modal model; Step 3, input the industrial anomaly image and the text prompt into a mask generator to generate a set of anomaly candidate regions, and input the set of anomaly candidate regions into an anomaly candidate region generator to generate multiple segmentation mask images and corresponding confidence scores; Step 4, divide each segmentation mask image into multiple sub-regions, map each sub-region into a one-dimensional vector through linear mapping, and combine it with position embedding to generate a query Q. At the same time, extract features from the divided image through a mask encoder to generate a key K and a value V. Calculate the sub-weights of each sub-region based on the Q-K-V self-attention mechanism and take the average. Take the average value of the sub-weights greater than the set weight upper limit as the anomaly weight, indicating that there is an anomaly in the sub-region corresponding to the anomaly candidate region. Initially screen out the anomaly candidate regions irrelevant to the anomaly, thereby adjusting the mask encoder, and generate a set of anomaly segmentation mask candidates corresponding to all segmentation mask images based on the adjusted mask encoder; Step 5, input the text prompt, as well as the pre-set overall prompt and detail prompt, into the CLIP text encoder, and add the output text overall feature and text detail feature to obtain the text feature; Step 6, input the industrial anomaly image and the set of anomaly segmentation mask candidates into the CLIP image encoder respectively, and add the output image feature and mask feature to obtain the visual feature; Step 7, calculate the similarity between the text feature and the visual feature, re-correct the confidence score corresponding to each anomaly candidate region in the set of anomaly segmentation mask candidates based on the similarity calculation result, and sort the corrected confidence scores in descending order. Take the top K anomaly candidate regions corresponding to the confidence scores and aggregate them to obtain the final anomaly map.
2. The industrial anomaly image segmentation method based on a large model according to claim 1, wherein, In step two, the multimodal model adopts the Blip+ChatGPT method. For industrial anomaly images, Blip is used to generate a text description containing nouns related to the anomaly object, and ChatGPT performs comparative reasoning based on the received text description and the pre-provided query statement to generate a target noun T that conforms to the anomaly concept. o .
3. The industrial anomaly image segmentation method based on a large model according to claim 1, wherein In Step 3, the mask generator uses the Grounding-Dino model pre-trained on ImageNet to generate a set of anomaly candidate regions with bounding boxes and their corresponding confidence scores; The anomaly region generator is an open-world visual segmentation basic model constructed using SAM to generate pixel-level segmentation mask images.
4. The industrial anomaly image segmentation method based on a large model according to claim 1, wherein, In Step 4, in the process of calculating the sub-weights of each sub-region based on the Q-K-V self-attention mechanism and taking the average, specifically: Configure a set number of fixed keys K for each query Q to focus on a small group of attention points around the query point; Remove the top-level MLP of the deformable attention mechanism and apply a projection layer for dimension expansion to calculate the sub-weights of each sub-region, and perform an average operation on the sub-weights of all sub-regions.
5. The industrial anomaly image segmentation method based on a large model according to claim 4, wherein Divide each segmentation mask image into 9 sub-regions, and the weight calculation process of the segmentation mask image is as follows: Q 1-9 = Proj.[Pos; p1;...; p9]; K = V = Proj.(γ clip ·(M c ·Q)); W′ = Avg(W 1-9 ) Among them, Q 1-9 is the query Q for 1 to 9 sub-regions, Pos represents the positional embedding, M c is the segmentation mask image, W n is the sub-weight of the nth sub-region, n ∈ [1, 9], i represents the ith segmentation mask image, I is the total number of segmentation mask images, that is, the industrial anomaly image, Attn. is the self-attention mechanism, Q n is the query Q of the nth sub-region, K i is the key K of the ith segmentation mask image, Avg is the average operation, W′ is the weight of the segmentation mask image, that is, the average value of the sub-weights, p1;...; p9 respectively represent 9 sub-regions, γ clip represents the CLIP text encoder.
6. An industrial anomaly image segmentation system based on a large model, characterized in that, It includes an acquisition module, a text acquisition module, a mask image generation module, an anomaly segmentation mask candidate set generation module, a text feature generation module, a visual feature generation module, and an anomaly map generation module; The acquisition module is used to obtain industrial anomaly images; The text acquisition module is used to obtain the text prompt corresponding to the industrial anomaly image based on the multi-modal model; The mask image generation module is used to input the industrial anomaly image and the text prompt into the mask generator to generate a set of anomaly candidate regions, and input the set of anomaly candidate regions into the anomaly candidate region generator to generate multiple segmentation mask images and corresponding confidence scores; The anomaly segmentation mask candidate set generation module is used to divide each segmentation mask image into multiple sub-regions, map each sub-region into a one-dimensional vector through linear mapping, and combine it with the position embedding to generate the query Q. At the same time, the divided image is subjected to feature extraction through the mask encoder to generate the key K and the value V. Based on the Q-K-V self-attention mechanism, the sub-weights of each sub-region are calculated and averaged. The average value of the sub-weights greater than the set weight upper limit is used as the anomaly weight, indicating that the sub-region of the corresponding anomaly candidate region has an anomaly. The anomaly candidate regions unrelated to the anomaly are preliminarily screened out, thereby adjusting the mask encoder, and based on the adjusted mask encoder, an anomaly segmentation mask candidate set corresponding to all segmentation mask images is generated; The text feature generation module is used to input the text prompt, as well as the preset overall prompt and detail prompt, into the CLIP text encoder, and add the output text overall feature and text detail feature to obtain the text feature; The visual feature generation module is used to input the industrial anomaly image and the anomaly segmentation mask candidate set into the CLIP image encoder respectively, and add the output image feature and mask feature to obtain the visual feature; The anomaly map generation module is used to calculate the similarity between the text feature and the visual feature, and re-correct the confidence score corresponding to each anomaly candidate region in the anomaly segmentation mask candidate set based on the similarity calculation result, and sort the corrected confidence scores in descending order, and take the top K anomaly candidate regions corresponding to the confidence scores and aggregate them to obtain the final anomaly map.
7. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the industrial anomaly image segmentation method according to any one of claims 1-6.
8. A computer-readable storage medium, characterized in that, A number of classification programs are stored on the computer-readable storage medium, and the number of classification programs are used to be called by the processor and execute the industrial anomaly image segmentation method according to any one of claims 1-6.
Citation Information
Patent Citations
Image defect detection method, system and device
CN117575997A
Training method of anomaly detection model, and object anomaly detection method and device
CN117992898A
Cited By
Conditional ai generation image detection method and system based on semantic understanding
CN122637405A