Zero sample 2D anomaly detection method and device based on CLIP
By designing learnable general normal and abnormal semantic texts on the CLIP model, and combining global and local feature extraction, the problem of low detection accuracy of zero-sample 2D anomaly in the prior art is solved, and high-precision detection and segmentation effects are achieved.
Patent Information
- Application Number
- CN202510470058.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing zero-sample 2D abnormality detection methods vary greatly in different fields, resulting in low detection accuracy and making it difficult to achieve high-precision zero-sample 2D abnormality detection.
Using the zero-sample 2D anomaly detection method based on CLIP, by designing learnable general normal and abnormal semantic text, combining CLIP's visual encoder and text encoder, the global and local features of the image are extracted, and the global and local loss functions are defined for training, so as to realize zero-sample 2D anomaly detection.
It realizes high-precision zero-sample 2D abnormality detection and segmentation, with high detection and segmentation accuracy and great practical application value.
Smart Images

Figure CN119992237A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of anomaly detection technology, and in particular to a CLIP-based zero-sample 2D anomaly detection method and device. Background Art
[0002] Zero-shot anomaly detection requires the use of a detection model trained with auxiliary data, which can detect anomalies without any training samples in the target dataset. This type of task is important when training data is not available. However, since the trained model needs to be generalized to different fields with zero-shot, and the foreground, background, and abnormal areas of objects in different fields may be significantly different, this type of task is very challenging. Recently, pre-trained visual language multimodal models have demonstrated strong zero-shot learning capabilities in a variety of downstream tasks, providing new ideas for solving the zero-shot 2D anomaly detection problem. Summary of the invention
[0003] The purpose of this invention is to further explore and improve the deficiencies of existing 2D field anomaly detection methods and realize high-precision zero-sample 2D anomaly detection. This method proposes a learnable universal normal and abnormal semantic text and extracts features from both coarse-grained and fine-grained aspects of the image, and finally realizes zero-sample 2D anomaly detection, which is highly innovative; at the same time, the detection and segmentation accuracy of this method is high, and it has great practical application value.
[0004] The object of the present invention is achieved through the following technical solutions: In a first aspect, the present invention provides a zero-sample 2D anomaly detection method based on CLIP, the method comprising the following steps:
[0005] Step 1, data set acquisition and processing: obtain a 2D image data set of the object under test;
[0006] Step 2, design of learnable text: design a learnable text that is independent of the object, which is divided into two categories: normal general semantic text and abnormal general semantic text;
[0007] Step 3, obtaining text features: Use the original text encoder of CLIP to extract features from the learnable text obtained in step 2;
[0008] Step 4, diagonal attention mechanism and acquisition of local visual features: the self-attention mechanism in the CLIP visual encoder is changed to a diagonal attention mechanism, retaining the original parameters; the visual encoder is used to obtain the local image features output by each layer of the encoder;
[0009] Step 5, acquisition of global visual features: use CLIP's original visual encoder to obtain the global image features of the image;
[0010] Step 6, define the loss function: the loss function is obtained by adding the global loss based on the global image features and the local loss based on the local image features;
[0011] Step 7, training: Use the loss function in step 6 for training. During training, freeze the parameters of CLIP's original visual encoder and text encoder as well as the diagonal attention mechanism, and update the learnable vector in the learnable text in step 2.
[0012] Step 8, reasoning: During reasoning, use the trained learnable text and the original encoder for reasoning; obtain the anomaly score and anomaly score map, judge the anomaly based on the anomaly score, and segment the abnormal area based on the anomaly score map.
[0013] Furthermore, the detected images and the corresponding true values are obtained to construct a 2D image dataset. If the true value of a point is abnormal, it is represented as 1, otherwise it is represented as 0.
[0014] Furthermore, the learnable text in step 2 is a mixed representation learning, and the learnable vectors of the normal general semantic text and the abnormal general semantic text are updated during the training process, specifically:
[0015] Normal semantically learnable text:
[0016]
[0017] Abnormal semantics can be learned from text:
[0018]
[0019] in, is a learnable vector for normal general semantic text, is a learnable vector for unusual general semantic context.
[0020] Furthermore, the original learnable text in step 3 includes a fixed part and a learnable part. Based on the output of the two parts by each layer of the original CLIP text encoder, the learnable part output by the layer is discarded during the update, and the learnable part is reinitialized for training the next layer; the fixed part output by the last layer is used as the text feature, specifically:
[0021] use represents the fixed part of the original learnable text, and uses represents the learnable part of the original learnable text; represents the original CLIP text encoder Layer, total layer; Indicates The fixed part of the layer output, Indicates The learnable part of the layer output, where ;but:
[0022]
[0023] Afterwards, discard , reinitialize the learnable part to ;
[0024] Repeat the above steps, that is:
[0025]
[0026] in, ;
[0027] The last layer Output and ,Will As text features, that is:
[0028]
[0029] in, Represents the features extracted from normal general semantic text. The same method is used to extract the features of abnormal general semantic text. .
[0030] Furthermore, the diagonal attention mechanism and the local visual feature acquisition method in step 4 are specifically as follows:
[0031] Diagonal attention refers to using values as both the query and the key in the attention mechanism, that is:
[0032]
[0033] in, represents the value matrix in the calculated attention mechanism, It is vectors, express The dimension of the vector, ;
[0034] Afterwards, the self-attention mechanism structure in the original CLIP visual encoder is changed to a diagonal attention mechanism, while the parameters remain unchanged; the encoder is used to encode the image to obtain the local features of each point , indicating the picture Points in In the visual encoder The features of the layer output.
[0035] Furthermore, the loss function in step 6 is specifically:
[0036] The global loss is used to ensure that the object-independent text features match the global features of images of various objects, capturing the normal and abnormal semantics from the perspective of global features; it is calculated using cross entropy, namely:
[0037]
[0038] in, represents the number of samples, , Indicates normal semantic text or abnormal semantic text, Represents the softmax() normalized result of cosine similarity;
[0039] The local loss function is to The fine-grained image features output by the intermediate layer are used to judge abnormalities; first, the true value Rewrite as , Representing images of There is no abnormality at the point, otherwise it means there is an abnormality; use Indicates that the image is in the visual encoder The output image of the layer Points normal score, use accordingly represents the anomaly score; that is:
[0040]
[0041]
[0042] Then the local loss function is:
[0043]
[0044] in, express loss, express loss, represents upsampling, represents the true value, Represents the visual encoder Anomaly score map of layer output;
[0045] at last, .
[0046] Furthermore, the reasoning method in step 8 is specifically as follows:
[0047] For pictures , first use the original visual encoder of self-attention to extract global features , the anomaly score is calculated as:
[0048]
[0049] Next, the original visual encoder with diagonal attention is used to extract local features at each layer , calculate the local normal and abnormal scores, and merge the normal and abnormal scores of all layers to get the final normal score map and anomaly score plot ;Right now:
[0050]
[0051]
[0052] Finally, the score map is calculated as:
[0053]
[0054] in, Represents Gaussian filtering.
[0055] In a second aspect, the present invention provides a CLIP-based zero-sample 2D anomaly detection device, comprising a memory and one or more processors, wherein the memory stores executable code, and when the processor executes the executable code, a CLIP-based zero-sample 2D anomaly detection method as described is implemented.
[0056] In a third aspect, the present invention provides a computer-readable storage medium having a program stored thereon, and when the program is executed by a processor, the CLIP-based zero-sample 2D anomaly detection method is implemented.
[0057] In a fourth aspect, the present invention provides a computer program product, including a computer program, which, when executed by a processor, implements the CLIP-based zero-sample 2D anomaly detection method.
[0058] Compared with the prior art, the present invention has the following innovative advantages and significant effects:
[0059] 1) Using learnable text to learn common text with normal and abnormal semantics, it achieves zero-sample learning and is highly innovative;
[0060] 2) It integrates the coarse-grained global features and fine-grained local features of the image, achieving high-precision zero-sample anomaly detection and has high field application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 It is a schematic diagram of the overall process of the present invention.
[0062] Figure 2 It is a schematic diagram of the learnable text training process of the present invention.
[0063] Figure 3 It is a schematic diagram of the reasoning process of the present invention.
[0064] Figure 4 It is a structural diagram of a CLIP-based zero-sample 2D anomaly detection device provided by the present invention. DETAILED DESCRIPTION
[0065] The specific implementation method and working principle of the present invention are described in detail below in conjunction with the accompanying drawings:
[0066] This embodiment uses the industrial anomaly dataset MVTec-AD as a test set. The dataset contains 5354 normal and abnormal 2D images of 15 categories of objects and contains annotation information. The details are shown in Table 1. During training, this embodiment uses the test samples in the VisA dataset for fine-tuning. The dataset contains 10821 images of 12 different objects, of which 1200 are abnormal images and the rest are normal images. Its domain is quite different from that of MVTec-AD.
[0067] Table 1 Detailed information of the MVTec-AD dataset
[0068]
[0069] In this embodiment, a zero-sample 2D anomaly detection method based on CLIP is trained and tested on the above two data sets. The method results are to achieve zero-sample 2D anomaly detection. The detailed implementation steps are as follows:
[0070] Step 1: Use the test data in the VisA dataset as the training set, denoted as ,in Indicates The original images to be detected, represents the corresponding true value, a certain point is abnormal and is represented as 1, otherwise it is represented as 0. The data in MVTec-AD is used as the test set, denoted as ;
[0071] Step 2: Design two object-independent learnable texts:
[0072] Normal general semantics learnable text:
[0073]
[0074] Abnormal general semantics learnable text:
[0075]
[0076] in, is a learnable vector for normal text, is a learnable vector of abnormal text, which is updated during training. Indicates the specific description object and is not updated during training;
[0077] Step 3: Use the original text encoder of CLIP to get the and Encode and extract the features as and , specifically:
[0078] use Indicates the original The fixed part in Indicates the original The learnable part of represents the original CLIP text encoder Layer, total layer; Indicates The fixed part of the layer output, Indicates The learnable part of the layer output, where .but:
[0079]
[0080] Afterwards, discard , reinitialize the learnable part to .
[0081] Repeat the above steps, that is:
[0082]
[0083] in, .
[0084] The last layer Output and ,Will As text features, we get:
[0085]
[0086] Use the same method to get;
[0087] Step 4: Rewrite the self-attention part in the original CLIP visual encoder into a diagonal attention mechanism with the same parameters. The diagonal attention mechanism refers to using the value as both the query and the key in the attention mechanism, that is:
[0088]
[0089] in, represents the value matrix in the calculated attention mechanism, It is vectors, express The dimension of the vector, .
[0090] Use this encoder to Encode the images in the image to obtain the local features of each point in each image , indicating the picture Points in In the visual encoder Features of layer outputs;
[0091] Step 5: Use the original CLIP visual encoder to convert The image in is directly encoded to obtain the global features ; Assume the original visual encoder is , calculate the global eigenvector as:
[0092]
[0093] Step 6: Define the loss function ;
[0094] The global loss is to ensure that the object-independent text features match the global features of the images of various objects, which can effectively capture the normal and abnormal semantics from the perspective of global features. It is calculated using cross entropy, that is:
[0095]
[0096] represents the number of samples, . Indicates normal semantic text or abnormal semantic text, Represents the softmax() normalized result of cosine similarity.
[0097] The local loss function is to The fine-grained image features output by the intermediate layer are used to judge abnormalities. First, the true value Rewrite as , Representing images of If the point is normal, it indicates abnormality. Indicates that the image is in the visual encoder The output image of the layer Points normal score, use accordingly represents the anomaly score. That is:
[0098]
[0099] in:
[0100]
[0101]
[0102] in, express loss, express loss, represents upsampling, represents the true value, Represents the visual encoder Anomaly score map of the layer output.
[0103] Step 7: Use the loss function The training goal is to minimize During training, the parameters of CLIP's original text encoder and visual encoder are frozen, and the parameters of the learnable part of the learnable text are updated;
[0104] Step 8: Use the trained learnable text pairs 5354 images in the test, for the image , first use the original visual encoder of self-attention to extract global features ;
[0105] The anomaly score is calculated as:
[0106]
[0107] Determine whether the detected object is abnormal based on the calculated anomaly score.
[0108] Original visual encoder using diagonal attention to extract local features at each layer , calculate the local normal and abnormal scores, and merge the normal and abnormal scores of all layers to get the final normal score map and anomaly score plot .Right now:
[0109]
[0110]
[0111] Finally, the score map is calculated as:
[0112]
[0113] in, Represents Gaussian filtering.
[0114] The AUROC value obtained in the final test is 91.5%, which exceeds other methods. At the same time, the abnormal area can be segmented according to the abnormal score map to realize the visualization of the abnormal area.
[0115] The present invention discloses a zero-sample 2D anomaly detection method based on CLIP. With the help of the powerful zero-sample learning ability of the multimodal model CLIP, zero-sample 2D anomaly detection and segmentation are achieved by designing and training a learnable text, using a self-attention mechanism to extract coarse-grained global information of the image, and using a diagonal attention mechanism to extract fine-grained local information of the image. Figure 1 It is a schematic diagram of the overall process of the present invention, Figure 2 is a schematic diagram of the learnable text training process of the present invention, Figure 3 After detecting and segmenting the abnormal area, it can be further applied to industry, providing a guarantee for subsequent abnormal detection in industrial production processes.
[0116] Corresponding to the aforementioned embodiment of a zero-sample 2D anomaly detection method based on CLIP, the present invention further provides an embodiment of a zero-sample 2D anomaly detection device based on CLIP.
[0117] See also Figure 4 A CLIP-based zero-sample 2D anomaly detection device provided in an embodiment of the present invention includes a memory and one or more processors, wherein the memory stores executable code, and when the processor executes the executable code, it is used to implement a CLIP-based zero-sample 2D anomaly detection method in the above embodiment.
[0118] The embodiment of a CLIP-based zero-sample 2D anomaly detection device provided by the present invention can be applied to any device with data processing capabilities, and the device with data processing capabilities can be a device or apparatus such as a computer. The device embodiment can be implemented through software, or through hardware or a combination of software and hardware. Taking software implementation as an example, as a device in a logical sense, it is formed by the processor of any device with data processing capabilities in which it is located reading the corresponding computer program instructions in the non-volatile memory into the internal memory for execution. From the hardware level, if Figure 4As shown in FIG. 1 , a hardware structure diagram of a CLIP-based zero-sample 2D anomaly detection device provided by the present invention is provided in any device with data processing capability, except Figure 4 In addition to the processor, memory, network interface, and non-volatile memory shown, any device with data processing capabilities in which the apparatus in the embodiments is located may also include other hardware, generally based on the actual functions of the device with data processing capabilities, which will not be described in detail.
[0119] The implementation process of the functions and effects of each unit in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, and will not be repeated here.
[0120] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The device embodiment described above is only schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of the present invention. Ordinary technicians in this field can understand and implement it without paying creative work.
[0121] An embodiment of the present invention further provides a computer-readable storage medium on which a program is stored. When the program is executed by a processor, a zero-sample 2D anomaly detection method based on CLIP in the above embodiment is implemented.
[0122] The computer-readable storage medium may be an internal storage unit of any device with data processing capability described in any of the aforementioned embodiments, such as a hard disk or a memory. The computer-readable storage medium may also be an external storage device of any device with data processing capability, such as a plug-in hard disk, a smart media card (SMC), an SD card, a flash card, etc. equipped on the device. Furthermore, the computer-readable storage medium may also include both an internal storage unit and an external storage device of any device with data processing capability. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capability, and may also be used to temporarily store data that has been output or is to be output.
[0123] The present invention also provides a computer program product, comprising a computer program, and when the computer program is executed by a processor, the zero-sample 2D anomaly detection method based on CLIP is implemented.
[0124] The above embodiments are used to illustrate the present invention rather than to limit the present invention. Any modification and change made to the present invention within the spirit of the present invention and the protection scope of the claims shall fall within the protection scope of the present invention.
Claims
1. A CLIP-based zero-sample 2D anomaly detection method, characterized in that: The method comprises the following steps: Step 1, data set acquisition and processing: obtain a 2D image data set of the object under test; Step 2, design of learnable text: design a learnable text that is independent of the object, which is divided into two categories: normal general semantic text and abnormal general semantic text; Step 3, obtaining text features: Use the original text encoder of CLIP to extract features from the learnable text obtained in step 2; Step 4, diagonal attention mechanism and acquisition of local visual features: the self-attention mechanism in the CLIP visual encoder is changed to a diagonal attention mechanism, retaining the original parameters; the visual encoder is used to obtain the local image features output by each layer of the encoder; Step 5, acquisition of global visual features: use CLIP's original visual encoder to obtain the global image features of the image; Step 6, define the loss function: the loss function is obtained by adding the global loss based on the global image features and the local loss based on the local image features; Step 7, training: Use the loss function in step 6 for training. During training, freeze the parameters of CLIP's original visual encoder and text encoder as well as the diagonal attention mechanism, and update the learnable vector in the learnable text in step 2. Step 8, reasoning: During reasoning, use the trained learnable text and the original encoder for reasoning; obtain the anomaly score and anomaly score map, judge the anomaly based on the anomaly score, and segment the abnormal area based on the anomaly score map.
2. A CLIP-based zero-sample 2D anomaly detection method according to claim 1, characterized in that: The detected images and their corresponding true values are obtained to construct a 2D image dataset. If the true value of a point is abnormal, it is represented as 1, otherwise it is represented as 0.
3. The CLIP-based zero-sample 2D anomaly detection method according to claim 1, characterized in that: The learnable text in step 2 is a mixed representation learning, and the learnable vectors of normal general semantic text and abnormal general semantic text are updated during the training process.
4. The CLIP-based zero-sample 2D anomaly detection method according to claim 1, characterized in that: The text feature acquisition method in step 3 is specifically as follows: the original learnable text includes a fixed part and a learnable part. Based on the output of the two parts at each layer of the original CLIP text encoder, the learnable part output by the layer is discarded during updating, and the learnable part is reinitialized for training the next layer; the fixed part output by the last layer is used as the text feature.
5. The CLIP-based zero-sample 2D anomaly detection method according to claim 1, characterized in that: The diagonal attention mechanism and the local visual feature acquisition method in step 4 are specifically as follows: the diagonal attention mechanism refers to using the value as both the query and the key in the attention mechanism; then, the self-attention mechanism structure in the CLIP original visual encoder is changed to the diagonal attention mechanism, while the parameters remain unchanged; the encoder is used to encode the image to obtain the local features of each point, which represents the features of the points in the image output by the visual encoder.
6. The CLIP-based zero-sample 2D anomaly detection method according to claim 1, characterized in that: The loss function in step 6 is specifically: The global loss is used to ensure that the object-independent text features match the global features of images of various objects, capturing the normal and abnormal semantics from the perspective of global features, and is calculated using cross entropy; The local loss function is to judge anomalies based on the fine-grained image features output by the intermediate layer of the visual encoder; First, the true value is 0 to indicate that there is no abnormality in the point on the image, otherwise it indicates that there is an abnormality; based on the visual encoder output, the scores of normal and abnormal points in the image are obtained; The local loss function is represented based on focal loss and dice loss.
7. The CLIP-based zero-sample 2D anomaly detection method according to claim 1, characterized in that: The reasoning method in step 8 is specifically as follows: First, use the original visual encoder with self-attention to extract the global features of the image and calculate the anomaly score; Next, the original visual encoder with diagonal attention is used to extract local features of each layer, calculate local normal and abnormal scores, and merge the normal and abnormal scores of all layers to obtain the final normal score map and abnormal score map; Finally, the score map is calculated based on Gaussian filtering.
8. A CLIP-based zero-sample 2D anomaly detection device, comprising a memory and one or more processors, wherein the memory stores executable code, characterized in that: When the processor executes the executable code, a CLIP-based zero-sample 2D anomaly detection method according to any one of claims 1 to 7 is implemented.
9. A computer-readable storage medium having a program stored thereon, characterized in that: When the program is executed by a processor, a CLIP-based zero-sample 2D anomaly detection method according to any one of claims 1 to 7 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the CLIP-based zero-sample 2D anomaly detection method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Scenario graph large model illusion detection method, system and equipment for scene layout anomaly perception
CN122115962A