Zero sample anomaly detection method, device and equipment based on vision and language model

By using a visual and language model-based method in zero-sample anomaly detection, using the relationship between the same batch of images for abnormal detection, the problems of large computing resources and weak generalization capabilities in the prior art are solved, and efficient and accurate zero-sample anomaly detection is achieved.

CN119941644AActive Publication Date: 2025-05-06浙江大学宁波国际科创中心
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202411946095.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-06
Estimated Expiration
2044-12-27

AI Technical Summary

Technical Problem

The prior art has problems such as large computing resource consumption and weak generalization ability in zero-sample anomaly detection, especially in the absence of samples in specific fields, it is difficult to achieve effective anomaly detection.

Method used

The zero-sample anomaly detection method based on visual and language models is adopted. By obtaining inference images and reference images from the same batch of images, the first multimodal model is used for dual-modal detection, and the initial mask is generated; then the features are extracted and feature aggregation is performed through the second multimodal model, noise features are filtered, abnormal scores are calculated, and mask refinement is performed, and the final inference mask is finally obtained.

Benefits of technology

It realizes efficient and accurate zero-sample anomaly detection without additional training data, reducing computing resource consumption, and improving the robustness and detection accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941644A_ABST
    Figure CN119941644A_ABST
Patent Text Reader

Abstract

The invention provides a zero sample anomaly detection method, device and equipment based on a vision and language model, and relates to the technical field of anomaly detection, and the method comprises the steps: obtaining a reasoning image and a reference image from the same batch of images; performing bimodal detection on the reasoning image and the reference image through a first multi-modal model to obtain an initial mask; respectively extracting features of the same batch of images through a second multi-modal model, and respectively carrying out feature aggregation to obtain an aggregation reasoning feature and an aggregation reference feature; performing noise feature filtering on the aggregation reference feature according to the initial mask to obtain a non-abnormal aggregation reference feature, and obtaining an initial abnormal score according to the non-abnormal aggregation reference feature and the aggregation reasoning feature; and performing mask refining according to the initial anomaly score and the initial mask to obtain a final inference mask for realizing zero sample anomaly detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of anomaly detection technology, and in particular to a zero-sample anomaly detection method, device and equipment based on vision and language models. Background Art

[0002] Anomaly detection is a key topic in computer vision and is widely used in industrial visual inspection. The model compares the difference between normal images and images to be detected in feature space to detect anomalies in the latter. In related technologies, anomaly detection methods mainly adopt unsupervised paradigms and still require a large number of normal samples for training. These methods consume a lot of computing resources and have weak generalization capabilities.

[0003] With the emergence of zero-shot and few-shot visual language models, such models are applied to zero-shot anomaly detection, and rare classes become important in many application scenarios. The original QK attention mechanism affects the local semantics and limits the ability of fine-grained anomaly segmentation. In order to strengthen the local semantic representation, the self-attention mechanism is introduced. Although the improvement of the self-attention mechanism has brought performance improvement, the local semantic representation is still insufficient, and the entire test data set needs to be used for feature matching for anomaly detection. In actual production environments, tests are generally performed on a batch of samples, requiring the prior knowledge of the distribution of the test data set to be optimized in advance, but it is not practical to use the entire test data set in real industrial scenarios. The method of implementing zero-shot anomaly detection is still imperfect in related technologies. Summary of the invention

[0004] The problem solved by the present invention is how to achieve zero-sample anomaly detection.

[0005] To solve the above problems, the present invention provides a zero-sample anomaly detection method, device and equipment based on vision and language models.

[0006] In a first aspect, the present invention provides a zero-sample anomaly detection method based on a vision and language model, comprising: obtaining an inference image and a reference image from the same batch of images, wherein the inference image represents an image to be detected for anomalies, and the reference image includes other images in the same batch of images except the inference image;

[0007] Using a first multimodal model, performing bimodal detection on the reasoning image and the reference image respectively to obtain an initial mask, wherein the initial mask includes an initial reasoning anomaly mask and an initial reference anomaly mask;

[0008] Extracting features of the same batch of images respectively through a second multimodal model and performing feature aggregation respectively to obtain aggregated inference features and aggregated reference features;

[0009] According to the initial mask, filtering the aggregated reference features for noise features to obtain non-abnormal aggregated reference features, and obtaining an initial anomaly score according to the non-abnormal aggregated reference features and the aggregated reasoning features, wherein the initial anomaly score includes a reasoning anomaly score and a reference anomaly score;

[0010] Mask refinement is performed according to the initial anomaly score and the initial mask to obtain a final reasoning mask, wherein the final reasoning mask is used to represent anomaly detection reasoning results.

[0011] Optionally, before performing bimodal detection on the reasoning image and the reference image respectively through the first multimodal model to obtain an initial mask, the method further includes:

[0012] Determine the initial multimodal model;

[0013] Removing residual connections and feed-forward modules in the initial multimodal model;

[0014] The attention features of the intermediate layers of the initial multimodal model are weighted to replace the attention features of the last layer to obtain the first multimodal model, wherein the intermediate layers include specific layers in the initial multimodal model.

[0015] Optionally, before performing bimodal detection on the reasoning image and the reference image respectively through the first multimodal model to obtain an initial mask, the method further includes:

[0016] Construct normal text prompt templates and abnormal text prompt templates;

[0017] The normal text features and the abnormal text features in the normal text prompt template and the abnormal text prompt template are respectively extracted through the first multimodal model.

[0018] Optionally, the text features include the normal text features and the abnormal text features, and the step of performing bimodal detection on the reasoning image and the reference image respectively through the first multimodal model to obtain the initial mask includes:

[0019] extracting, by means of the first multimodal model, inference image features and reference image features of the inference image and the reference image respectively;

[0020] respectively calculating the inference similarity and the reference similarity between the inference image feature and the text feature, and between the reference image feature and the text feature;

[0021] The initial mask is determined according to the inferred similarity and the reference similarity.

[0022] Optionally, extracting features of the same batch of images respectively through the second multimodal model and performing feature aggregation respectively to obtain aggregated inference features and aggregated reference features includes:

[0023] For each image, extracting a stage patch feature through the second multimodal model, wherein the stage patch feature is obtained by outputting each layer of the visual encoder of the second multimodal model;

[0024] Performing feature aggregation on the patch features of the stage with the same image number and the same number of visual encoder layers within a preset neighborhood area to obtain aggregated features;

[0025] The aggregated inference feature is obtained according to the aggregated feature corresponding to the inference image, and the aggregated reference feature is obtained according to the aggregated feature corresponding to the reference image.

[0026] Optionally, filtering the aggregated reference features for noise features according to the initial mask to obtain non-abnormal aggregated reference features, and obtaining an initial abnormality score according to the non-abnormal aggregated reference features and the aggregated reasoning features includes:

[0027] filtering out the aggregated reference features having abnormal features according to the initial mask, and retaining non-abnormal aggregated reference features;

[0028] The feature distance between the aggregated reasoning feature and the non-abnormal aggregated reference feature is calculated, and the minimum feature distance is used as the initial abnormality score.

[0029] Optionally, performing mask refinement according to the initial anomaly score and the initial mask to obtain a final reasoning mask includes:

[0030] In the inference image and the reference image, an average anomaly score is generated according to the extracted stage patch features, wherein the average anomaly score includes an inference average anomaly score and a reference average anomaly score;

[0031] The average anomaly score is averaged with the initial anomaly score to obtain a refined reasoning anomaly score and a refined reference anomaly score for each neighborhood area;

[0032] Binarizing the refined inference anomaly score and the refined reference anomaly score that exceed a refinement threshold to obtain an inference intermediate mask and a reference intermediate mask;

[0033] The inference intermediate mask and the reference intermediate mask are respectively subjected to collaborative voting to obtain a final inference mask and a final reference mask.

[0034] Optionally, the performing collaborative voting on the inference intermediate mask and the reference intermediate mask respectively to obtain a final inference mask and a final reference mask includes:

[0035] Obtaining the inference intermediate mask and the reference intermediate mask under each preset neighborhood area;

[0036] Collaboratively voting on the inference intermediate mask under each preset neighborhood area to obtain the final inference mask;

[0037] Collaborative voting is performed on the reference intermediate mask under each preset neighborhood area to obtain the final reference mask.

[0038] In a second aspect, the present invention further provides a zero-sample anomaly detection device based on a vision and language model, comprising:

[0039] An acquisition module, used to acquire an inference image and a reference image from the same batch of images, wherein the inference image represents an image to be detected for abnormality, and the reference image includes other images in the same batch of images except the inference image;

[0040] A first detection module, configured to perform bimodal detection on the reasoning image and the reference image respectively through a first multimodal model to obtain an initial mask, wherein the initial mask includes an initial reasoning anomaly mask and an initial reference anomaly mask;

[0041] A second detection module is used to extract features of the same batch of images respectively through a second multimodal model and perform feature aggregation respectively to obtain aggregated inference features and aggregated reference features;

[0042] A noise feature filtering module, configured to perform noise feature filtering on the aggregated reference feature according to the initial mask to obtain a non-abnormal aggregated reference feature, and obtain an initial anomaly score according to the non-abnormal aggregated reference feature and the aggregated reasoning feature, wherein the initial anomaly score includes an inference anomaly score and a reference anomaly score;

[0043] A refining module is used to perform mask refinement according to the initial anomaly score and the initial mask to obtain a final reasoning mask, wherein the final reasoning mask is used to represent anomaly detection reasoning results.

[0044] In a third aspect, the present invention further provides an electronic device, comprising a memory and a processor;

[0045] The memory is used to store computer programs;

[0046] The processor is used to implement the zero-sample anomaly detection method based on vision and language model as described in the first aspect when executing the computer program.

[0047] The beneficial effects of the zero-sample anomaly detection method based on vision and language model of the present invention are:

[0048] The inference image and reference image are obtained from the same batch of images. The similarity between the images is used to directly perform anomaly detection without additional training data, which is conducive to zero-shot anomaly detection in the absence of samples in a specific field. The first multimodal model is used to perform bimodal detection on the inference image and the reference image to generate an initial mask, which combines multimodal information such as images and text to enhance local semantic representation, facilitate fine-grained anomaly segmentation, and provide a basis for subsequent refinement. The second multimodal model is used to extract and aggregate features of the same batch of images to capture information at different scales and levels, improve the comprehensiveness and accuracy of detection, and enhance the model's ability to understand complex patterns. According to the initial mask, the noise in the aggregated reference features is filtered, the information of normal patterns is retained, the influence of noise is reduced, the robustness and reliability of the model are improved, the model can focus on real anomalies, and the detection accuracy is improved. The inference image is compared with the purified reference features to obtain a reliable anomaly score, evaluate the possibility of anomalies in each pixel or region, and provide a basis for further optimization. Mask refinement is performed based on the anomaly score and the initial mask, and finally an inference mask that accurately reflects the anomaly detection results is obtained, achieving efficient and accurate zero-shot anomaly detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 Schematic diagram of a zero-sample anomaly detection method based on a visual and language model according to an embodiment of the present invention;

[0050] Figure 2 4 is a flowchart of step S200 of the zero-sample anomaly detection method based on vision and language model according to an embodiment of the present invention;

[0051] Figure 3 4 is a flowchart of steps S300-S500 of a zero-sample anomaly detection method based on a vision and language model according to an embodiment of the present invention;

[0052] Figure 4 An exemplary diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0053] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below in conjunction with the accompanying drawings. Although certain embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be interpreted as being limited to the embodiments described herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present invention. It should be understood that the drawings and embodiments of the present invention are only for exemplary purposes and are not intended to limit the scope of protection of the present invention.

[0054] It should be understood that the various steps described in the method embodiments of the present invention may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.

[0055] The term "including" and its variations used in this document are open inclusions, that is, "including but not limited to"; the term "based on" means "based at least in part on"; the term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one other embodiment"; the term "some embodiments" means "at least some embodiments"; the term "optionally" means "optional embodiments". The relevant definitions of other terms will be given in the following description. It should be noted that the concepts of "first", "second", etc. mentioned in the present invention are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0056] It should be noted that the modifications of "one" and "plurality" mentioned in the present invention are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise clearly indicated in the context, it should be understood as "one or more".

[0057] The names of the messages or information exchanged between multiple devices in the embodiments of the present invention are only used for illustrative purposes, and are not used to limit the scope of these messages or information.

[0058] Anomaly detection is widely used in different fields and plays an important role. It is used to identify anomalies in samples and determine their locations. Traditional anomaly detection methods mainly adopt unsupervised paradigms and rely on a large number of normal samples for training. However, these methods consume a lot of computing resources and have poor generalization capabilities.

[0059] Zero-shot anomaly detection has begun to attract attention and can be directly inferred without target domain data. The development of visual-language pre-trained models and their strong zero-shot generalization capabilities have promoted the application of such technologies in zero-shot anomaly detection. Multimodal models are used for anomaly detection to estimate anomaly probabilities by the similarity of image and text features. However, some studies have pointed out that the original qk attention mechanism affects local semantics and limits the ability of fine-grained anomaly segmentation. In order to strengthen the local semantic representation, the vv attention mechanism was introduced.

[0060] Although the improvement of the self-attention mechanism has improved performance, it has weakened the generalization ability of the model to some extent. Since most of the pixels in the test dataset are normal, the entire test dataset is used for feature matching for anomaly detection. In actual production environments, testing is generally performed on a batch of samples. It is not practical to use the entire test dataset and requires a lot of computing resources. In addition, the prior knowledge of the distribution of the test dataset before testing requires post-optimization. Mainstream anomaly detection methods usually involve fine-tuning the model on one dataset and then testing it on another dataset. This approach may pose a risk of data leakage.

[0061] In order to solve the above problems, the present invention proposes a zero-sample anomaly detection method without training, which combines the principle of zero-sample anomaly detection with the idea of ​​feature matching. By considering images in the same batch as mutual references, limited reference information can be collected. This method utilizes the relationship between images in the same batch, avoids the traditional training process, and ensures the effectiveness of anomaly detection. In other words, with the help of comparison between images in the batch, the anomaly detection method of the present invention can complete the anomaly detection task without additional training.

[0062] In view of the problems existing in the above-mentioned related technologies, this embodiment provides a zero-sample anomaly detection method, device and equipment based on vision and language models.

[0063] like Figure 1 As shown, an embodiment of the present invention provides a zero-sample anomaly detection method based on a vision and language model, comprising:

[0064] Step S100, obtaining an inference image and a reference image from the same batch of images, wherein the inference image represents an image to be detected for abnormality, and the reference image includes other images in the same batch of images except the inference image.

[0065] From the same batch of images, the inference image and the reference image to be detected are obtained, wherein, to ensure zero-sample detection, the inference image and the reference image in the embodiment of the present invention are both selected from the same batch of images. For example, when there are five images in the same batch, when image 1 is used as the inference image, at least one of image 2, image 3, image 4 and image 5 is used as the reference image; when image 2 is used as the inference image, at least one of image 1, image 3, image 4 and image 5 is used as the reference image.

[0066] like Figure 2 As shown, in step S200, the inference image I is respectively u and the reference image I v Perform bimodal detection to obtain an initial mask, wherein the initial mask includes an initial reasoning anomaly mask M uand the initial reference exception mask M v .

[0067] The multimodal model refers to a model that can process and integrate input information from two or more different data types (ie, modalities). In an embodiment of the present invention, the multimodal model is used to process input data of text type and image type.

[0068] The inference image and the reference image are detected by the first multimodal model, and the corresponding initial inference anomaly mask and initial reference anomaly mask are obtained respectively. The initial inference anomaly mask and the initial reference anomaly mask are used to represent the probability of the inference image and the reference image being classified as normal or abnormal. The proxy layer is used to restore the semantic relevance between patches, and its attention weight comes from the intermediate layer. u and M v denote the inference mask and reference mask respectively, and VisualEncoder denotes the visual encoder.

[0069] Optionally, the first multimodal model is an improved CLIP model, namely, a SeCLIP model.

[0070] like Figure 3 As shown, in step S300, features of the same batch of images are extracted respectively through a second multimodal model and feature aggregation is performed respectively to obtain aggregated inference features and aggregated reference features.

[0071] The image data is processed using an optimized or adjusted second multimodal model to extract visual features from the same batch of images. These features include not only global information but also local information at different levels, which helps to understand the image content more comprehensively and better capture local details and semantic information in the image.

[0072] Since the outputs of different layers of the multimodal model have different attentions, in order to understand the local semantics and increase the generalization of the model, feature aggregation is performed on local information of different scales at different stages (outputs of different layers) to generate richer feature representations to enhance the robustness and richness of feature representations, thereby improving the accuracy of anomaly detection. The reference image provides a benchmark for the normal mode. By comparing the aggregated features of the inference image and the reference image, abnormal areas can be more effectively identified.

[0073] Optionally, the second multimodal model is an improved CLIP model, namely, a FiCLIP model.

[0074] Step S400: According to the initial mask, the aggregated reference features are subjected to noise feature filtering to obtain non-abnormal aggregated reference features, and an initial anomaly score is obtained according to the non-abnormal aggregated reference features and the aggregated inference features, wherein the initial anomaly score includes an inference anomaly score and a reference anomaly score.

[0075] When applying feature matching methods for anomaly detection, it is necessary to ensure that the reference image does not contain any patches similar to the anomaly patches in the reasoning image. However, when applied to zero-shot anomaly detection, the uncertainty of whether the reference feature is normal may introduce some noise. The similarity between the anomaly patches in the reference and reasoning images leads to a lower anomaly score for the reasoning image, which may lead to an increase in the false negative rate in anomaly detection. Therefore, the abnormal features in the aggregated reference features are filtered out, and the filtered non-anomalous aggregated reference features are used as references to calculate the initial anomaly score.

[0076] The initial mask is used to filter out abnormal features in the reference image so that the reference features only contain information about normal patterns, thereby reducing the impact of noise on subsequent calculations. By comparing the inference image with clean, non-abnormal reference features, abnormal areas can be identified more accurately, avoiding false positives and false negatives. After removing the noise features, the model can better capture the true abnormal information, improving the robustness and reliability of the overall detection. The initial anomaly score provides anomaly probability distribution, which helps to more accurately locate and describe abnormal areas in subsequent steps.

[0077] Step S500, performing mask refinement according to the initial anomaly score and the initial mask to obtain a final reasoning mask, wherein the final reasoning mask is used to represent anomaly detection reasoning results.

[0078] The initial anomaly score is calculated by comparing the difference between the inference image and the non-anomalous reference features, indicating the possibility of anomaly at each pixel or patch location. This score provides a preliminary estimate of potential abnormal areas in the image, helping to quickly locate areas that may have problems and lay the foundation for subsequent analysis. The generated initial mask serves as an important input for the subsequent refinement step to guide how to optimize the anomaly detection results.

[0079] Combining anomaly scores of different neighborhood sizes and at different stages can capture multi-scale and hierarchical information, which helps to enhance the comprehensiveness and accuracy of detection. The binary mask generated based on the preliminary anomaly score indicates which areas may be abnormal, guiding the noise feature filtering and refinement process. By filtering out abnormal features in the reference image, the reference features only contain information about normal patterns, reducing the impact of noise on subsequent calculations and improving the robustness of the model.

[0080] The initial mask provides a starting point for the refinement step, allowing the model to focus on the areas most likely to be abnormal, thereby optimizing more efficiently. Through preliminary screening, most of the obviously normal areas can be excluded, reducing the probability of false positives and making the final results more reliable. This approach not only simplifies the processing flow, but also improves detection efficiency and the quality of results.

[0081] In this embodiment, by obtaining the inference image and the reference image from the same batch of images, the similarity between the images is used to directly perform anomaly detection without additional training data, which is conducive to achieving zero-sample anomaly detection in the absence of samples in a specific field.

[0082] like Figure 2 As shown, the inference image I is respectively processed by the first multimodal model u and the reference image I v Perform bimodal detection, obtain the initial mask, and use a multimodal model to combine image and text features (i.e. Figure 2 The generated initial mask provides a preliminary estimate of the anomaly location, laying the foundation for subsequent refinement steps and reducing computing resource consumption.

[0083] The second multimodal model extracts the features of the same batch of images and performs feature aggregation. Feature aggregation can capture information of different scales and levels, improving the comprehensiveness and accuracy of detection. This process strengthens the model's ability to understand complex patterns in images, which is conducive to more accurate identification of anomalies.

[0084] According to the initial mask, the aggregated reference features are filtered for noise features to obtain non-abnormal aggregated reference features. After filtering out abnormal features in the reference image, the aggregated reference features only contain information about normal patterns, which reduces the impact of noise on subsequent calculations and improves the robustness and reliability of the model. This step enables the model to focus on real abnormal situations and improve detection accuracy.

[0085] An initial anomaly score is obtained based on the non-abnormal aggregated reference features and the aggregated inference features. By comparing the inference image with the purified reference features, a more reliable anomaly score, namely the inference anomaly score and the reference anomaly score, can be obtained. These scores are used to evaluate the probability of anomaly in each pixel or region, providing a basis for further optimization.

[0086] The mask is refined according to the initial anomaly score and the initial mask to obtain the final inference mask. The refining process further optimizes the anomaly detection results, excludes most of the obviously normal areas, reduces the probability of false positives, and makes the final results more reliable. The final inference mask accurately represents the inference results of anomaly detection, and realizes efficient and accurate zero-shot anomaly detection.

[0087] Optionally, before performing bimodal detection on the reasoning image and the reference image respectively through the first multimodal model to obtain an initial mask, the method further includes:

[0088] Determine the initial multimodal model.

[0089] The residual connections and feed-forward modules in the initial multimodal model are removed.

[0090] The attention features of the intermediate layers of the initial multimodal model are weighted to replace the attention features of the last layer to obtain the first multimodal model, wherein the intermediate layers include specific layers in the initial multimodal model.

[0091] In one embodiment, the initial multimodal model is a CLIP model. Since CLIP is trained for classification tasks, it pays more attention to global semantic information. However, abnormal classification and segmentation require understanding of local semantics. Therefore, it is necessary to restore its semantic relevance to local patterns. The ViT-based CLIP visual encoder consists of a series of attention blocks. Each block generates a feature representation X i ∈R B ×(HW+1)×D , where i represents the layer index, B represents the batch size, D represents the dimension, and HW represents the number of local patch markers.

[0092] To maintain the generalization ability of CLIP, the changes are limited to the last layer. By visualizing the attention map of CLIP, it can be found that the attention map of the middle layer has better local semantics. inter Replace the attention map Attn of the last layer -1 :

[0093] Attn -1 =Attn inter , X attn =Proj(Attn inter·v-1 ),

[0094] Among them, Proj represents the projection layer.

[0095] The residual connection has a significant negative impact on the performance of dense segmentation tasks, while the feed-forward module has little effect. In an embodiment of the invention, the output X of the last block layer final ∈R B×(HW+1)×DIt can be defined as:

[0096] X final =X attn ,

[0097] Among them, X attn Represents the feature representation after being processed by the modified attention mechanism.

[0098] Optionally, before performing bimodal detection on the reasoning image and the reference image respectively through the first multimodal model to obtain an initial mask, the method further includes:

[0099] Construct normal text prompt templates and abnormal text prompt templates.

[0100] The normal text features and the abnormal text features in the normal text prompt template and the abnormal text prompt template are respectively extracted through the first multimodal model.

[0101] Optionally, the text features include the normal text features and the abnormal text features, and the step of performing bimodal detection on the reasoning image and the reference image respectively through the first multimodal model to obtain the initial mask includes:

[0102] The reasoning image features and the reference image features of the reasoning image and the reference image are respectively extracted through the first multimodal model.

[0103] The inference similarity and the reference similarity between the inference image feature and the text feature, and between the reference image feature and the text feature are calculated respectively.

[0104] The initial mask is determined according to the inferred similarity and the reference similarity.

[0105] In order to utilize CLIP's multimodal capabilities for zero-shot anomaly detection, the similarity between text features and visual features is calculated to determine whether an anomaly is present. Anomaly detection involves the concepts of normal and abnormal, so the text prompt template can be designed as "an [extra information] of [abnormal category] [image category] image", where the extra information represents additional information about the image, such as rotation or cropping; the abnormal category represents whether the image is abnormal, including abnormal or normal; the image category represents the category of the image, such as human image and object image.

[0106] The text features of the text prompt template are extracted using the text encoder of CLIP to obtain normal text features and abnormal text features respectively, and the final text feature F is obtained by averaging multiple normal text features and multiple abnormal text features respectively. t ∈R 2×C , where C represents the dimension of the feature. For anomaly classification, the global visual feature F of the original CLIP is usedc ∈R B×C (i.e., the inference image features and the reference image features. When the processed image is an inference image, F c is the inference image feature), the probability Cls that the image is classified as normal or abnormal prob ∈R B×2 It can be expressed as:

[0107] CLS prob =Softmax(F c ·F t ),

[0108] Anomaly classification score CLS score ∈R B Indicates the probability that the image category is abnormal. For abnormal segmentation, remove the [CLS] tag and map the dimension to C, and finally obtain the local patch feature F s ∈R B×HW×C . Abnormal segmentation probability SEG prob ∈R B×H×W×2 It can be expressed as:

[0109] SEG prob =Softmax(F s ·F t ),

[0110] Abnormal segmentation score Seg score ∈R B×HW Indicates the probability that a patch is an anomaly.

[0111] For two images in the same batch, SeCLIP-AD, the first multimodal model, calculates the similarity between text and visual features to generate the corresponding anomaly score CLS. prob and mask SEG prob , mask SEG prob This is the initial mask.

[0112] Alternatively, if Figure 3 As shown, the extracting features of the same batch of images respectively through the second multimodal model and performing feature aggregation respectively to obtain aggregated inference features and aggregated reference features includes:

[0113] For each image, a stage patch feature is extracted by the second multimodal model, wherein the stage patch feature is obtained by each layer output of the visual encoder of the second multimodal model.

[0114] The patch features of the stage with the same image number and the same number of visual encoder layers within a preset neighborhood area are aggregated to obtain aggregated features.

[0115] The aggregated inference feature is obtained according to the aggregated feature corresponding to the inference image, and the aggregated reference feature is obtained according to the aggregated feature corresponding to the reference image.

[0116] The second multimodal model is used to simultaneously extract B unlabeled test images D = {I u}, where u=1,...,B. For a CLIP encoder with L layers, a multi-stage patch feature is used where i∈{0,1,...,L} represents the layer of the CLIP visual encoder.

[0117] By aggregating neighboring features, the detection accuracy of anomalies of various sizes is improved. Specifically, patch marking It can be reshaped into H×W×D. Average pooling is used to aggregate the stage patch features within the r×r neighborhood of the current position to obtain the aggregated features after aggregation and Then reshape it into and Among them, u represents the number of the inference image, v represents the number of the reference image, i represents the number of visual encoder layers, and r represents the neighborhood area.

[0118] Among them, the stage patch features represent the features obtained by different layers of the visual encoder and In one embodiment, there are multiple preset neighborhood areas, for example, the preset neighborhood areas are r=1, 3, 5. That is, the preset neighborhood areas are 1, 3, and 5 to obtain patch marking features within the r×r neighborhood of the current position respectively.

[0119] Optionally, filtering the aggregated reference features for noise features according to the initial mask to obtain non-abnormal aggregated reference features, and obtaining an initial abnormality score according to the non-abnormal aggregated reference features and the aggregated reasoning features includes:

[0120] The aggregated reference features having abnormal features are filtered out according to the initial mask, and non-abnormal aggregated reference features and corresponding masks are retained.

[0121] The feature distance between the aggregated reasoning feature and the non-abnormal aggregated reference feature is calculated, and the minimum feature distance is used as the initial abnormality score.

[0122] The aggregated reference features with abnormal features are filtered out by the initial mask, which is expressed as:

[0123]

[0124] Where M∈R HW represents the mask, P a represents the abnormal probability, Pn represents normal probability.

[0125] Normal characteristics can be defined as:

[0126]

[0127] in, represents the filtered aggregate reference feature, N' represents the length of the normal patch, and M represents the mask after filtering.

[0128] The initial anomaly score is expressed as:

[0129]

[0130] in, represents the initial anomaly score.

[0131] Alternatively, if Figure 3 As shown, performing mask refinement according to the initial anomaly score and the initial mask to obtain a final reasoning mask includes:

[0132] In the inference image and the reference image, an average anomaly score is generated according to the extracted phase patch features, wherein the average anomaly score includes an inference average anomaly score and a reference average anomaly score.

[0133] The average anomaly score is averaged with the initial anomaly score to obtain a refined inference anomaly score and a refined reference anomaly score for each neighborhood area.

[0134] The refined inference anomaly score and the refined reference anomaly score exceeding a refinement threshold are binarized to obtain an inference intermediate mask and a reference intermediate mask.

[0135] The inference intermediate mask and the reference intermediate mask are respectively subjected to collaborative voting to obtain a final inference mask and a final reference mask.

[0136] Features extracted from multiple stages and different aggregation types And use the anomaly score to refine the anomaly mask M. Specifically, the patch features of all stages are used to generate an average anomaly score, that is, a refined inference anomaly score and a refined reference anomaly score. For example, in an embodiment of the present invention, the average anomaly score is generated by the patch features of four stages, which is expressed as:

[0137]

[0138] Among them, m represents the number of stages, u represents the number of inference images, v represents the number of reference images, and i represents the number of layers of the multimodal model, with a total of L layers. represents the refined inference anomaly score, Denotes the refined reference anomaly score.

[0139] The intermediate mask includes the inference intermediate mask and the reference intermediate mask. The intermediate mask is defined as:

[0140]

[0141] Among them, M inter Represents the intermediate mask, It represents the abnormal score of the image in the neighborhood r, μ represents the hyperparameter, i.e., the refinement threshold, and its value is set according to specific needs.

[0142] Optionally, the hyperparameter μ=0.57.

[0143] Then collaborative voting and the intermediate mask Minter are applied to refine the final inference mask M u and the final reference mask M v .

[0144] Optionally, the performing collaborative voting on the inference intermediate mask and the reference intermediate mask respectively to obtain a final inference mask and a final reference mask includes:

[0145] The inference intermediate mask and the reference intermediate mask under each preset neighborhood area are obtained.

[0146] Collaborative voting is performed on the inference intermediate mask under each preset neighborhood area to obtain the final inference mask.

[0147] Collaborative voting is performed on the reference intermediate mask under each preset neighborhood area to obtain the final reference mask.

[0148] The inference intermediate mask and the reference intermediate mask are calculated for each preset neighborhood area respectively, and the inference intermediate mask and the reference intermediate mask under each neighborhood area are collaboratively voted, and the final inference mask and the final reference mask are iterated in sequence.

[0149] An embodiment of the present invention provides a zero-sample anomaly detection device based on a vision and language model, comprising:

[0150] An acquisition module, used to acquire an inference image and a reference image from the same batch of images, wherein the inference image represents an image to be detected for abnormality, and the reference image includes other images in the same batch of images except the inference image;

[0151] A first detection module, configured to perform bimodal detection on the reasoning image and the reference image respectively through a first multimodal model to obtain an initial mask, wherein the initial mask includes an initial reasoning anomaly mask and an initial reference anomaly mask;

[0152] A second detection module is used to extract features of the same batch of images respectively through a second multimodal model and perform feature aggregation respectively to obtain aggregated inference features and aggregated reference features;

[0153] A noise feature filtering module, configured to perform noise feature filtering on the aggregated reference feature according to the initial mask to obtain a non-abnormal aggregated reference feature, and obtain an initial anomaly score according to the non-abnormal aggregated reference feature and the aggregated reasoning feature, wherein the initial anomaly score includes an inference anomaly score and a reference anomaly score;

[0154] A refining module is used to perform mask refinement according to the initial anomaly score and the initial mask to obtain a final reasoning mask, wherein the final reasoning mask is used to represent anomaly detection reasoning results.

[0155] like Figure 4 As shown, an electronic device 400 provided by an embodiment of the present invention includes a memory 410 and a processor 420; the memory 410 is used to store a computer program; the processor 420 is used to implement the xxx method described above when executing the computer program.

[0156] In other words, an electronic device 400 includes a memory 410 and a processor 420 coupled to the memory 410; the memory 410 is configured to store a computer program; and the processor 420 is configured to perform the following operations when executing the computer program:

[0157] Acquire a reasoning image and a reference image from the same batch of images, wherein the reasoning image represents an image to be detected for abnormality, and the reference image includes other images in the same batch of images except the reasoning image;

[0158] Using a first multimodal model, performing bimodal detection on the reasoning image and the reference image respectively to obtain an initial mask, wherein the initial mask includes an initial reasoning anomaly mask and an initial reference anomaly mask;

[0159] Extracting features of the same batch of images respectively through a second multimodal model and performing feature aggregation respectively to obtain aggregated inference features and aggregated reference features;

[0160] According to the initial mask, filtering the aggregated reference features for noise features to obtain non-abnormal aggregated reference features, and obtaining an initial anomaly score according to the non-abnormal aggregated reference features and the aggregated reasoning features, wherein the initial anomaly score includes a reasoning anomaly score and a reference anomaly score;

[0161] Mask refinement is performed according to the initial anomaly score and the initial mask to obtain a final reasoning mask, wherein the final reasoning mask is used to represent anomaly detection reasoning results.

[0162] An embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the xxx method described above is implemented.

[0163] In other words, a non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor performs the following operations:

[0164] Acquire a reasoning image and a reference image from the same batch of images, wherein the reasoning image represents an image to be detected for abnormality, and the reference image includes other images in the same batch of images except the reasoning image;

[0165] Using a first multimodal model, performing bimodal detection on the reasoning image and the reference image respectively to obtain an initial mask, wherein the initial mask includes an initial reasoning anomaly mask and an initial reference anomaly mask;

[0166] Extracting features of the same batch of images respectively through a second multimodal model and performing feature aggregation respectively to obtain aggregated inference features and aggregated reference features;

[0167] According to the initial mask, filtering the aggregated reference features for noise features to obtain non-abnormal aggregated reference features, and obtaining an initial anomaly score according to the non-abnormal aggregated reference features and the aggregated reasoning features, wherein the initial anomaly score includes a reasoning anomaly score and a reference anomaly score;

[0168] Mask refinement is performed according to the initial anomaly score and the initial mask to obtain a final reasoning mask, wherein the final reasoning mask is used to represent anomaly detection reasoning results.

[0169] An electronic device 400 that can be used as a server or client of the present invention will now be described, which is an example of a hardware device that can be applied to various aspects of the present invention. The electronic device 400 is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device 400 can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.

[0170] The electronic device 400 includes a computing unit, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) or a computer program loaded from a storage unit into a random access memory (RAM). In the RAM, various programs and data required for the operation of the device can also be stored. The computing unit, ROM, and RAM are connected to each other via a bus. An input / output (I / O) interface is also connected to the bus.

[0171] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, the storage medium can be a disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc. In the present application, the unit described as a separate component may or may not be physically separated, and the component displayed as a unit may or may not be a physical unit, that is, it may be located in one place, or it may be distributed on multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment of the present invention. In addition, each functional unit in each embodiment of the present invention can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0172] Although the present invention is disclosed as above, the protection scope of the present invention is not limited thereto. Those skilled in the art may make various changes and modifications without departing from the spirit and scope of the present invention, and these changes and modifications will fall within the protection scope of the present invention.

Claims

1. A zero-sample anomaly detection method based on vision and language model, characterized in that: include: Acquire a reasoning image and a reference image from the same batch of images, wherein the reasoning image represents an image to be detected for abnormality, and the reference image includes other images in the same batch of images except the reasoning image; Using a first multimodal model, performing bimodal detection on the reasoning image and the reference image respectively to obtain an initial mask, wherein the initial mask includes an initial reasoning anomaly mask and an initial reference anomaly mask; Extracting features of the same batch of images respectively through a second multimodal model and performing feature aggregation respectively to obtain aggregated inference features and aggregated reference features; According to the initial mask, filtering the aggregated reference features for noise features to obtain non-abnormal aggregated reference features, and obtaining an initial anomaly score according to the non-abnormal aggregated reference features and the aggregated reasoning features, wherein the initial anomaly score includes a reasoning anomaly score and a reference anomaly score; Mask refinement is performed according to the initial anomaly score and the initial mask to obtain a final reasoning mask, wherein the final reasoning mask is used to represent anomaly detection reasoning results.

2. The zero-sample anomaly detection method based on vision and language model according to claim 1, characterized in that: Before the first multimodal model is used to perform bimodal detection on the reasoning image and the reference image to obtain an initial mask, the method further includes: Determine the initial multimodal model; Removing residual connections and feed-forward modules in the initial multimodal model; The attention features of the intermediate layers of the initial multimodal model are weighted to replace the attention features of the last layer to obtain the first multimodal model, wherein the intermediate layers include specific layers in the initial multimodal model.

3. The zero-sample anomaly detection method based on vision and language model according to claim 1, characterized in that: Before the first multimodal model is used to perform bimodal detection on the reasoning image and the reference image to obtain an initial mask, the method further includes: Construct normal text prompt templates and abnormal text prompt templates; The normal text features and the abnormal text features in the normal text prompt template and the abnormal text prompt template are respectively extracted through the first multimodal model.

4. The zero-sample anomaly detection method based on vision and language model according to claim 3, characterized in that: The text features include the normal text features and the abnormal text features. The step of performing bimodal detection on the reasoning image and the reference image respectively through the first multimodal model to obtain the initial mask includes: extracting, by means of the first multimodal model, inference image features and reference image features of the inference image and the reference image respectively; respectively calculating the inference similarity and the reference similarity between the inference image feature and the text feature, and between the reference image feature and the text feature; The initial mask is determined according to the inferred similarity and the reference similarity.

5. The zero-sample anomaly detection method based on vision and language model according to claim 1, characterized in that: The extracting features of the same batch of images respectively through the second multimodal model and performing feature aggregation respectively to obtain aggregated inference features and aggregated reference features comprises: For each image, extracting a stage patch feature through the second multimodal model, wherein the stage patch feature is obtained by outputting each layer of the visual encoder of the second multimodal model; Performing feature aggregation on the patch features of the stage with the same image number and the same number of visual encoder layers within a preset neighborhood area to obtain aggregated features; The aggregated inference feature is obtained according to the aggregated feature corresponding to the inference image, and the aggregated reference feature is obtained according to the aggregated feature corresponding to the reference image.

6. The zero-sample anomaly detection method based on vision and language model according to claim 1, characterized in that: The step of filtering the aggregated reference features according to the initial mask to obtain a non-abnormal aggregated reference feature, and obtaining an initial abnormality score according to the non-abnormal aggregated reference feature and the aggregated reasoning feature comprises: filtering out the aggregated reference features having abnormal features according to the initial mask, and retaining non-abnormal aggregated reference features; The feature distance between the aggregated reasoning feature and the non-abnormal aggregated reference feature is calculated, and the minimum feature distance is used as the initial abnormality score.

7. The zero-sample anomaly detection method based on vision and language model according to claim 1, characterized in that: The performing mask refinement according to the initial anomaly score and the initial mask to obtain a final reasoning mask includes: In the inference image and the reference image, an average anomaly score is generated according to the extracted stage patch features, wherein the average anomaly score includes an inference average anomaly score and a reference average anomaly score; The average anomaly score is averaged with the initial anomaly score to obtain a refined reasoning anomaly score and a refined reference anomaly score for each neighborhood area; Binarizing the refined inference anomaly score and the refined reference anomaly score that exceed the refinement threshold to obtain an inference intermediate mask and a reference intermediate mask; The inference intermediate mask and the reference intermediate mask are respectively subjected to collaborative voting to obtain a final inference mask and a final reference mask.

8. The zero-sample anomaly detection method based on vision and language model according to claim 7, characterized in that: The step of performing collaborative voting on the inference intermediate mask and the reference intermediate mask to obtain a final inference mask and a final reference mask comprises: Obtaining the inference intermediate mask and the reference intermediate mask under each preset neighborhood area; Perform collaborative voting on the inference intermediate mask under each preset neighborhood area to obtain the final inference mask; Collaborative voting is performed on the reference intermediate mask under each preset neighborhood area to obtain the final reference mask.

9. A zero-sample anomaly detection device based on vision and language model, characterized in that: include: An acquisition module, used to acquire an inference image and a reference image from the same batch of images, wherein the inference image represents an image to be detected for abnormality, and the reference image includes other images in the same batch of images except the inference image; A first detection module, configured to perform bimodal detection on the reasoning image and the reference image respectively through a first multimodal model to obtain an initial mask, wherein the initial mask includes an initial reasoning anomaly mask and an initial reference anomaly mask; A second detection module is used to extract features of the same batch of images respectively through a second multimodal model and perform feature aggregation respectively to obtain aggregated inference features and aggregated reference features; A noise feature filtering module, configured to perform noise feature filtering on the aggregated reference feature according to the initial mask to obtain a non-abnormal aggregated reference feature, and obtain an initial anomaly score according to the non-abnormal aggregated reference feature and the aggregated reasoning feature, wherein the initial anomaly score includes an inference anomaly score and a reference anomaly score; A refining module is used to perform mask refinement according to the initial anomaly score and the initial mask to obtain a final reasoning mask, wherein the final reasoning mask is used to represent anomaly detection reasoning results.

10. An electronic device, characterized in that: including memory and processor; The memory is used to store computer programs; The processor is used to implement the zero-sample anomaly detection method based on vision and language model according to any one of claims 1 to 8 when executing the computer program.

Citation Information

Patent Citations

  • Knowledge distillation method based on multi-teacher model

    CN118379568A

  • Zero sample anomaly detection method and device based on image-text pre-training model

    CN118864876A

  • Zero sample anomaly detection method based on multi-mode learnable prompt

    CN118865000A

  • Zero sample image anomaly detection method and device

    CN119130931A

  • Training method for image classification model, image classification method and related apparatus

    WO2024188017A1