Multi-modal few-shot data-driven anomaly detection method, system and storage medium
By transforming the test images without abnormal diffusion model and combining multiple comparison methods, the problem of visual promotion difficulties and unstable results of small sample anomaly detection is solved, and a more accurate and stable abnormal detection effect is achieved.
Patent Information
- Application Number
- CN202411687552.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-25
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2044-11-25
AI Technical Summary
The existing small sample anomaly detection methods are difficult to generalize visually to other complex fields, and due to the small number of reference images, the accuracy and stability of the results are insufficient.
The test image is converted into a normal image with a series of normal distributions through an anomaly-free diffusion model, and multiple comparisons are performed in combination with text description and normal image sample sets, and the comparison scores are fused to obtain the final prediction results.
It realizes more accurate and stable abnormality detection under small sample data, and can be applied to multiple complex fields, such as medical diagnosis.
Smart Images

Figure CN119205736B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of medical image processing, and particularly relates to an abnormal detection method, system and storage medium driven by multi-modal small sample data. Background Art
[0002] Anomaly detection (AD) has attracted wide attention due to its wide applicability in various fields such as industrial defect detection, medical diagnosis, video surveillance, and manufacturing inspection. This high level of interest is mainly attributed to the fact that it only relies on positive samples and adopts an unsupervised learning paradigm, in which it usually learns from the distribution of positive samples and identifies anomalies by detecting outliers. Traditional AD methods include those based on autoencoders, GANs, and knowledge-based methods, etc. In addition, some diffusion-based methods have also emerged recently. Although most of these methods do not require annotated data, they require a large number of normal samples to be provided during the training phase to effectively capture the distribution of normal samples. This requirement for a large amount of data also severely restricts its development in various fields.
[0003] Recently, many works have been explored in the small sample field. Especially with the rise of large vision-language pre-trained models (VLM), many works on small sample anomaly detection based on VLM have achieved state-of-the-art results. WinCLIP first uses the pre-trained CLIP to complete small sample anomaly detection by carefully designing text prompts and image feature comparison. AnomalyGPT eliminates the drawback of requiring manual threshold setting and supports multi-round conversations. InCTRL realizes general anomaly detection by leveraging in-context learning.
[0004] However, the latest such small sample anomaly detection works in the visual aspect are all about making rough direct matches and feature comparisons between several normal images (reference images) and test samples, which will affect the accuracy of the results, and thus it is difficult to generalize the samples to other complex fields. Because the direct comparison between the test image and several small sample reference images may have large differences, the prediction results obtained for each test sample will be affected by the position, angle, size, color, etc. of the randomly selected reference images (several small sample normal images). Secondly, due to the small number of reference images, just a few images as independent references are not representative enough, and of course, their feature distributions are not deeply mined, making it difficult to meet the accurate feature comparison and prone to unstable results.
[0005] Therefore, for the anomaly detection task, how to design models and methods to expand the sample size of small sample data so that the model can use small sample data as reference images to obtain accurate results is still an important research topic in this field. Summary of the Invention
[0006] In view of the problems of the prior art, the present invention provides a multi-modal few-shot data-driven anomaly detection method, system, and storage medium.
[0007] A multi-modal few-shot data-driven anomaly detection system, comprising:
[0008] An input module, configured to input a test image and a positive sample text description of the corresponding category;
[0009] An image customization module, configured to use a no-anomaly diffusion model to transform the test image into a series of normally distributed normal images to obtain customized images;
[0010] A prediction module, configured to compare the test image with the customized image, compare the test image with the text description, and compare the test image with a normal image sample set respectively to obtain three comparison scores, and fuse the three comparison scores to obtain a final prediction result.
[0011] Preferably, the test image and the normal image are medical diagnostic images.
[0012] Preferably, the no-anomaly diffusion model is a model parameterized by θ for category o , and is constrained and fine-tuned by a set C of normal image-text demonstration pairs o to train the no-anomaly diffusion model, and the formula is as follows:
[0013]
[0014] where θ is the parameter of the customized no-anomaly diffusion model for category o, o is the category, x n,t is the input normal reference image the latent version after t noise addition steps, c is the short text description of the normal image x n , C o is the image-text demonstration pair, is the expected value, ([[]] , ) is the normal image and text pair, is the standard Gaussian noise added to the noisy image, and t is the noise addition time step.
[0015] Preferably, in the image customization module, the method for transforming the test image into a series of normally distributed normal images includes the following steps:
[0016] Step 1, passing the test image through the image encoder of the no-anomaly diffusion model and then through a noise addition process to obtain a noisy data z t ;
[0017] Step 2, add the positive sample text description of the corresponding category as a condition to guide the non-abnormal diffusion model to denoise z t gradually;
[0018] Step 3, convert the denoised z t into the customized image through an image decoder.
[0019] Preferably, in the prediction module, three comparison scores are obtained through the CLIP model.
[0020] Preferably, the comparison score S obtained by comparing the test image with the customized image p is calculated as follows:
[0021]
[0022] where, l represents the number of layers, and n represents the number of feature extraction blocks. represents the cosine similarity function, represents the x th l layer feature extracted from the test image represents the feature extracted from the customized image;
[0023] And / or, extract multi-level features from the normal sample set using the same CLIP image encoder, and store these features in the repository M; the comparison score S obtained by comparing the test image with the normal image sample set N is calculated as follows:
[0024]
[0025] where, m represents the sample features sampled from the repository M, represents the l th layer feature;
[0026] And / or, the comparison score S obtained by comparing the test image with the normal image sample set text is calculated as follows:
[0027]
[0028] where, represents obtaining a set of normal and abnormal prompts through the CLIP text encoder, and softmax(·) is the average softmax score across multiple levels.
[0029] Preferably, the formula for fusing the three comparison scores is:
[0030]
[0031] Among them, is the fused score, and α and β are weights respectively.
[0032] Preferably, the average value of the normal image is used as the threshold to obtain the final prediction result.
[0033] The present invention also provides a method for anomaly detection by applying the above multi-modal small sample data-driven anomaly detection system, including the following steps:
[0034] Input the test image into the anomaly-free diffusion model to obtain a text description;
[0035] Convert the test image into a series of normally distributed normal images to obtain a customized image;
[0036] Compare the test image with the customized image, compare the test image with the text description, and compare the test image with the normal image sample set respectively to obtain three comparison scores, and fuse the three comparison scores to obtain the final prediction result.
[0037] The present invention also provides a computer-readable storage medium, on which there is stored: a computer program for implementing the above method.
[0038] The present invention constructs a preferred model structure and method for small sample data-driven anomaly detection. Specifically, more series of normally distributed normal images are constructed through test images, thus effectively expanding the normal image samples that can be compared. In the process of obtaining the final anomaly detection result, the present invention combines three aspects of information, respectively comparing the test image with the customized image, comparing the test image with the text description, and comparing the test image with the normal image sample set, and finally fusing the three comparison results to obtain the final prediction result.
[0039] According to the technical solution of the present invention, the anomaly detection of small sample volume tasks can be more accurately realized, and due to the use of triple comparison anomaly reasoning, the method has better stability and robustness. For example, when applied in the field of medical diagnosis, the present invention can use only a few (small samples) unlabeled medical normal images for the model to learn, and can efficiently, accurately and stably identify different modalities (such as: MRI, CT, OCT), different categories of medical diseases (such as: fundus, brain, lungs), and can accurately segment their parts. Therefore, the present invention has good application prospects.
[0040] Obviously, based on the above content of the present invention, according to the common general technical knowledge and conventional means in the art, without departing from the above basic technical idea of the present invention, various other forms of modifications, substitutions or changes can also be made.
[0041] The following is a further detailed description of the above content of the present invention in the form of specific embodiments. However, this should not be construed as limiting the scope of the above subject matter of the present invention to the following examples. Any technology implemented based on the above content of the present invention falls within the scope of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is a schematic flowchart of an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] It should be specifically noted that the algorithms for steps such as data acquisition, transmission, storage, and processing not specifically described in the embodiments, as well as the hardware structures and circuit connections not specifically described, can be implemented through the content disclosed in the prior art.
[0044] Embodiment: Anomaly Detection Method and System Driven by Small-Sample Data
[0045] The system of this embodiment includes:
[0046] An input module, configured to input a test image and a positive sample text description of the corresponding category;
[0047] An image customization module, configured to use a non-anomaly diffusion model to convert the test image into a series of normal images with a normal distribution to obtain a customized image;
[0048] A prediction module, configured to compare the test image with the customized image, compare the test image with the text description, and compare the test image with a normal image sample set respectively to obtain three comparison scores, and fuse the three comparison scores to obtain a final prediction result.
[0049] The system of this embodiment can be used for anomaly detection in various fields. Taking the detection of medical images as an example, only a small number of (small-sample) unlabeled medical normal images can be used for the model to learn, and different modalities (such as MRI, CT, OCT) and different categories of medical diseases (such as fundus, brain, lungs) can be efficiently, accurately, and stably identified, and their parts can be accurately segmented.
[0050] Specifically, the method for performing anomaly detection using the above system is as Figure 1 shown, including the following steps:
[0051] Step S1, input the test image into a non-anomaly diffusion model to obtain a text description:
[0052] This method is first based on the stable diffusion model, and then a diffusion model without anomalies is customized based on Dreambooth (trained using a small number of normal images and their text descriptions) to mine the distribution of reference images. To better align the category with its text description, a series of demonstrations Co are provided for the diffusion model, pairing small samples of reference images with the corresponding text of a specific class.
[0053] The specific steps are as follows: For each class, a series of small-sample reference images (i.e., 2 - 8 normal images) and their matching text c are prepared. For example: A normal [category] picture, where [category] can be replaced by retina, brain, skin, etc., depending on the category of the input picture, for the customized training of the diffusion model.
[0054] The diffusion model without anomalies is a model parameterized by θ for category o , and is constrained and fine-tuned by a series of normal image-text demonstration pairs C o to train the diffusion model without anomalies, and the formula is as follows:
[0055]
[0056] where θ is the parameter of the customized diffusion model without anomalies for category o, o is the category, x n,t is the input normal reference image the latent version after t noise addition steps, c is the short text description of the normal image x n , C o is the image-text demonstration pair, is the expected value, ([[]] , ) is the normal image and text pair, ϵ is the standard Gaussian noise added to the noisy image, and t is the noise addition time step.
[0057] In summary of this embodiment, the dimensions of the image and the latent image are the same throughout the process. Therefore, the diffusion model without anomalies learns the distribution of normal images under the text prompt c, and can thus be used in the subsequent steps to transform customized images.
[0058] Step S2, convert the test image into a series of normal images with a normal distribution to obtain a customized image:
[0059] In this embodiment, a normal custom conversion strategy is proposed to convert a test image into a positive sample distribution similar to it. It adaptively preserves the features of its original normal part, while the abnormal part gradually transforms into the normal part. These processes are all completed using the anomaly-free diffusion model obtained in the previous step. Specifically, first, text prompts are designed to maximize the retention of information in the normal regions of the test image while converting the abnormal regions into a normal state. To reduce the influence of other factors (such as contrast, image quality), these prompts are used to simulate all potential normal states, and a template list is planned for various potential image physical states.
[0060] The specific conversion process is as follows:
[0061] Step 1, pass the test image through the image encoder of the anomaly-free diffusion model, and then through a noise addition process to obtain a noisy data z t ;
[0062] Step 2, add the positive sample text description of the corresponding category as a condition to guide the anomaly-free diffusion model to gradually denoise z t ;
[0063] Step 3, pass the denoised z t through the image decoder to convert it into the custom image.
[0064] Step S3, compare the test image with the custom image, compare the test image with the text description, and compare the test image with the normal image sample set respectively to obtain three comparison scores, and fuse the three comparison scores to obtain the final prediction result:
[0065] This step obtains three comparison scores through the CLIP model.
[0066] Comparison between the test image and the custom image: In this embodiment, the image encoder of the CLIP model is divided into n multi-feature extraction blocks to extract multi-layer features, and then the comparison results are compared separately on the multi-layer features. For the test image x and its corresponding custom image , this CLIP image encoder extracts their features and respectively, and finally the comparison score is calculated as follows:
[0067]
[0068] Among them, l represents the number of layers, and n represents the number of feature extraction blocks. represents the cosine similarity function, represents the x extracted from the test imagel Layer features represent features extracted from the customized image;
[0069] Comparison between the test image and the normal image sample set: The data pool (normal image sample set) of normal images is a non - abnormal sample pool, which consists of normal reference images and generated normal images, better representing the distribution of positive samples. Among them, the generated normal images are normal samples generated from the above non - abnormal diffusion model and are incorporated into the data pool for this step of the prediction task because many previous studies have demonstrated the ability of diffusion models to synthesize high - fidelity images. Then, multi - level features are extracted from the non - abnormal samples in the data pool using the same CLIP image encoder, and these features are stored in repository M. The prediction score S between the test image and the data pool samples N can be expressed as:
[0070]
[0071] where m represents the sample features sampled from repository M, represents the features of the l layer;
[0072] Comparison between the test image and the text description: To calculate the anomaly score between the test image and the text prompt, the text prompt is divided into two types: normal text and abnormal text. The goal is to cover more possible states for these objects to better simulate various potential conditions of the image. Specifically, the text feature F text ∈R2×d is a set of normal and abnormal prompts obtained through the CLIP text encoder and then compared with the features in the test image:
[0073]
[0074] where represents a set of normal and abnormal prompts obtained through the CLIP text encoder, and softmax(·) is the average softmax score across multiple levels.
[0075] The final prediction result is obtained by combining the scores of the above three branches. The formula for fusing the three comparison scores is:
[0076]
[0077] where is the fused score, and α and β are weights respectively. As a preferred method, by default .
[0078] In this way, each test image corresponds to an anomaly score , The higher the anomaly score, the more likely it is an abnormal (lesion) image. The higher the pixel-level anomaly value in a test image, the greater the likelihood of a lesion area in that region. In this embodiment, the average anomaly value of the normal sample set images is used as a threshold to obtain the final lesion detection and segmentation results.
[0079] This embodiment further verifies the effectiveness of the above system and method in the medical field, industrial field, and semantic field. In the medical field, we cover datasets of various modalities, such as magnetic resonance imaging (MRI), computed tomography (CT), and optical coherence tomography (OCT). Specifically, the medical datasets include OCT2017, BrainMRI, HeadCT, and RESC. In the industrial field, we utilize multiple datasets, including MVTec-AD, Visa, KSDD, AFID, and ELPV. For semantic anomaly detection, we use two datasets: MNIST and CIFAR-10.
[0080] The performance of anomaly detection is evaluated by calculating the value of the area under the receiver operating characteristic curve (AUROC). The larger the AUROC value, the better the model performance.
[0081] We compared the method of this embodiment we proposed with the state-of-the-art anomaly detection (AD) methods reported in various existing literatures. The comparison methods include: PaDiM, PatchCore, RegAD, CoOp, WinCLIP, and InCTRL.
[0082] This embodiment verified the test set under the setting of selecting 2, 4, and 8 samples as the training set, with the resolution set to 240. All experiments were completed using PyTorch on an NVIDIA GeForce RTX 4090 GPU.
[0083] The experimental results are as follows:
[0084] 1. Compared with PaDiM, PatchCore, RegAD, WinCLIP, and InCTRL, the AUROC results of the method of this embodiment on the medical datasets (OCT2017, BrainMRI, HeadCT, and RESC) are as follows:
[0085] Table 1
[0086]
[0087] It can be concluded from the table that the method of this embodiment has improvements on all 4 medical datasets.
[0088] 2. The AUROC results of the method in this embodiment compared with the methods PaDiM, PatchCore, RegAD, WinCLIP, and InCTRL on industrial datasets (MVTec-AD, Visa, KSDD, AFID, and ELPV) are as follows:
[0089] Table 2
[0090]
[0091] It can be seen from the table that the method in this embodiment is the best on 5 industrial datasets in most cases in the industrial dataset.
[0092] 3. The AUROC results of the method in this embodiment compared with other methods on semantic datasets (MNIST and CIFAR-10) are as follows:
[0093] Table 3
[0094]
[0095] It can be seen from the table that the method in this embodiment is the best under various settings in the semantic data.
[0096] As can be seen from the above embodiments, the present invention constructs an anomaly detection method and system for small sample size tasks, which can achieve more accurate prediction performance on the premise of only applying a small number of normal image samples. Therefore, the present invention has good application prospects.
Claims
1. A multimodal small sample data driven anomaly detection system, characterized in that: include: The input module is configured to input a test image and a positive sample text description of the corresponding category; An image customization module is configured to transform the test image into a series of normal images with normal distribution using a non-abnormal diffusion model to obtain a customized image; A prediction module is configured to respectively compare the test image with the customized image, compare the test image with the text description, and compare the test image with a normal image sample set to obtain three comparison scores, and fuse the three comparison scores to obtain a final prediction result; The anomaly-free diffusion model is a model D parameterized by θ on class o. θ (x n,t ,c), through a series of normal image-text demonstration pairs composed of C o To constrain and fine-tune the training of the anomaly-free diffusion model, the formula is as follows: Where θ is the parameter of the customized anomaly-free diffusion model on category o, o is the category, x n,t is the input normal reference image x n After t noise addition steps, the potential version c is the normal image x n A brief text description of C o For image-text demonstration pairs, is the expected value, (x n ,c) is a normal image and text pair, ∈ is the standard Gaussian noise added to the noise image, and t is the noise adding time step.
2. The multimodal small sample data driven anomaly detection system according to claim 1, characterized in that: The test image and the normal image are medical diagnosis images.
3. The multimodal small sample data driven anomaly detection system according to claim 1, characterized in that: In the image customization module, the method for converting the test image into a series of normal images with normal distribution comprises the following steps: Step 1: Pass the test image through the image encoder of the non-abnormal diffusion model, and then go through the noise adding process to obtain a noisy data z t ; Step 2: Add the positive sample text description of the corresponding category as a condition to guide the non-abnormal diffusion model to z t Perform progressive denoising; Step 3: The denoised z t The image is converted into the customized image through the image decoder.
4. The multimodal small sample data driven anomaly detection system according to claim 1, characterized in that: In the prediction module, three comparison scores are obtained through the CLIP model.
5. The multimodal small sample data driven anomaly detection system according to claim 4, characterized in that: The comparison score S obtained by comparing the test image with the custom image p The calculation method is as follows: Where l represents the number of layers, n represents the number of feature extraction blocks, <·> represents the cosine similarity function, and F x,l represents the l-th layer feature extracted from the test image x, represents the features extracted from the custom image; And / or, using the same CLIP image encoder to extract multi-level features from the normal sample set, and storing these features in a repository M; comparing the test image with the normal image sample set to obtain a comparison score S N The calculation method is as follows: Among them, m represents the sample features sampled from the repository M, m l Represents the features of layer l; and / or, comparing the test image with a normal image sample set to obtain a comparison score S text The calculation method is as follows: Among them, F text represents a set of normal and abnormal prompts obtained through the CLIP text encoder, and softmax(·) is the average softmax score across multiple levels.
6. The multimodal small sample data driven anomaly detection system according to claim 5, characterized in that: The formula for fusing the three contrast scores is: in, is the score after fusion, and α and β are weights respectively.
7. The multimodal small sample data driven anomaly detection system according to claim 6, characterized in that: The normal image The average value of is used as the threshold to get the final prediction result.
8. A method for performing anomaly detection using the multimodal small sample data driven anomaly detection system according to any one of claims 1 to 7, characterized in that: The steps include: Inputting the test image into a non-abnormal diffusion model to obtain a text description; Converting the test image into a series of normal images with normal distribution to obtain a customized image; The test image is respectively compared with the customized image, the test image is compared with the text description, and the test image is compared with a normal image sample set to obtain three comparison scores, and the three comparison scores are fused to obtain a final prediction result.
9. A computer-readable storage medium, characterized in that: Stored thereon is: a computer program for implementing the method described in claim 8.