Metalearning and multi-modal large model-based few-sample anomaly detection method and system
Through the semantic-exception network based on meta-learning and multimodal large model, the semantic-exception network is constructed, which solves the problem of poor generalization ability of traditional methods in cross-category anomaly detection, and achieves higher detection accuracy and robustness.
Patent Information
- Application Number
- CN202411791492.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-06
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-06
AI Technical Summary
Traditional anomaly detection methods have poor generalization ability in cross-category anomaly detection, and cannot effectively convert semantic information into abnormal information, resulting in insufficient detection accuracy and robustness in the case of small sample data.
Using a small sample anomaly detection method based on meta-learning and multimodal large model, a semantic-anomaly network is constructed by dividing the auxiliary data set into multiple tasks, which includes a visual encoder, a language encoder, an image-image exception discriminator, an image-text anomaly detector and a fusion module, using these modules to capture exception information in multi-layer features and multimodal representations.
It improves the accuracy and robustness of image anomaly detection, enhances the model's ability to detect cross-category anomaly, and can achieve effective anomaly detection with less sample data.
Smart Images

Figure CN119939445A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to a few-sample anomaly detection method and system based on meta-learning and multimodal large models. Background Art
[0002] Anomaly detection plays a vital role in human daily life. It is widely used in various fields such as industrial inspection, medical analysis, security monitoring, etc. Traditional research usually proposes a single anomaly detection method for a specific category. Although these methods have achieved good results, they are limited to specific categories and require a large amount of training data. Since abnormal samples are often difficult to obtain, such methods are impractical in real-world applications. Therefore, it is very important to develop a general anomaly detection method.
[0003] The essence of general anomaly detection is to enable the model to achieve robust cross-category anomaly detection. In order to achieve the goal of general anomaly detection, existing work has tried various aspects such as feature extraction and anomaly learning. For example, the prior art proposed an exploration of general anomaly detection through category-agnostic few-shot through image alignment. With the development of pre-trained models, some large visual language models, such as CLIP, have shown strong few-shot capabilities. For another example, Winclip introduced a window-based industrial detection method based on the large pre-trained visual language model CLIP. By comparing visual data and language data, the powerful semantic information of CLIP is converted into anomaly information, which significantly improves the performance of Winclip in industrial heritage detection. Subsequently, InCTRL solved the limitation of Winclip that it is only applicable to one field. Residual learning is used to convert the powerful semantic information of CLIP into anomaly information, and a general anomaly detection model applicable to multiple fields is proposed.
[0004] Although these methods have achieved some results, there are still two major challenges in building effective general anomaly detection for various categories. The first challenge is how to enhance the ability of cross-category anomaly detection with less sample data. Since there are significant differences in anomaly information between different categories, the model tends to overfit the anomaly information of a single or specific category and cannot generalize across multiple categories, especially when there is little reference data. Winclip and InCTRL cannot achieve good results on every category. The second challenge is how to make full use of the semantic information in the pre-trained model while achieving a balance with the capture of abnormal information. Since large visual-language models are usually trained to extract features that represent general semantic information rather than abnormal information, directly using large visual-language models may not be helpful for anomaly detection. Only by solving the above two challenges can effective general anomaly detection be achieved. Summary of the invention
[0005] Technical problem to be solved by the present invention: In view of the above-mentioned problems in the prior art, a few-sample anomaly detection method and system based on meta-learning and multimodal large models are provided. The present invention aims to solve the problems of poor generalization ability of cross-category anomaly detection in traditional anomaly detection methods and inability to effectively convert semantic information into anomaly information, thereby improving the accuracy and robustness of image anomaly detection and the ability to detect cross-category anomalies.
[0006] In order to solve the above technical problems, the technical solution adopted by the present invention is: A few-sample anomaly detection method based on meta-learning and multimodal large model includes the following steps: S1: Divide the auxiliary dataset D={X,Y} containing N categories of detection target objects into N tasks according to the categories of the detection target objects. ~ The set of tasks , where X is the image sample and Y is the label. Divide into training set and test set , where the training set and test set Each query sample Each of them has normal prompt text, abnormal prompt text and reference samples for reference ; S2, constructing a semantic-anomaly network based on a multimodal large model, wherein the semantic-anomaly network includes a visual encoder, a language encoder, an image-image anomaly discriminator, an image-text anomaly detector and a fusion module. The visual encoder is used to convert the query sample And its reference samples The visual encoding is performed respectively to obtain image embedding, the language encoder is used to perform language encoding on the normal prompt text and the abnormal prompt text to obtain text embedding, and the image-image abnormality discriminator is used to perform language encoding on the normal prompt text and the abnormal prompt text to obtain text embedding according to the query sample. And its reference samples The image embedding is used to identify anomalies and obtain anomaly scores The image-text anomaly detector is used to perform anomaly discrimination based on image embedding and text embedding to obtain anomaly scores. , the fusion module is used to score the anomaly and anomaly score Add and fuse to get the final anomaly score ; S3, using task collections The training set for each task in and test set Training the semantic-anomaly network to obtain a trained semantic-anomaly network; S4, task collection Test sets for each task in Sample query in And its corresponding normal prompt text, abnormal prompt text and reference samples Input the trained semantic-anomaly network to obtain the anomaly score, and determine the query sample to be tested based on whether the anomaly score exceeds the preset threshold. Is it abnormal?
[0007] Optionally, the normal prompt text and the abnormal prompt text are query samples. The name of the detected target object replaces the query sample The detected target object category is embedded in the prompt text.
[0008] Optionally, the query sample And its reference samples The image embedding is obtained by performing visual encoding respectively, including: the visual encoder is the query sample Extract query samples separately Low-level features ,intermediate and advanced features , get the query sample Feature pyramid ; Visual encoder is the reference sample Extract reference samples separately Low-level features ,intermediate and advanced features , get the reference sample Feature pyramid .
[0009] Optionally, the language encoding of the normal prompt text and the abnormal prompt text to obtain text embedding includes: language encoding the normal prompt text by a language encoder to obtain a text embedding of the normal prompt text. , language encoding the exception prompt text to obtain the text embedding of the exception prompt text .
[0010] Optionally, the query sample And its reference samples The image embedding is used to identify anomalies and obtain anomaly scores Includes: For query samples Low-level features and reference samples Low-level features ,in Used to represent dimensions, and Respectively represent the number and dimension of features, For the reference sample quantity, the query sample Low-level features Each feature in Calculate its difference with the reference sample Low-level features similarity to identify k The most similar feature c, and then find the feature c and feature The cosine similarity between them is used to obtain the features The anomaly score is finally taken k The average of the largest anomaly scores is used as the anomaly score of the low-level feature. ; For query samples Intermediate features and reference samples Intermediate features ,in Used to represent dimensions, and Respectively represent the number and dimension of features, For the reference sample quantity, the query sample Intermediate features Each feature in Calculate its difference with the reference sample Intermediate features similarity to identify k The most similar feature c, and then find the feature c and feature The cosine similarity between them is used to obtain the features The anomaly score is finally taken k The average of the largest anomaly scores is used as the anomaly score of the intermediate feature. ; For query samples Advanced features of and reference samples Advanced features of ,in Used to represent dimensions, and Respectively represent the number and dimension of features, For the reference sample quantity, the query sample Advanced features of Each feature in Calculate its difference with the reference sample Advanced features of similarity to identify k The most similar feature c, and then find the feature c and feature The cosine similarity between them is used to obtain the features The anomaly score is finally taken k The average of the largest anomaly scores is used as the anomaly score of the high-level feature ; Finally, the abnormality score of the low-level features is , anomaly score of mid-level features and anomaly scores for advanced features Add up to get the anomaly score .
[0011] Optionally, the image-text anomaly detector includes a feature differentiator and a multi-layer perceptron, and the anomaly score is obtained by performing anomaly discrimination based on image embedding and text embedding. Including: Through the feature differentiator, first calculate the query sample Advanced features of and reference samples Advanced features of The difference between them is then mapped to the same latitude as the text embedding through a multi-layer perceptron to obtain the differential feature D, and then the normalized exponential function softmax is used to obtain the text embedding of the differential feature D belonging to the normal prompt text according to the following formula: and text embedding of exception prompt text Probability : , In the above formula, is an exponential function, is the cosine similarity function; the differential feature D is attributed to the text embedding of the normal prompt text and text embedding of exception prompt text Probability Obtaining anomaly scores through multi-layer perceptron .
[0012] Optionally, in step S3, the task set is used The training set for each task in and test set Training the semantic-anomaly network includes training for S rounds, and the training in each round includes: S3.1, initialize the category variable i to 0; S3.2, the tasks of the i-th category The training set Input into the semantic-anomaly network and calculate the loss according to the loss function; calculate the gradient of the loss and update the model parameters of the semantic-anomaly network internally , the model parameters of the semantic-anomaly network are updated once It refers to temporarily updating the model parameters of the semantic-anomaly network; S3.3, the tasks of the i-th category The test set Input to use model parameters The semantic-anomaly network of ; S3.4, add 1 to the category variable i, and determine whether the category variable i is less than N. If not, jump to step S3.2 to continue the training in this round; otherwise, calculate all the cumulative losses. The gradient of the semantic-anomaly network is updated externally once. , the external update once semantic-anomaly network model parameters It means actually updating the model parameters of the semantic-anomaly network once and saving them so that the model parameters saved in this round can be used in the next round of training of the semantic-anomaly network. .
[0013] In addition, the present invention also provides a few-sample anomaly detection system based on meta-learning and multimodal large models, comprising a microprocessor and a memory connected to each other, wherein the microprocessor is programmed or configured to execute the few-sample anomaly detection method based on meta-learning and multimodal large models.
[0014] In addition, the present invention also provides a computer-readable storage medium, which stores a computer program or instruction, and the computer program or instruction is programmed or configured to execute the few-sample anomaly detection method based on meta-learning and multimodal large model through a processor.
[0015] In addition, the present invention also provides a computer program product, including a computer program or instructions, which are programmed or configured to execute the few-sample anomaly detection method based on meta-learning and multimodal large model through a processor.
[0016] Compared with the prior art, the present invention mainly has the following advantages: the few-sample anomaly detection method based on meta-learning and multimodal large model of the present invention comprises dividing the auxiliary data set into a task set consisting of N tasks according to the category of the detection target object, dividing each task into a training set and a test set, constructing a semantic-anomaly network based on the multimodal large model, the semantic-anomaly network comprises a visual encoder, a language encoder, an image-image anomaly discriminator, an image-text anomaly detector and a fusion module, training the semantic-anomaly network with the training set and the test set of each task in the task set, and using the semantic-anomaly network to predict the anomaly score of the query sample to determine whether the query sample is abnormal, the present invention uses the image-image anomaly discriminator to analyze the difference between the query sample and the reference sample from multi-layer features to improve the accuracy and robustness of image anomaly detection; the present invention uses the image-text anomaly detector to discard the concept of category from the original prompt text, so that the model can extract abnormal information unrelated to the category, thereby improving the model's ability to detect cross-category anomalies. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 Schematic diagram of the basic flow of the method of the embodiment of the present invention.
[0018] Figure 2 1 is a network structure diagram of the semantic-anomaly network MetaSAN in an embodiment of the present invention.
[0019] Figure 3 It is a data flow diagram of the semantic-anomaly network MetaSAN in an embodiment of the present invention.
[0020] Figure 4 This is a training flow chart of the semantic-anomaly network MetaSAN in an embodiment of the present invention.
[0021] Figure 5 It is a test flow chart of the semantic-anomaly network MetaSAN in an embodiment of the present invention. DETAILED DESCRIPTION
[0022] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0023] The few-sample anomaly detection method based on meta-learning and multimodal large models in this invention aims to solve the problems of poor generalization ability of cross-category anomaly detection in traditional anomaly detection methods and inability to effectively convert semantic information into anomaly information, thereby improving the accuracy and robustness of image anomaly detection and the ability to detect cross-category anomalies. Figure 1 As shown, the few-sample anomaly detection method based on meta-learning and multimodal large model in this embodiment includes the following steps: S1: Divide the auxiliary dataset D={X,Y} containing N categories of detection target objects into N tasks according to the categories of the detection target objects. ~ The set of tasks , where X is the image sample and Y is the label. Divide into training set and test set , where the training set and test set Each query sample Each of them has normal prompt text, abnormal prompt text and reference samples for reference ; S2, constructing a semantic-anomaly network (MetaSAN) based on a multimodal large model; S3, using task collections The training set for each task in and test set Training the semantic-anomaly network to obtain a trained semantic-anomaly network; S4, task collection Test sets for each task in Sample query in And its corresponding normal prompt text, abnormal prompt text and reference samples Input the trained semantic-anomaly network to obtain the anomaly score, and determine the query sample to be tested based on whether the anomaly score exceeds the preset threshold. Is it abnormal?
[0024] Step S1 is used for task and data construction of anomaly detection meta-learning. Since it is few-sample learning, it is different from traditional supervised methods. It uses training sets for training and test sets for evaluation. Before training, we first constructed auxiliary training data through the MVTec dataset, and then selected a small number of normal samples on other datasets for reference prompts. All our auxiliary set data are used for training tasks. In order to ensure that the model considers the optimization direction of different situations during the training process of anomaly detection, the auxiliary dataset D={X, Y} is identified, where X is the image sample and Y is the label. Anomaly detection meta-learning will be divided into N tasks according to the N categories of objects in the auxiliary dataset, denoted as . Each task Divide into training set and test set ,in For internal updates, For external updates. Our goal is to capture abnormal information, not semantic information. For the convenience of representation, both the training set and the test set include query samples. (used to query whether the sample is abnormal) and k normal reference samples (called k-shot, used as a reference for query samples), as well as normal and abnormal text prompts. The present invention takes meta-learning in image anomaly detection as a premise that the model can be applied to multiple categories in few-sample learning. On the basis of meta-learning, anomaly detection meta-learning is proposed, and tasks and data dedicated to anomaly detection are constructed. , so that the model can be optimized in the direction of multi-category anomaly detection and enhance the model's robustness to different categories of anomalies. Each task contains image samples and prompt text samples, capturing abnormal information from both image-image and image-text aspects, and enhancing the model's ability to detect multi-category anomalies.
[0025] Step S2 is used to construct a semantic-anomaly network (MetaSAN) based on a large multimodal model; the purpose of the semantic-anomaly network is to capture anomaly information in the multimodal representation from a large visual language model. It mainly consists of an image-image anomaly discriminator and an image-text anomaly detector, which are used to learn anomalies between multiple feature levels (i.e., feature pyramids) and multiple modal dimensions (i.e., image-image, image-text). Figure 2 As shown, the semantic-anomaly network in this embodiment includes a visual encoder, a language encoder, an image-image anomaly discriminator, an image-text anomaly detector and a fusion module. The visual encoder is used to convert the query sample And its reference samples The visual encoding is performed respectively to obtain image embedding, the language encoder is used to perform language encoding on the normal prompt text and the abnormal prompt text to obtain text embedding, and the image-image abnormality discriminator is used to perform language encoding on the normal prompt text and the abnormal prompt text to obtain text embedding according to the query sample. And its reference samples The image embedding is used to identify anomalies and obtain anomaly scores The image-text anomaly detector is used to perform anomaly discrimination based on image embedding and text embedding to obtain anomaly scores. , the fusion module is used to score the anomaly and anomaly score Add and fuse to get the final anomaly score . Among them, the visual encoder is mainly responsible for processing the image data of the sample, extracting features from the image, and obtaining the image feature encodings of the query sample and the reference sample. The language encoder is mainly responsible for processing the prompt text data of the sample. Features are extracted from the prompt text to obtain normal and abnormal text encodings. The image-image anomaly discrimination module is responsible for discriminating the difference between the query sample image and the normal reference sample. The greater the difference between the two, the greater the abnormal value of the query sample, and the smaller the difference between the two, the smaller the abnormal value of the query sample. The image-text anomaly detection module is responsible for detecting the abnormal value of the query sample from the image-text level, and detecting the degree of abnormality of the query sample by obtaining the attribution relationship between the query sample image and the normal and abnormal text.
[0026] In order to capture more abundant abnormal information, this embodiment is inspired by the feature pyramid and proposes an image-image abnormality discriminator, which extracts multi-level features to capture more abundant abnormal information, making it adaptable to abnormality detection of different categories. Specifically, in this embodiment, the query sample And its reference samples The image embedding is obtained by performing visual encoding respectively, including: the visual encoder is the query sample Extract query samples separately Low-level features ,intermediate and advanced features , get the query sample Feature pyramid ; Visual encoder is the reference sample Extract reference samples separately Low-level features ,intermediate and advanced features , get the reference sample Feature pyramid . These feature pyramids enrich the representation of query samples and reference samples, effectively increasing the exposure of abnormal information. Once these richer representations are obtained, we input them into the image-image anomaly discriminator. It should be noted that the above-mentioned visual encoder does not depend on the specific type of visual encoder. For example, as an optional implementation, the visual encoder in this embodiment specifically adopts the visual encoder in the CLIP model. In addition, other types of visual encoders may also be used as needed.
[0027] After obtaining the outliers of the query sample from the image-image information, we consider how to obtain the outliers from the image-text information. In the prior art, methods such as WinCLIP and InCTRL realize anomaly detection by directly calculating the similarity between class tokens and text embeddings. For example, their text prompts are "a perfect {category} photo", "a damaged {category} photo", where {category} can refer to different categories of detected objects, such as chewing gum, capsules, bottles, etc. However, both class tokens and text contain redundant category information. In order to avoid the excessive influence of category information on anomaly detection, we discard the category information in the image-text anomaly detector: in this embodiment, the normal prompt text and the abnormal prompt text are the query samples. The name of the detected target object replaces the query sample The prompt text embedded in the detection target object category. For example, as an optional implementation, the normal prompt text in this embodiment is: "a perfect {object} photo... "a perfect {object} photo"; the abnormal prompt text = "a damaged {object} photo,..." a flawed {object} photo.
[0028] In this embodiment, language encoding the normal prompt text and the abnormal prompt text to obtain text embedding includes: language encoding the normal prompt text through a language encoder to obtain the text embedding of the normal prompt text , language encoding the exception prompt text to obtain the text embedding of the exception prompt text The above-mentioned language encoder does not depend on the specific language encoder type. For example, as an optional implementation, the language encoder in this embodiment specifically adopts the language encoder in the CLIP model. In addition, other types of language encoders can also be used as needed. The text encoder of CLIP is used to obtain text embedding. and Represent normal and abnormal prompts, respectively. Similarly, to ensure that the visual embedding is not affected by the category, no class label is used in this embodiment. Unlike image-image anomaly detection, which requires multi-level features to expose anomalies, image-text anomaly detection only uses high-level features combined with prompt text, because high-level features contain more detailed semantics and can be aligned with text information.
[0029] The image-image anomaly discriminator calculates the difference between the query sample and the reference sample based on the three-level features of the query sample and the reference sample to measure the difference between the query samples. The low-level features are used as an example to illustrate how to convert features into anomaly information. The underlying features of the query sample and the corresponding k-shot reference sample are recorded as and ,in and Respectively represent the number and dimension of features. For the features of the query sample Each feature in , we calculate its We then calculate the cosine similarity between c and the query sample feature e to obtain the anomaly score of the query sample. Finally, we take the mean of the maximum anomaly scores of the k features as the anomaly score of the underlying feature Similarly, we obtained the anomaly scores of mid-level features and high-level features respectively and Therefore, the anomaly scores of these three levels of features are added together to obtain the final anomaly score Specifically, in this embodiment, according to the query sample And its reference samples The image embedding is used to identify anomalies and obtain anomaly scores Includes: For query samples Low-level features and reference samples Low-level features ,in Used to represent dimensions, and Respectively represent the number and dimension of features, For the reference sample quantity, the query sample Low-level features Each feature in Calculate its difference with the reference sample Low-level features similarity to identify k The most similar feature c, and then find the feature c and feature The cosine similarity between them is used to obtain the features The anomaly score is finally taken k The average of the largest anomaly scores is used as the anomaly score of the low-level feature. ; For query samples Intermediate features and reference samples Intermediate features ,in Used to represent dimensions, and Respectively represent the number and dimension of features, For the reference sample quantity, the query sample Intermediate features Each feature in Calculate the difference between it and the reference sample Intermediate features similarity to identifyk The most similar feature c, and then find the feature c and feature The cosine similarity between them is used to obtain the features The anomaly score is finally taken k The average of the largest anomaly scores is used as the anomaly score of the intermediate feature. ; For query samples Advanced features of and reference samples Advanced features of ,in Used to represent dimensions, and Respectively represent the number and dimension of features, For the reference sample quantity, the query sample Advanced features of Each feature in Calculate the difference between it and the reference sample Advanced features of similarity to identify k The most similar feature c, and then find the feature c and feature The cosine similarity between them is used to obtain the features The anomaly score is finally taken k The average of the largest anomaly scores is used as the anomaly score of the high-level feature. ; Finally, the abnormality score of the low-level features is , anomaly score of mid-level features and anomaly scores for advanced features Add up to get the anomaly score .
[0030] Figure 2 This is a data flow diagram of the semantic-anomaly network MetaSAN in an embodiment of the present invention. Figure 3 First, the present invention constructs a sample consisting of a query sample image and a reference sample image, as well as a normal prompt text and an abnormal prompt text. Then, the sample is passed through the visual encoder and language encoder of the multimodal large model CLIP to extract the features of the query sample image, the reference sample image, the normal prompt text and the abnormal prompt text, and then these features are input into the image-image anomaly discriminator and the image-text anomaly detector to continue to calculate the anomaly score of the query sample. Next, the image-image anomaly discriminator performs feature-by-feature anomaly discrimination on the query sample and the reference sample features to obtain image-image anomaly scores at all levels, and the image-text anomaly detector models the image and abnormal text through an image feature differentiator and a multi-layer perceptron. It is expected that the higher the similarity between the abnormal sample and the abnormal text, the higher the image-text anomaly score. Finally, we fuse the various anomaly scores to obtain the final anomaly score. As Figure 3 As shown, the image-text anomaly detector in this embodiment includes a feature differentiator and a multi-layer perceptron, and the anomaly score is obtained by performing anomaly discrimination based on image embedding and text embedding. Including: Through the feature differentiator, first calculate the query sample Advanced features of and reference samples Advanced features of The difference between them is then mapped to the same latitude as the text embedding through a multi-layer perceptron to obtain the differential feature D, and then the normalized exponential function softmax is used to obtain the text embedding of the differential feature D belonging to the normal prompt text according to the following formula: and text embedding of exception prompt text Probability : , In the above formula, is an exponential function, is the cosine similarity function; the differential feature D is attributed to the text embedding of the normal prompt text and text embedding of exception prompt text Probability Obtaining anomaly scores through multi-layer perceptron .
[0031] Both the training set and the test set contain query samples and reference samples , as well as normal and abnormal prompt text. By inputting these data into the semantics-to-anomaly model, we obtain query samples The final anomaly score of , based on this score, we can get a set through the binary classification loss function The loss of this set can be the training set and test set Next, we introduce a meta-learning process for anomaly detection, which consists of three main steps: First, we train the training set of each task Input into the model. We update the model once internally based on the loss calculation gradient to obtain Internal update means that the model parameters are not actually updated. In this case, a reference weight for the category is obtained for the test set Secondly, we will test the set of each task Input into the model and use the previously obtained Compute the loss for each task. Test set The losses are accumulated. Finally, the cumulative loss of the test set of all tasks is used Perform an external update to actually update the model parameters , the external update is to update the actual parameters of the model in the optimization direction considering each category. This ensures that the gradient of each model optimization step simultaneously considers the anomaly detection update between N categories. Figure 4 As shown, in step S3 of this embodiment, the task set is used The training set for each task in and test set Training the semantic-anomaly network includes training for S rounds, and the training in each round includes: S3.1, initialize the category variable i to 0; S3.2, the tasks of the i-th category The training set Input into the semantic-anomaly network and calculate the loss according to the loss function; calculate the gradient of the loss and update the model parameters of the semantic-anomaly network internally , the model parameters of the semantic-anomaly network are updated once It refers to temporarily updating the model parameters of the semantic-anomaly network; S3.3, the tasks of the i-th category The test set Input to use model parameters The semantic-anomaly network of ; S3.4, add 1 to the category variable i, and determine whether the category variable i is less than N. If not, jump to step S3.2 to continue the training in this round; otherwise, calculate all the cumulative losses. The gradient of the semantic-anomaly network is updated externally once. , the external update once semantic-anomaly network model parameters It means actually updating the model parameters of the semantic-anomaly network once and saving them so that the model parameters saved in this round can be used in the next round of training of the semantic-anomaly network. .
[0032] Step S3 uses the task set The training set for each task in and test set Training the semantic-anomaly network includes a training phase and a testing phase. Figure 4As shown, in the training phase, firstly, by constructing the tasks and data for anomaly detection meta-learning, which are the training set and the test set respectively, then, the training set is sampled by the small batch random sampling method, sampling N tasks at a time, and each task contains batch_size query samples, and k-shot normal reference samples, as well as normal prompt text and abnormal prompt text. The sampled data is then input into the semantic-anomaly network to obtain the total loss of the training set, and the model parameters are obtained by internal update. The model parameters are then used as reference parameters to calculate the total loss of the test set, and then the total loss of N tasks is calculated, and then an external update is performed to actually update the model parameters, so that the model will be updated according to the optimization direction of different types of anomaly detection during optimization. When the training process reaches the specified number of training rounds, the present invention will stop training and take the best model on the verification set as the initial intrusion detection model. As shown Figure 5 As shown, in the test phase, this embodiment polls the test set, and each sample still contains four parts: query sample, reference sample, normal prompt text and abnormal prompt text. These four data are simultaneously input into the semantic-anomaly network to obtain the anomaly score. After polling all samples, the total accuracy is obtained.
[0033] In order to verify the few-sample anomaly detection method based on meta-learning and multimodal large models in this embodiment, the two most common indicators in anomaly detection, the area under the ROC curve (AUROC) and the area under the PR curve (AUPRC), are used as comparison indicators in this embodiment, and then compared with the existing WinCLIP method and InCTRL method on the Visa and ELPV data sets. The few-sample method is set to 2-shot. The comparative experimental results are shown in Table 1.
[0034] Table 1: Comparative experimental results of the method of this embodiment and the existing WinCLIP method and InCTRL method
[0035] As shown in Table 1, the semantic-anomaly network (MetaSAN) in this embodiment achieves good performance in both AUROC and AUPRC indicators on the Visa and ELPV datasets, surpassing the existing WinCLIP method and InCTRL method.
[0036] In addition, an ablation experiment is set up in this embodiment to verify the effectiveness of each module of MetaSAN. After removing the image-image anomaly discriminator (discriminator), image-text anomaly detector (detector) and anomaly meta-learning (meta-learning) respectively, the results are shown in Table 2.
[0037] Table 2: Ablation experiment results
[0038] As shown in Table 2, after removing the image-image anomaly discriminator (discriminator), image-text anomaly detector (detector) and anomaly meta-learning (meta-learning), the semantic-anomaly network (MetaSAN) in this embodiment has decreased in both AUROC and AUPRC indicators on the Visa and ELPV data sets, which shows that these three modules have all taken effect, and removing any one model will affect the performance of the semantic-anomaly network (MetaSAN).
[0039] In summary, the few-sample anomaly detection method based on meta-learning and multimodal large models in this embodiment includes dividing the auxiliary data set into a task set consisting of N tasks according to the category of the detection target object, dividing each task into a training set and a test set, constructing a semantic-anomaly network based on the multimodal large model, the semantic-anomaly network includes a visual encoder, a language encoder, an image-image anomaly discriminator, an image-text anomaly detector and a fusion module, using the training sets and test sets of each task in the task set to train the semantic-anomaly network and use the semantic-anomaly network to predict the anomaly score of the query sample to determine whether the query sample is abnormal, and the image-image anomaly discriminator is used in this embodiment to analyze the difference between the query sample and the reference sample from multi-layer features to improve the accuracy and robustness of image anomaly detection; the image-text anomaly detector is used in this embodiment to discard the concept of category from the original prompt text, so that the model can extract abnormal information that is not related to the category, thereby improving the model's ability to detect cross-category anomalies.
[0040] In addition, the present embodiment also provides a few-sample anomaly detection system based on meta-learning and multimodal large models, comprising a microprocessor and a memory connected to each other, wherein the microprocessor is programmed or configured to execute the few-sample anomaly detection method based on meta-learning and multimodal large models. In addition, the present embodiment also provides a computer-readable storage medium, in which a computer program or instruction is stored, and the computer program or instruction is programmed or configured to execute the few-sample anomaly detection method based on meta-learning and multimodal large models through a processor. In addition, the present embodiment also provides a computer program product, comprising a computer program or instruction, which is programmed or configured to execute the few-sample anomaly detection method based on meta-learning and multimodal large models through a processor.
[0041] Those skilled in the art should understand that the technical solutions provided by the embodiments of the present application may be in the form of methods, systems, or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes. The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the process Figure 1 A process or multiple processes and / or boxes Figure 1 These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing device to work in a specific way, so that the instructions stored in the computer-readable memory produce a product including an instruction device, which implements the functions specified in the process. Figure 1 A process or multiple processes and / or boxes Figure 1 These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide for implementing the process in the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0042] The above is only a preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions under the concept of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technicians in this technical field, some improvements and modifications without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.
Claims
1. A few-sample anomaly detection method based on meta-learning and multimodal large model, characterized in that: The steps include: S1: Divide the auxiliary dataset D={X,Y} containing N categories of detection target objects into N tasks according to the categories of the detection target objects. ~ The set of tasks , where X is the image sample and Y is the label. Divide into training set and test set , where the training set and test set Each query sample Each of them has normal prompt text, abnormal prompt text and reference samples for reference ; S2, constructing a semantic-anomaly network based on a multimodal large model, wherein the semantic-anomaly network includes a visual encoder, a language encoder, an image-image anomaly discriminator, an image-text anomaly detector and a fusion module. The visual encoder is used to convert the query sample And its reference samples The visual encoding is performed respectively to obtain image embedding, the language encoder is used to perform language encoding on the normal prompt text and the abnormal prompt text to obtain text embedding, and the image-image abnormality discriminator is used to perform language encoding on the normal prompt text and the abnormal prompt text to obtain text embedding according to the query sample. And its reference samples The image embedding is used to identify anomalies and obtain anomaly scores The image-text anomaly detector is used to perform anomaly discrimination based on image embedding and text embedding to obtain anomaly scores. , the fusion module is used to score the anomaly and anomaly score Add and fuse to get the final anomaly score ; S3, using task collections The training set for each task in and test set Training the semantic-anomaly network to obtain a trained semantic-anomaly network; S4, task collection Test sets for each task in Sample query in And its corresponding normal prompt text, abnormal prompt text and reference samples Input the trained semantic-anomaly network to obtain the anomaly score, and determine the query sample to be tested based on whether the anomaly score exceeds the preset threshold. Is it abnormal? 2. The method for detecting anomalies with a small number of samples based on meta-learning and multimodal large models according to claim 1, characterized in that: The normal prompt text and abnormal prompt text are the query samples The name of the detected target object replaces the query sample The detected target object category is embedded in the prompt text.
3. The method for detecting anomalies with a small number of samples based on meta-learning and multimodal large models according to claim 1, characterized in that: The query sample And its reference samples The image embedding is obtained by performing visual encoding respectively, including: the visual encoder is the query sample Extract query samples separately Low-level features ,intermediate and advanced features , get the query sample Feature pyramid ; Visual encoder is the reference sample Extract reference samples separately Low-level features ,intermediate and advanced features , get the reference sample Feature pyramid .
4. The method for detecting anomalies with a small number of samples based on meta-learning and multimodal large models according to claim 3, characterized in that: The language encoding of the normal prompt text and the abnormal prompt text to obtain text embedding includes: language encoding the normal prompt text by a language encoder to obtain the text embedding of the normal prompt text , language encoding the exception prompt text to obtain the text embedding of the exception prompt text .
5. The method for detecting anomalies with a small number of samples based on meta-learning and multimodal large models according to claim 4, characterized in that: According to the query sample And its reference samples The image embedding is used to identify anomalies and obtain anomaly scores Includes: For query samples Low-level features and reference samples Low-level features ,in Used to represent dimensions, and Respectively represent the number and dimension of features, For the reference sample quantity, the query sample Low-level features Each feature in Calculate the difference between it and the reference sample Low-level features similarity to identify k The most similar feature c, and then find the feature c and feature The cosine similarity between them is used to obtain the features The anomaly score is finally taken k The average of the largest anomaly scores is used as the anomaly score of the low-level feature. ; For query samples Intermediate features and reference samples Intermediate features ,in Used to represent dimensions, and Respectively represent the number and dimension of features, For the reference sample quantity, the query sample Intermediate features Each feature in Calculate the difference between it and the reference sample Intermediate features similarity to identify k The most similar feature c, and then find the feature c and feature The cosine similarity between them is used to obtain the features The anomaly score is finally taken k The average of the largest anomaly scores is used as the anomaly score of the intermediate feature. ; For query samples Advanced features of and reference samples Advanced features of ,in Used to represent dimensions, and Respectively represent the number and dimension of features, For the reference sample quantity, the query sample Advanced features of Each feature in Calculate the difference between it and the reference sample Advanced features of similarity to identify k The most similar feature c, and then find the feature c and feature The cosine similarity between them is used to obtain the features The anomaly score is finally taken k The average of the largest anomaly scores is used as the anomaly score of the high-level feature. ; Finally, the abnormality score of the low-level features is , anomaly score of mid-level features and anomaly scores for advanced features Add up to get the anomaly score .
6. The method for detecting anomalies with a small number of samples based on meta-learning and multimodal large models according to claim 5, characterized in that: The image-text anomaly detector includes a feature differentiator and a multi-layer perceptron. The anomaly discrimination is performed based on image embedding and text embedding to obtain an anomaly score. Including: Through the feature differentiator, first calculate the query sample Advanced features of and reference samples Advanced features of The difference between them is then mapped to the same latitude as the text embedding through a multi-layer perceptron to obtain the differential feature D, and then the normalized exponential function softmax is used to obtain the text embedding of the differential feature D belonging to the normal prompt text according to the following formula: and text embedding of exception prompt text Probability : , In the above formula, is an exponential function, is the cosine similarity function; the differential feature D is attributed to the text embedding of the normal prompt text and text embedding of exception prompt text Probability Obtaining anomaly scores through multi-layer perceptron .
7. The method for detecting anomalies with a small number of samples based on meta-learning and multimodal large models according to claim 1, characterized in that: Step S3 uses the task set The training set for each task in and test set Training the semantic-anomaly network includes training for S rounds, and the training in each round includes: S3.1, initialize the category variable i to 0; S3.2, the tasks of the i-th category The training set Input into the semantic-anomaly network and calculate the loss according to the loss function; calculate the gradient of the loss and update the model parameters of the semantic-anomaly network internally , the model parameters of the semantic-anomaly network are updated once internally It refers to temporarily updating the model parameters of the semantic-anomaly network; S3.3, the tasks of the i-th category The test set Input to use model parameters The semantic-anomaly network of ; S3.4, add 1 to the category variable i, and determine whether the category variable i is less than N. If not, jump to step S3.2 to continue the training in this round; otherwise, calculate all the cumulative losses. The gradient of the semantic-anomaly network is updated externally once. , the external update once semantic-anomaly network model parameters It means actually updating the model parameters of the semantic-anomaly network once and saving them so that the model parameters saved in this round can be used in the next round of training of the semantic-anomaly network. .
8. A few-sample anomaly detection system based on meta-learning and multimodal large models, comprising a microprocessor and a memory connected to each other, characterized in that: The microprocessor is programmed or configured to execute the few-sample anomaly detection method based on meta-learning and multimodal large model as described in any one of claims 1 to 7.
9. A computer-readable storage medium having a computer program or instruction stored therein, characterized in that: The computer program or instruction is programmed or configured to execute the few-sample anomaly detection method based on meta-learning and multimodal large model described in any one of claims 1 to 7 through a processor.
10. A computer program product comprising a computer program or instructions, characterized in that The computer program or instruction is programmed or configured to execute the few-sample anomaly detection method based on meta-learning and multimodal large model described in any one of claims 1 to 7 through a processor.
Citation Information
Patent Citations
Zero sample anomaly detection method based on multi-mode learnable prompt
CN118865000A
Method of segmenting abnormal robust for complex autonomous driving scenes and system thereof
US20240071096A1
AU2020103905A4
Cited By
A meta-learning anomaly detection method and system under a data full life cycle
CN120579100B
Experimental record text anomaly classification method and device based on image feature processing
CN121147952A
Experimental record text anomaly classification method and device based on image feature processing
CN121147952B
Abnormal sample generation system and anomaly detection device for quality control before manufacturing
CN121167312A
Abnormal sample generation system and abnormal detection device for pre-manufacturing quality control
CN121167312B