Zero sample image anomaly detection method based on prompt learning tuning
By adopting a cue learning tuning method in the zero-sample image anomaly detection model, combining the adaptive cue template and V-V attention mechanism, the existing models have solved the problems of low cue construction efficiency and poor abnormal positioning adaptability in the cue construction, achieving more efficient anomaly detection and more comprehensive anomaly evaluation.
Patent Information
- Application Number
- CN202510053241.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-13
AI Technical Summary
The existing zero-sample image anomaly detection model has problems such as low cue efficiency and poor adaptability of CLIP models in abnormal positioning tasks.
Using a method based on prompt learning tuning, a zero-sample anomaly detection model including a text encoder, an image encoder, a prompt learning module and a multi-scale feature extraction module is constructed. Category-independent prompt text is generated through the adaptive prompt template, and combined with the V-V attention mechanism and multi-scale feature extraction module, the abnormal area is accurately positioned.
It significantly improves the quality of text features, improves the precise positioning ability of abnormal areas, solves the problem of difficulty in detecting irregular abnormalities, and uses large visual language models to describe abnormalities to form a more comprehensive abnormality evaluation.
Smart Images

Figure CN119991588A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image anomaly detection, and in particular to a zero-sample image anomaly detection method based on prompt learning and tuning. Background Art
[0002] Image anomaly detection is a key task in computer vision and is widely used in industrial visual inspection, intelligent manufacturing and other fields. It includes two subtasks: anomaly classification and anomaly localization. The former is used to determine whether the target has an anomaly, and the latter is used to identify the pixel-level anomaly location. Compared with conventional target detection tasks, image anomaly detection faces many challenges, such as scarcity of abnormal samples, low visibility of anomalies, irregular shapes and uncertain types. In addition, with the widespread application of advanced manufacturing technology in the industrial field, abnormal samples are becoming increasingly scarce, but abnormal situations are becoming more complex and diverse. Based on these task characteristics, image anomaly detection places higher requirements on the multi-class recognition ability and zero-sample recognition ability of the anomaly detection model.
[0003] Pre-trained visual language models such as CLIP have shown strong zero-shot recognition capabilities in image anomaly detection tasks. They are combined with cue engineering to learn knowledge from cues and transfer it to anomaly detection tasks. However, cue engineering relies on manual design, and its generalization ability is limited by the design of cue templates, which is inconsistent with the automation requirements of real-world scenarios. In addition, CLIP is a classification model. Knowledge related to visual anomalies is not injected into the model, but it detects unknown anomalies through its strong generalization ability. However, anomaly detection includes a classification task and a segmentation and localization task. Without fine-tuning, CLIP has poor adaptability in cue-guided image localization tasks. Summary of the invention
[0004] In order to overcome the technical defects of low efficiency in constructing prompts in existing zero-shot image anomaly detection models and poor adaptability of CLIP models in anomaly localization tasks, the present invention provides a zero-shot image anomaly detection method based on prompt learning and tuning.
[0005] The present invention is implemented by adopting the following technical solution: a zero-sample image anomaly detection method based on prompt learning and tuning, comprising:
[0006] S1. Obtain the image data to be detected and perform preprocessing;
[0007] S2. Build a zero-shot anomaly detection model, including a text encoder, an image encoder, a hint learning module, a multi-scale feature extraction module, etc.
[0008] S3. Fine-tune the zero-shot anomaly detection model using a typical dataset in the field of images to be detected;
[0009] S4. Use the fine-tuned zero-shot anomaly detection model to detect the image to be detected, and use a basic large model to describe the anomaly information to obtain the anomaly classification, anomaly location and anomaly description results of the image.
[0010] Optionally, in step S1, the image to be detected is collected using a high-quality imaging device, and may be an image of a certain field collected manually in real time, or may be an existing high-resolution data set in a certain field, acquired by local reading or network transmission.
[0011] Optionally, in step S2, the main structure of the zero-sample anomaly detection model adopts the CLIP model, which includes two branches: an image encoder and a text encoder. In addition to the main structure, it also includes a prompt learning module and a multi-scale feature extraction module. The prompt learning module generates a pair of positive and negative sample prompts for each image, and the multi-scale feature extraction module processes local features of the image to accurately locate anomalies of different sizes and shapes.
[0012] Optionally, in step S2, the visual encoding subnetwork of the CLIP model adopts dual-path calculation to extract global visual features for abnormal classification in the multi-head attention path. The original multi-head attention is expressed as follows:
[0013] Attn(Q,K,V)=softmax(Q·K T ·scale)·V
[0014] Where Q, K, and V represent the Query, Key, and Value in the multi-head attention mechanism. After the inner product of Q and K and the scaling calculation, the attention weight is obtained through the softmax function, and the final attention output is obtained after the inner product with V.
[0015] Optionally, in step S2, the VV attention path extracts and aggregates the visual features of a specific layer of the image encoder for abnormality location and generation of an abnormality map. The specific expression is as follows:
[0016] Attn(V l ,V l ,V l )=softmax(V·V T ·scale)·V
[0017] Z l =Proj·(Attn(V l ,V l ,V l ))+Z l-1
[0018] In the formula, Z l-1represents the local fine-grained perception output at the (l-1)th layer, Z l Denotes the local fine-grained perception output of layer 1, and Proj denotes the output projection. Compared with the original QKV attention calculation, VV attention replaces Q and K with V and removes the MLP at the top layer.
[0019] Optionally, in step S2, the prompt learning module automatically learns a pair of class-independent prompt texts through an adaptive prompt template, thereby effectively encoding the prompt context of the positive and negative samples. The adaptive prompt learning template is expressed as follows:
[0020]
[0021] Where Prompt+ and Prompt- represent text prompts for modeling normal and abnormal conditions, respectively, each including N learnable prompt contexts; domain prompt words include relevant vocabulary in the field to be detected; object and damage object represent image states.
[0022] Optionally, in step S2, the multi-scale feature extraction module uses convolution kernels of different sizes and shapes to process local visual features, so that the model can adapt to anomalies of different sizes and shapes, enrich local visual features, and solve the problem that existing models have difficulty in detecting irregular anomalies.
[0023] Optionally, in step S3, the dataset used to fine-tune the zero-sample anomaly detection model comes from the same field as the image to be detected. The dataset can be an image stored in a local server or cloud. For any type of image in the dataset, it should include a high-resolution original image and its corresponding pixel-level annotated image, and be obtained by local reading or network transmission.
[0024] Optionally, in step S3, when fine-tuning the zero-sample anomaly detection model, the cross entropy loss is used to optimize the anomaly classification task, and the Focal loss and Dice loss are used to optimize the anomaly localization task. During the entire fine-tuning process, all CLIP model parameters are frozen, and only the parameters of the prompt learning module and the multi-scale feature extraction module are optimized. The calculation expressions of Focalloss and Dice loss are:
[0025] Focal t )=-α t *(1-p t ) γ *log(P t )
[0026]
[0027] In the formula, pt is the probability distribution predicted by the model, α t is a balancing factor used to adjust the weights of positive and negative samples, and γ is a regulating factor used to reduce the weights of easy-to-classify samples. i is the probability of the i-th pixel predicted by the model, g i is the true label of the ith pixel, and ∈ is a smoothing term to prevent the denominator from being zero.
[0028] Optionally, in step S4, when detecting the image to be detected, the fine-tuned zero-sample anomaly detection model is used to perform anomaly classification and anomaly location. In the anomaly classification task, the cosine similarity of the global visual features and text features of the image to be detected is calculated to obtain the anomaly classification result; in the anomaly location task, the cosine similarity of the processed local visual features and text features of the image to be detected is calculated and interpolated to obtain an anomaly map, thereby outputting the anomaly location result.
[0029] Optionally, in step S4, the ChatGLM-4V model is used to perform the anomaly description. During this process, any unrelated high-resolution image without anomalies needs to be provided for reference, and the model is instructed to output an anomaly description of the image to be detected through task-related prompts and output format-related prompts, thereby forming a more comprehensive anomaly assessment.
[0030] The technical solution provided by the present invention has the following advantages compared with the prior art:
[0031] The present invention provides a zero-shot image anomaly detection method based on prompt learning and tuning. By learning contextual prompts and domain prompt words that are not related to categories, the positive and negative sample prompt contexts are effectively encoded, and the quality of text features is significantly improved. The method also uses the VV attention mechanism and the multi-scale feature extraction module to accurately locate the abnormal area, solving the problem that the existing model has difficulty in detecting irregular anomalies. Finally, with the help of the prior knowledge of advanced large models (such as ChatGLM-4V), the detection image is described abnormally to form a comprehensive detection and evaluation of image anomalies, providing a richer judgment basis for industrial image anomaly detection, and filling the gap in the traditional method in the detailed description of abnormal information. Compared with other zero-shot image anomaly detection models, this method solves the key pain points in the anomaly detection task and significantly improves the model's zero-shot image anomaly detection capability. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0033] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0034] Figure 1 A flowchart of a zero-sample anomaly detection method based on prompt learning and tuning in an embodiment of the present invention is shown.
[0035] Figure 2 It shows the overall structure diagram of the zero-sample anomaly detection method based on prompt learning and tuning in an embodiment of the present invention.
[0036] Figure 3 A diagram showing the dual-path structure of the CLIP image encoder according to an embodiment of the present invention.
[0037] Figure 4 A prompt input template diagram of ChatGLM-4V in an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0038] In order to more clearly understand the above-mentioned objectives, features and advantages of the present invention, the scheme of the present invention will be further described below. It should be noted that the embodiments of the present invention and the features in the embodiments can be combined with each other without conflict.
[0039] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present invention, rather than all of the embodiments.
[0040] A specific embodiment of the present invention is described in detail below with reference to the accompanying drawings.
[0041] Reference Figure 1 ,This embodiment provides a zero-sample industrial image anomaly detection method based on prompt learning and tuning, including steps S1 to S4.
[0042] S1. Obtain image data to be detected and perform preprocessing.
[0043] Specifically, the image to be detected is a high-resolution image obtained by using a high-quality imaging device, including a foreground and a background with high distinction, and the image to be detected may or may not include an abnormality.
[0044] More specifically, the abnormal area is usually part of the object to be detected in the foreground, so the accurate segmentation of each area of the image to be detected is conducive to the rapid detection of subsequent abnormalities and reduces the computational complexity.
[0045] S101. Perform uniform preprocessing on the image data to be detected, such as resizing, center cropping, and standardization, to ensure the consistency of the input data and facilitate model evaluation.
[0046] S102. Perform denoising on the input data to filter out noise to make the image smooth.
[0047] In some embodiments, the image to be detected may be a real-time scanned image, or may be an image stored in a local server or cloud, and may be obtained by reading or transmitting over a network.
[0048] In some implementations, the input image denoising employs a Gaussian smoothing filter method to smooth the image.
[0049] In a specific embodiment of the present invention, the images are uniformly adjusted to a size of 518*518.
[0050] S2. Build a zero-shot anomaly detection model, including a text encoder, an image encoder, a hint learning module, a multi-scale feature extraction module, etc.
[0051] Reference Figure 2 , a zero-sample image anomaly detection model established in one embodiment of the method is shown in the figure. Its main structure adopts the CLIP-VIT / L-14 model, which includes two branches: an image encoder and a text encoder. The text encoder includes 12 layers of text encoding subnetworks, and each layer of the text encoding subnetwork adopts the Transformer model. The image encoder includes 24 layers of visual encoding subnetworks, and each layer of the visual encoding subnetwork adopts the VIT model. In addition to the main structure, it also includes modules such as a prompt learning module and a multi-scale feature extraction module.
[0052] More specifically, the prompt learning module automatically learns a pair of class-independent prompt texts through an adaptive prompt template, thereby effectively encoding the prompt context of positive and negative samples and inputting them into the text encoder to generate text features. The prompt template in the prompt learning module is as follows:
[0053]
[0054] Where Prompt+ and Prompt- represent the text prompts for modeling normal and abnormal conditions, respectively. Each text prompt includes N learnable prompt contexts. The domain prompt word (domain) includes some words related to the domain description. Object and damage object represent the image status.
[0055] More specifically, see Figure 3 ,The visual encoding subnetwork adopts dual-path calculation to extract global visual features for anomaly classification in the multi-head attention path. The original multi-head attention is expressed as follows:
[0056] Attn(Q,K,V)=softmax(Q·K T ·scale)·V
[0057] Where Q, K, and V represent the Query, Key, and Value in the multi-head attention mechanism. After the inner product of Q and K and the scaling calculation, the attention weight is obtained through the softmax function, and the final attention output is obtained after the inner product with V.
[0058] More specifically, the VV attention path extracts and aggregates local visual features of a specific layer in the image encoder for anomaly location and generates anomaly maps. The specific expression is as follows:
[0059] Attn(V l ,V l ,V l )=softmax(V·V T ·scale)·V
[0060] Z l =Proj·(Attn(V l ,V l ,V l ))+Z l-1
[0061] In the formula, Z l-1 represents the local fine-grained perception output at the (l-1)th layer, Z l Denotes the local fine-grained perception output of layer 1, and Proj denotes the output projection. Compared with the original QKV attention calculation, VV attention replaces Q and K with V and removes the MLP at the top layer.
[0062] In a specific embodiment of the present invention, a multi-head attention mechanism is used to obtain the global visual features of the image to be detected, and a VV attention mechanism is used to extract the visual features of the 6th, 12th, 18th, and 24th layers and aggregate them to obtain local visual features.
[0063] More specifically, a multi-scale feature extraction module is used to process local visual features with convolution kernels of different sizes and shapes to comprehensively locate abnormal regions of different shapes and sizes.
[0064] In a specific embodiment of the present invention, the convolution kernels of different sizes used include 6 sizes: 3*3, 1*1, 7*7, 5*5, 3*1, and 1*3.
[0065] S3. Fine-tune the zero-shot anomaly detection model using a typical dataset in the field of images to be detected.
[0066] S301. Perform data enhancement on image data in typical data sets, such as random cropping, color jittering, etc., to increase data diversity and improve the generalization ability of the model.
[0067] S302: De-noising the image in the dataset to filter out the noise and make the image smooth.
[0068] Specifically, a typical dataset used for fine-tuning a model is an image dataset stored in a local server or cloud. The dataset includes high-resolution original images and their corresponding pixel-level annotated images.
[0069] In some embodiments, image data denoising uses a Gaussian smoothing filter method to smooth the image.
[0070] More specifically, when fine-tuning the model, the parameters of the CLIP backbone model are frozen, and its pre-trained parameters are used. The learnable parameters only include the parameters of the prompt learning module and the multi-scale feature extraction module. During the training process, the cross entropy loss is used to optimize the anomaly classification task, and the Focal loss and Dice loss are used to optimize the anomaly localization task. Focal loss assigns lower weights to samples that are easy to classify and higher weights to samples that are difficult to classify; DiceLoss is an indicator that measures the similarity between the model prediction and the true label, which can be used in multi-classification problems. The calculation expressions of Focalloss and Dice loss are:
[0071] Focal t )=-α t *(1-p t ) γ *log(P t )
[0072]
[0073] In the formula, p t is the probability distribution predicted by the model, α t is a balancing factor used to adjust the weights of positive and negative samples, and γ is a regulating factor used to reduce the weights of easy-to-classify samples. i is the probability of the i-th pixel predicted by the model, g i is the true label of the ith pixel, and ∈ is a smoothing term to prevent the denominator from being zero.
[0074] In a specific embodiment of the present invention, the selected data sets are MvTec AD data set and VisA data set, which are stored in a local server, obtained by reading, and uniformly adjusted to a size of 518*518 for input model. In the fine-tuning process, the Adam optimizer is used, and the specific training parameters are as follows: a fixed learning rate of 0.001, a bias betas set to (0.5, 0.999), a training batch size of 8, and a number of iterations of 15 cycles. Correspondingly, a model fine-tuned by MvTec AD and a model fine-tuned by the VisA data set are obtained.
[0075] S4. Use the fine-tuned zero-shot anomaly detection model to detect the image to be detected, and use a basic large model to describe the anomaly information to obtain the comprehensive detection and evaluation results of the image anomaly.
[0076] Specifically, the preprocessed image in S1 is input into the zero-shot image anomaly detection model obtained by fine-tuning in S3. In the image encoder, the global visual features extracted by multi-head attention and the local visual features obtained by aggregating the 6th, 12th, 18th, and 24th layer features extracted by VV attention are obtained; in the text encoder, the text features of the adaptive prompt embedding learned by the prompt learning module are obtained.
[0077] More specifically, the classification scores of global visual features and text features are calculated to obtain the anomaly classification results; the cosine similarity of local visual features and text features is calculated and interpolated to obtain a similarity map, and the anomaly map is obtained by similarity measurement to obtain the anomaly localization result.
[0078] More specifically, during the anomaly detection process, refer to Figure 2 , use the ChatGLM-4V model to describe the anomaly. During this process, any irrelevant high-resolution image without anomaly needs to be provided for reference, and the model is instructed to output the anomaly description of the image to be detected through task-related prompts and output format-related prompts, thereby forming a more comprehensive anomaly assessment.
[0079] More specifically, see Figure 4 ,ChatGLM-4V’s prompts include task-related prompt templates and output format-related templates. Task-related prompts define the anomaly detection task, while format-related prompts instruct the model to output anomaly detection in binary (0, 1) form and provide corresponding reasoning for the prediction in the format of Python strings.
[0080] In a specific embodiment of the present invention, when the zero-shot image anomaly detection model fine-tuned by the MvTec AD dataset is used to test the VisA dataset, the image-level AUROC index and the pixel-level AUROC index achieve performances of 82.5 and 95.5, respectively, in the anomaly classification and anomaly localization tasks; in the anomaly description, given a normal image of the MvTec AD dataset, ChatGLM-4V can well evaluate the anomalies in the VisA dataset. When the MvTec AD dataset is tested using the zero-shot image anomaly detection model fine-tuned by the VisA dataset, the image-level AUROC index and the pixel-level AUROC index achieve performances of 91.8 and 90.9, respectively, in the anomaly classification and anomaly localization tasks; in the anomaly description, given a normal image of the VisA dataset, ChatGLM-4V can well evaluate the anomalies in the MvTec AD dataset. The experimental results of this embodiment show that the method used in the present invention constructs more comprehensive contextual cues that are more adaptable to anomaly detection tasks and significantly improves the positioning performance of the model. In addition, combining a large visual language model with the method of the present invention enables the model to perform more comprehensive anomaly evaluation when performing anomaly detection.
[0081] The above is only a specific implementation of the present invention, which enables those skilled in the art to understand or implement the present invention. Although detailed descriptions are given with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the aforementioned embodiments, or replace some or all of the technical features therein by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments, and they should all be covered by the protection scope of the claims.
Claims
1. A zero-shot image anomaly detection method based on prompt learning and tuning, characterized in that: The steps include: S1. Obtain the image data to be detected and perform preprocessing; S2. Build a zero-shot anomaly detection model, including a text encoder, an image encoder, a hint learning module, and a multi-scale feature extraction module; S3. Fine-tune the zero-shot anomaly detection model using a typical dataset in the field of images to be detected; S4. Use the fine-tuned zero-shot anomaly detection model to detect the image to be detected, and use a basic large model to describe the anomaly information to obtain anomaly classification, anomaly location and anomaly description results.
2. The zero-shot image anomaly detection method based on prompt learning and tuning according to claim 1, characterized in that: The image to be detected in step S1 is collected by using an imaging device, which can be an image of a certain field collected manually in real time, or an existing data set in a certain field, which is obtained by local reading or network transmission; the preprocessing method includes resizing, center cropping, and standardization to ensure the consistency of the input data; and denoising is performed to filter out noise to make the image smooth.
3. The zero-shot image anomaly detection method based on prompt learning and tuning according to claim 1, characterized in that: In step S2, the main structure of the zero-sample anomaly detection model adopts the CLIP model, which includes two branches: an image encoder and a text encoder. In addition to the main structure, it also includes a prompt learning module and a multi-scale feature extraction module. The prompt learning module generates a pair of positive and negative sample prompts for each image, and the multi-scale feature extraction module processes the local visual features of the image to accurately locate anomalies of different sizes and shapes.
4. The zero-shot image anomaly detection method based on prompt learning and tuning according to claim 3, characterized in that: In step S2, the image encoder adopts dual-path calculation, including a multi-head attention path and a VV attention path. In the original multi-head attention path, the expression is as follows: Attn(Q,K,V)=softmax(Q·K T ·scale)·V In the formula, Q, K, and V represent the Query, Key, and Value in the multi-head attention mechanism. After the inner product of Q and K and the scaling calculation, the attention weight is obtained by the softmax function, and the final attention output is obtained by further inner product with V. The expression in the VV attention path is as follows: Attn(V l ,V l ,V l )=softmax(V·V T ·scale)·V From l =Proj·(Attn(V l ,V l ,V l ))+Z l-1 In the formula, Z l-1 represents the local fine-grained perception output at layer l-1, Z l Represents the local fine-grained perception output of the lth layer, Proj represents the output projection, and V is usually used to provide comprehensive information for each position or area in the image. Here, V is used to replace Q and K of the original attention calculation, which enhances the model's attention to local fine-grained features without destroying the original structure.
5. The zero-shot image anomaly detection method based on prompt learning and tuning according to claim 3, characterized in that: In step S2, the prompt learning module automatically learns a pair of class-independent prompt texts through an adaptive prompt template, thereby effectively encoding the prompt context of positive and negative samples; the adaptive prompt template is expressed as follows: Among them, Prompt+ and Prompt- represent the text prompts for modeling normal and abnormalities, respectively, each including N learnable prompt contexts; the domain prompt word domain includes some words representing the image domain; object and damage object represent the abnormal state of the image.
6. The zero-shot image anomaly detection method based on prompt learning and tuning according to claim 3, characterized in that: In step S2, the multi-scale feature extraction module uses convolution kernels of different sizes and shapes to process local visual features, so that the model can adapt to anomalies of different sizes and shapes, enrich local visual features, and solve the problem that existing models have difficulty in detecting irregular anomalies.
7. The zero-shot image anomaly detection method based on prompt learning and tuning according to claim 3, characterized in that: In step S3, the dataset used to fine-tune the zero-sample anomaly detection model is in the same field as the image to be detected; the dataset can be an image stored in a local server or cloud. For any type of image in the dataset, it includes the original image and its corresponding pixel-level annotated image, and is obtained by local reading or network transmission; and the dataset images are further enhanced to increase data diversity.
8. The zero-shot image anomaly detection method based on prompt learning and tuning according to claim 7, characterized in that: In step S3, when fine-tuning the zero-sample anomaly detection model, the cross entropy loss is used to optimize the anomaly classification task, and Focalloss and Diceloss are used to constrain the anomaly localization task; during the entire fine-tuning process, all CLIP model parameters are frozen, and only the parameters of the prompt learning module and the multi-scale feature extraction module are optimized.
9. The zero-shot image anomaly detection method based on prompt learning and tuning according to claim 1, characterized in that: In step S4, when detecting the image to be detected, the fine-tuned zero-sample anomaly detection model obtained in claim 8 is used to perform anomaly classification and anomaly location; in the anomaly classification task, the cosine similarity of the global visual features and text features of the image to be detected is calculated to obtain the anomaly classification result; in the anomaly location task, the cosine similarity of the processed local visual features and text features of the image to be detected is calculated and interpolated to obtain an anomaly map to obtain an anomaly location result.
10. The zero-shot image anomaly detection method based on prompt learning and tuning according to claim 9, characterized in that: In step S4, the large basic model ChatGLM-4V model is used to perform the abnormal description. During this process, any irrelevant non-abnormal image is provided for reference, and the model is instructed to output the abnormal description of the image to be detected through task-related prompts and output format-related prompts, thereby forming a more comprehensive abnormality assessment.
Citation Information
Cited By
Zero sample anomaly detection method and device and electronic equipment
CN120635483A
Abnormity detection method based on multi-modal and multi-level features, electronic equipment, storage medium and program product
CN121191097A