Zero-sample industrial anomaly classification and segmentation system and method based on multi-source expert scoring

Through the multi-source feature extraction and cascade fusion methods, the problem of domain differences in zero-sample industrial anomaly classification and segmentation is solved, and the detection accuracy and robustness are achieved, which is suitable for industrial image detection.

CN120510449APending Publication Date: 2025-08-19SICHUAN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510723499.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-31
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

When facing different domain differences, the existing zero-sample industrial anomaly classification and segmentation methods are difficult to effectively alleviate domain differences, resulting in insufficient accuracy of industrial image anomaly classification and segmentation, and uncertain performance when detecting new product categories.

Method used

The multi-source feature extractor is used to extract image features through multiple pre-trained models, combined with the expert scoring module and the cascade fusion module, and the multi-source, multi-stage and multi-scale feature fusion is used to generate the final abnormal scoring map.

Benefits of technology

It effectively alleviates the differences between different domains, improves the accuracy and robustness of abnormal detection, ensures the inference speed, and improves the overall performance of industrial image detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120510449A_ABST
    Figure CN120510449A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of image recognition, and discloses a zero-sample industrial anomaly classification and segmentation system and method based on multi-source expert scoring, and the system comprises a multi-source feature extractor which is used for carrying out the feature extraction of a to-be-detected image through more than two pre-trained feature extraction models, obtaining more than two image source features; the expert scoring module is used for obtaining an abnormal scoring vector in each image source feature according to an expert delegate; and the cascading fusion module is used for carrying out cascading fusion processing on the abnormal score vector of each image source feature to obtain a fused abnormal score vector, and generating a final abnormal score graph according to the fused abnormal score vector, including abnormal classification scoring and abnormal segmentation output. According to the method, the difference of different domains is effectively relieved, the accuracy of system anomaly recognition is improved, the reasoning speed is ensured, the overall performance in anomaly detection is improved, and the method is expected to be widely used in industrial image detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image recognition technology and relates to a zero-sample industrial anomaly classification and segmentation technology, and in particular to a zero-sample industrial anomaly classification and segmentation system and method based on multi-source expert scoring. Background Art

[0002] Industrial Anomaly Classification and Segmentation (AC / AS) aims to identify abnormal regions within products, which is crucial for ensuring product quality during industrial inspection. Deep learning technology has been widely used in industrial visual inspection to address the low efficiency, high cost, and high error rates inherent in traditional manual inspection. However, supervised detection methods face a significant challenge: obtaining a sufficient number of abnormal samples and pre-defining all possible anomaly types.

[0003] To address these issues, many robust unsupervised methods have been developed. These methods do not require prior knowledge of anomalies and usually rely on a completely normal dataset for training, hence they are called full-sample methods. However, in real-world applications, collecting a large and diverse set of normal images for each product type is extremely challenging. Therefore, some few-shot learning methods have been proposed to address the problem of limited sample numbers. These methods rely on only a small number of normal images for training while still achieving satisfactory accuracy. In addition, real-world applications often require the detection of new industrial product categories. In this case, full-shot methods require retraining the model, which is not only time-consuming but also uncertain in terms of performance on new product categories.

[0004] To this end, zero-shot methods have been developed. These methods can classify and segment anomalies in industrial product categories without additional training. Recent advances in zero-shot AC / AS include WinCLIP and APRIL-GAN, which cleverly leverage textual cues for anomaly measurement, introducing new perspectives and possibilities. Furthermore, MuSc significantly improves zero-shot AC / AS by leveraging normal and anomaly cues in unlabeled test images without relying on additional information such as textual cues.

[0005] The success of existing zero-shot AC / AS methods is largely due to the excellent performance of multimodal models in capturing a wide range of conceptual representations in images. The completeness of the image representations extracted by pre-trained models is a key factor affecting the performance of these methods, as most existing methods rely on models pre-trained on large datasets to extract image representations for AC / AS. However, there is a domain difference between industrial product images and the training sets of these pre-trained models. Different pre-trained models focus on different aspects when extracting representations of industrial product images (e.g. Figure 1 ), which indicates that a single pre-trained model cannot fully represent these images.

[0006] Therefore, how to alleviate the differences between different domains, improve the accuracy of industrial image anomaly classification and segmentation, and achieve more reliable AC / AS is a key technical problem that needs to be solved urgently in this field. Summary of the Invention

[0007] The purpose of the present invention is to address the problems existing in the above-mentioned prior art and provide a zero-shot industrial anomaly classification and segmentation system and method based on multi-source expert scoring, which can extract more comprehensive industrial image representations, alleviate the differences between different domains, and improve the overall performance in anomaly detection.

[0008] In order to achieve the above objectives, the present invention adopts the following technical solutions to achieve them.

[0009] The present invention provides a zero-shot industrial anomaly classification and segmentation system based on multi-source expert scoring, which includes: A multi-source feature extractor is used to extract features from the image to be detected using two or more pre-trained feature extraction models, respectively, to obtain two or more image source features; each image source feature includes several patch block feature vectors corresponding to the image to be detected; The expert scoring module obtains the abnormality scoring vector in each image source feature based on the expert committee; The cascade fusion module is used to perform cascade fusion processing on the anomaly score vectors of each image source feature to obtain a fused anomaly score vector, and generate the final anomaly score map based on the fused anomaly score vector, including anomaly classification score and anomaly segmentation output.

[0010] In one implementation, the feature extraction model is an image encoder pre-trained using a large model such as CLIP, DINOV2, or DINO. The layers of the feature extraction module are divided into several stages, and the outputs of the stages are aggregated at different scales to obtain feature vectors for each patch at different stages and scales. Specifically, the patch features output by each stage of the feature extraction module are aggregated at different scales using 1×1, 3×3, and 5×5 as the local neighborhood aggregation ranges, respectively. This yields patch features at corresponding scales, known as locally aware patch features. For details, see Towards Total Recall in Industrial Anomaly Detection, Karsten Roth, Latha Pemula, Joaquin Zepeda, et al, 2022 IEEE / CVFConference on Computer Vision and Pattern Recognition (CVPR), 978-1-6654-6946-3 / 22 / $31.00 ©2022 IEEE, DOI 10.1109 / CVPR52688.2022.01392. The patch features at different stages and scales for the same patch form the corresponding patch feature vector.

[0011] In one feasible embodiment, the method for constructing the expert committee is as follows: for a number of samples of the same category as the image to be detected, feature extraction is performed through any feature extraction model in the multi-source feature extractor to obtain a number of sample image source features; the extracted features of each sample at any stage and any scale are scored against each other to obtain the initial anomaly scores of all sample extracted features, and the samples are sorted in ascending order according to the initial anomaly scores, and the samples whose initial anomaly scores meet the set threshold range are selected as normal samples to form the expert committee.

[0012] The expert scoring module calculates the minimum distance between each patch feature within the same image source feature and each expert committee member using a minimum distance function. It then performs an interval averaging operation to obtain an anomaly score for each patch feature within each image source feature. The anomaly scores of all patch features constitute the anomaly score vector for the corresponding image source feature. The minimum distance function can be Euclidean distance, Mahalanobis distance, or cosine similarity distance; the interval averaging operation averages the minimum distances between the patch features and the expert committee members that meet the minimum distance interval requirement.

[0013] In one implementable manner, the cascade fusion module includes a multi-source fusion expert, a multi-stage fusion expert and a multi-scale fusion expert; the multi-source fusion expert fuses the anomaly scores of the same patch block features from different image source features through a first single-class support vector machine to obtain a first fused anomaly score of each patch block feature; the multi-stage fusion expert averages the first fused anomaly scores of the patch block features of the same patch block at different stages at the same scale to obtain a second fused anomaly score of the same patch block at different scales; the multi-scale fusion expert obtains a third fused anomaly score of the corresponding patch block through a second single-class support vector machine based on the second fused anomaly scores of the same patch block at different scales; the third fused anomaly scores of all patch blocks constitute a fused anomaly score vector.

[0014] The above-mentioned cascade fusion module also includes an anomaly output module, which is used to reshape the fused anomaly score vector and then upsample the reshaped vector to match the resolution of the original image I to obtain an anomaly segmentation output; at the same time, the maximum anomaly score is obtained from the fused anomaly score vector as the image-level anomaly classification score of the image to generate the final anomaly score map.

[0015] The present invention also provides a zero-shot industrial anomaly classification and segmentation method based on multi-source expert scoring, which is performed using any of the above-mentioned implementable systems according to the following steps: Step 1: Perform multi-source feature extraction on the image to be detected to obtain two or more image source features; each image source feature includes several patch block feature vectors corresponding to the image to be detected; Step 2: Obtain anomaly score vectors in each image source feature based on the expert committee; Step 3: Perform cascade fusion processing on the anomaly score vectors of each image source feature to obtain a fused anomaly score vector, and generate the final anomaly score map based on the fused anomaly score vector, including anomaly classification score and anomaly segmentation output.

[0016] In step 1, the feature extraction module performs several stages of feature extraction on the image to be detected, and the outputs of the several stages are respectively subjected to neighborhood aggregation of different scales to obtain the patch block feature vectors at different stages and scales of each patch block.

[0017] Described step 3 comprises the following sub-steps: Step 31, fusing the anomaly scores of the same patch block feature from different image source features through a first single-class support vector machine to obtain a first fused anomaly score of each patch block feature; For the mth patch extracted by the multi-source feature extractor, the anomaly score vector from K different image source features is expressed as , is the anomaly score at the lth stage and the rth scale aggregation level; the multi-source fusion expert D S The first fused anomaly score of the patch is obtained by fusing the anomaly score vectors from K non-image source features through the first single-class support vector machine. : ; At the same scale, the first fusion anomaly score of the patch features of the same patch at different stages constitutes the anomaly score vector of the r-th scale aggregation level, which is expressed as ; Step 32: averaging the first fused anomaly scores of the patch features at different stages of the same patch at the same scale to obtain a second fused anomaly score of the same patch at different scales; Through multi-stage fusion expert D L The average operation of the anomaly score vector at the r-th scale aggregation level The fusion is performed to obtain the second fusion anomaly score of the same patch at this scale: ; The anomaly score vectors of the patch at different aggregation levels are expressed as ; Step 33: The second fused anomaly scores at different scales of the same patch are used to obtain a third fused anomaly score of the corresponding patch through a second single-class support vector machine; the third fused anomaly scores of all patches constitute a fused anomaly score vector; Multi-scale fusion expert D R The anomaly scoring vectors from different scale aggregation levels are analyzed by the second one-class support vector machine. The fusion is performed to obtain the third fusion anomaly score of the corresponding patch block: ; Repeat steps 31 to 33 above to calculate the third fused anomaly scores of all patches of the image to be detected. The third fused anomaly scores of all patches constitute a fused anomaly score vector. In step 34, the fused anomaly score vector is reshaped, and then the reshaped vector is upsampled to match the resolution of the original image I, thereby obtaining an anomaly segmentation output; at the same time, the maximum anomaly score is obtained from the fused anomaly score vector as the image-level anomaly classification score of the image, and the final anomaly score map is generated.

[0018] Compared with the existing technology, the zero-sample industrial anomaly classification and segmentation system and method based on multi-source expert scoring provided by the present invention has the following beneficial effects: 1) This paper proposes a novel zero-shot industrial anomaly classification and segmentation system, which includes a multi-source feature extractor, an expert scoring module, and a cascade fusion module. The multi-source feature extractor obtains image features from different sources, and the expert judge module then performs anomaly scoring on the feature images from different sources. Finally, the cascade fusion module fuses the anomaly scores of the feature images from different sources to obtain the final anomaly score of the image. This system not only effectively alleviates the differences between different domains, but also improves the accuracy of the system's anomaly recognition and ensures inference speed, thereby enhancing the overall performance of anomaly detection. It is expected to be widely used in industrial image detection. 2) The multi-source feature extractor proposed in this paper captures features from multiple sources, stages, and scales by integrating multiple pre-trained models, thereby extracting a more comprehensive representation of industrial images and effectively alleviating the domain gap between industrial images and model training datasets; 3) The expert scoring mechanism proposed in this paper can quickly select the most representative normal samples to form an expert committee. By combining multi-source features, the expert committee can evaluate the anomaly of all samples. This expert scoring module can effectively reduce the interference caused by abnormal samples in the scoring process. By isolating the impact of abnormal samples, the expert scoring mechanism improves the accuracy and inference speed of the model, thereby achieving more reliable AC / AS. 4) This paper introduces a cascade fusion strategy to fuse anomaly score vectors from multiple sources, scales, and stages from different pre-trained models. This strategy leverages the strengths of multiple models to ensure comprehensive information fusion, enhancing the accuracy and robustness of AC / AS. By combining anomaly score vectors from different sources, stages, and scales, the cascade fusion strategy improves the overall performance of the system in anomaly detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 Schematic diagram of the differences in the focus areas of industrial product image feature extraction for different pre-trained models; (a) corresponds to the real image, (b) corresponds to the DeiT model extraction result, (c) corresponds to the DINO model extraction result, (d) corresponds to the DINOV2 model extraction result, and (e) corresponds to the CLIP model extraction result; Figure 2 This is a schematic diagram of the principle of the zero-sample industrial anomaly classification and segmentation system based on multi-source expert scoring provided by the present invention; Figure 3 Schematic diagram of feature extraction principle for a single feature extraction model; Figure 4 A flowchart of the zero-sample industrial anomaly classification and segmentation method based on multi-source expert scoring provided by the present invention; Figure 5 Visual representation of the AS results of different methods on the VisA and MVTec AD benchmark datasets; Figure 6 The time variation curves of different methods for processing a single image at different numbers. DETAILED DESCRIPTION

[0020] The following will clearly and completely describe the technical solutions of various embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of the present invention.

[0021] Example 1

[0022] This embodiment provides a zero-sample industrial anomaly classification and segmentation system (MS-ExSc) based on multi-source expert scoring. Figure 2 As shown in the figure, it includes a multi-source feature extractor, an expert scoring module and a cascade fusion module.

[0023] (1) Multi-source feature extractor

[0024] The Multi-Source Features Extractor (MSFE) extracts features from the image to be detected using two or more pre-trained feature extraction models, generating two or more image source features. Each image source feature consists of several patch feature vectors corresponding to the image to be detected.

[0025] There are domain differences between industrial product images and the training sets of pre-trained models. Different pre-trained models focus on different areas of industrial products, resulting in different representations. To obtain a more comprehensive representation of industrial product images, this paper uses different pre-trained feature extraction models to extract different features for subsequent anomaly scoring.

[0026] In this embodiment, the feature extraction model is an image encoder pre-trained by a large model such as CLIP, DINOV2 or DINO, which is expressed as The multi-source feature extractor is represented as , , where K represents the number of different pre-trained feature extraction models. The input color image has three RGB channels, represented as , initially through Processing to extract features: (1); Here, H and W represent the height and width of the original image respectively; .

[0027] In order to obtain more comprehensive information, the layers of the feature extraction module are divided into several stages, and the outputs of the several stages are respectively subjected to neighborhood aggregation of different scales to obtain the feature vectors of the patches at different stages and scales. The layer is divided into L stages. To accommodate anomalies of varying sizes, the feature extraction module uses a multi-scale local neighborhood aggregation method (LNAMD) to obtain each aggregated patch. In this example, 1×1, 3×3, and 5×5 are used as the local neighborhood aggregation ranges, and the patch block features output by each stage of the feature extraction module are aggregated at different scales to obtain locally aware patch features. For detailed operations, see Towards Total Recall in Industrial Anomaly Detection, Karsten Roth, Latha Pemula, Joaquin Zepeda, et al., 2022 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 978-1-6654-6946-3 / 22 / $31.00 ©2022 IEEE, DOI 10.1109 / CVPR52688.2022.01392. The patch features of the same patch at different stages and scales constitute the corresponding patch feature vector.

[0028] For example, given an unlabeled test image I, we first obtain multiple image source features F of the image through formula (1). The feature extraction and optimization process is as follows Figure 3 As shown, the process of other feature extraction models is the same. The patch block feature of I is defined as , M represents the number of all patches in an image, and the image source features It can be expressed as .

[0029] (2) Expert Rating Module

[0030] The Expert Scoring Module (ESM) obtains the anomaly score vector in each image source feature based on the expert committee.

[0031] In this embodiment, the method for constructing the expert committee is as follows: for a number of samples of the same category as the image to be detected, the first feature extraction model is used to extract the image source feature F of each sample. 1 , according to the first stage, N is quickly selected from several sample extraction features under 1×1 size e The most representative normal samples were then used to form the expert committee D e :

[0032] (2).

[0033] In this embodiment, the features extracted from each sample at the first stage and 1×1 scale are subjected to mutual scoring (MSM, Mutual Scoring Mechanism) to obtain the initial anomaly scores of all sample extracted features. The initial anomaly scores are sorted in ascending order, and the samples whose initial anomaly scores meet the set threshold range (here, the top s% of samples with the lowest initial anomaly scores are selected) are selected as normal samples, and an expert committee D is formed. e . D e These images in D are likely to be normal and can provide a reliable basis for further scoring. e Each expert in is used to score the test image at multiple stages and scales to generate the corresponding anomaly score map.

[0034] The calculation method of the initial anomaly score of any sample extraction feature is as follows: the minimum distance between each patch feature of the sample and the patch feature of another sample is calculated by the minimum distance function, and the sum of the minimum distances of all patches is used as the sample anomaly score of the sample relative to the other sample. The sample anomaly scores of the sample relative to other samples except itself are arranged in ascending order, and the average of the smallest P% sample anomaly scores is taken as the initial anomaly score of the sample. Repeat the above operation to obtain the initial anomaly score of each sample. Then sort in ascending order according to the initial anomaly score, select the samples whose initial anomaly scores meet the set threshold range (here, select the samples with the lowest initial anomaly scores in the top s%) as normal samples, and form an expert committee D e .

[0035] The expert scoring module calculates the minimum distance between each patch block feature in the same image source feature and each expert committee member through the minimum distance function, and obtains the anomaly score of each patch block feature in each image source feature through interval averaging operation. The anomaly scores of all patch block features constitute the anomaly score vector of the corresponding image source feature.

[0036] More specifically, each aggregated patch in the image to be detected is assigned an abnormality score by each expert at each stage and each aggregation level (i.e., each scale). Taking the mth aggregated patch in the image to be detected as an example (m=1,2,…,M, M represents the number of patches), it is The features of the lth stage and rth scale aggregation level can be expressed as z k , and Expert I e The feature representation of all patches in is , so the anomaly score It is z k and The minimum distance between them is: (3); Among them, mindist represents the minimum distance function, and the distance can be Euclidean distance, Mahalanobis distance, cosine similarity distance, etc. Therefore, each expert D e The anomaly score vector in can be expressed as Next, Perform the interval averaging (IA) operation. The interval averaging operation is to take the average of the minimum distances between the patch block feature and the expert committee members that meet the minimum distance interval requirement (i.e., sort by minimum distance in ascending order and take the minimum distance of the top X%); it is expressed as: (4); Among them, IA means In ascending order of minimum distance, the minimum distances of the expert committee members in the first X% interval are averaged. In order to generate a detailed anomaly score for each image patch from different sources, the process outlined in Equations (3) and (4) is performed on the patch features under each combination of r and l.

[0037] The expert committee scores all image source features, and the obtained score vector is represented as A. , is the anomaly score vector from the k-th image source feature, .

[0038] (3) Cascade Fusion Module

[0039] The Cascade Fusion Module (CFM) is used to perform cascade fusion processing on the anomaly score vectors of each image source feature to obtain a fused anomaly score vector, and generate the final anomaly score map based on the fused anomaly score vector, including anomaly classification score and anomaly segmentation output.

[0040] The cascade fusion module includes multi-source fusion experts (D S ), multi-stage fusion experts (D L) and multi-scale fusion experts (D R ).

[0041] Multi-source fusion experts fuse the anomaly scores of the same patch block features from different image source features through the first one-class support vector machine (OCSVM) to obtain the first fused anomaly score of each patch block feature.

[0042] Specifically, for The anomaly score of the extracted mth patch at the lth stage and the rth scale aggregation level is The anomaly score vector from K different image source features is expressed as , then the multi-source fusion expert D S The first fused anomaly score of the patch is obtained by fusion through a learnable first one-class support vector machine (OCSVM). : (5).

[0043] The multi-stage fusion expert averages the first fusion anomaly scores of the patch features of the same patch at different stages at the same scale to obtain the second fusion anomaly scores of the same patch at different scales.

[0044] Specifically, from Equations (3), (4) and (5), we can obtain the abnormality score vector of the mth patch at the rth scale aggregation level, which is expressed as .

[0045] Through multi-stage fusion expert D L The average operation of the anomaly score vector at the r-th scale aggregation level The fusion is performed to obtain the second fusion anomaly score of the same patch at this scale: (6).

[0046] The anomaly score vectors of the patch at different aggregation levels are expressed as .

[0047] The multi-scale fusion expert uses the second fusion anomaly scores of the same patch at different scales to obtain the third fusion anomaly score of the corresponding patch through the second single-class support vector machine.

[0048] The anomaly scoring vectors from different scale aggregation levels are analyzed by the second one-class support vector machine. The fusion is performed to obtain the third fusion anomaly score of the corresponding patch block: (7).

[0049] The third fused anomaly scores of all patches constitute the fused anomaly score vector .

[0050] The above-mentioned cascade fusion module also includes an anomaly output module, which is used to reshape the fused anomaly score vector and then upsample the reshaped vector to match the resolution of the original image I, thereby obtaining an anomaly segmentation output; at the same time, according to the set threshold, the maximum anomaly score is obtained from the anomaly score vector As the image-level anomaly classification score of the image, the final anomaly score map is generated.

[0051] Reshaping the fused anomaly score vector is to convert the fused anomaly score vector into a two-dimensional representation, and then upsampling it to match the resolution of the original image I.

[0052] This embodiment also provides a zero-sample industrial anomaly classification and segmentation method based on multi-source expert scoring, such as Figure 4 As shown, using the above system, follow the steps below:

[0053] Step 1: Perform multi-source feature extraction on the image to be detected to obtain two or more image source features; each image source feature includes several patch block feature vectors corresponding to the image to be detected.

[0054] In this step, the feature extraction module performs several stages of feature extraction on the image to be detected, and the outputs of the several stages are respectively subjected to neighborhood aggregation of different scales to obtain the patch block feature vectors at different stages and scales of each patch block.

[0055] Step 2: Based on the expert committee, obtain the abnormality score vector in each image source feature.

[0056] In this step, the expert scoring module uses the minimum distance function to calculate the minimum distance between each patch block feature in the same image source feature and each expert committee member based on the expert committee constructed previously. Then, according to the minimum distance interval requirement, the interval averaging operation is performed to obtain the anomaly score of each patch block feature in each image source feature. The anomaly scores of all patch block features constitute the anomaly score vector of the corresponding image source feature.

[0057] Step 3: Perform cascade fusion processing on the anomaly score vectors of each image source feature to obtain a fused anomaly score vector, and generate the final anomaly score map based on the fused anomaly score vector, including anomaly classification score and anomaly segmentation output.

[0058] This step includes the following sub-steps: Step 31, fusing the anomaly scores of the same patch block feature from different image source features through a first single-class support vector machine to obtain a first fused anomaly score of each patch block feature; Step 32: averaging the first fused anomaly scores of the patch features at different stages of the same patch at the same scale to obtain a second fused anomaly score of the same patch at different scales; Step 33: The second fused anomaly scores at different scales of the same patch are used to obtain a third fused anomaly score of the corresponding patch through a second single-class support vector machine; the third fused anomaly scores of all patches constitute a fused anomaly score vector; Repeat steps 31 to 33 above to calculate the third fused anomaly scores of all patches of the image to be detected. The third fused anomaly scores of all patches constitute a fused anomaly score vector. In step 34, the fused anomaly score vector is reshaped, and then the reshaped vector is upsampled to match the resolution of the original image I, thereby obtaining an anomaly segmentation output; at the same time, the maximum anomaly score is obtained from the fused anomaly score vector as the image-level anomaly classification score of the image, and the final anomaly score map is generated.

[0059] Experimental example

[0060] The following describes the zero-shot industrial anomaly classification and segmentation system (MS-ExSc) and its methods based on multi-source expert scoring, as described in Example 1, using the MVTec AD dataset and the VisA dataset. The AC task uses three metrics: AUROC, AP, and F1; the AS task also includes the PRO metric to fairly evaluate the segmentation performance of anomalies at different scales.

[0061] The MVTec AD dataset, developed by MVTec, is widely used as a benchmark for unsupervised anomaly detection. It contains high-resolution RGB images (ranging from 700² to 1024² pixels) of industrial products and textures. The dataset encompasses 10 object categories and 5 texture categories, with normal images included in the training set for each category. Comprising 4096 normal samples and 1258 anomaly samples, the dataset is a valuable resource for evaluating anomaly detection methods in industrial settings.

[0062] The VisA dataset, released in 2021 by the Institute of Automation at Tsinghua University, is specifically designed for visual inspection tasks. The dataset contains 10,821 high-resolution RGB images (1000 × 1500 pixels) covering 12 different objects across three domains. Comprising 9,621 normal samples and 1,200 abnormal samples, the dataset is a valuable resource for evaluating anomaly detection methods in industrial settings.

[0063] Both of the above datasets specify a training set and a test set. In the test set, each abnormal sample is provided with a true label and segmentation result. This experiment is performed on the test set.

[0064] In this experimental example, the multi-source feature extractor includes two feature extraction models: ViT-L-14( ) and ViT-L-14-336( ), which are the image encoders of DINOV2 and CLIP respectively. ViT-L-14 is pre-trained by DINOV2, and ViT-L-14-336 is pre-trained by CLIP. Both of these pre-trained models consist of 24 layers, which are divided into 4 stages {1, 2, 3, 4}, corresponding to the layers {6, 12, 18, 24} respectively. This division enables the extraction of patch-level features from the relevant layers. Before inputting the image into the feature extraction model, the image is first resized to 518×518 pixels.

[0065] In the expert scoring module, the Euclidean distance is selected as the minimum distance function. Use to score the samples in the above two datasets against each other, and an expert committee is screened and formed by selecting the s% samples with the lowest initial anomaly scores (when n≤50, s% = 25%; when 50 < n ≤ 100, s% = 20%; when n≥150, s% = 15%, n represents the number of samples). In the execution of interval averaging (IA), the values within the lowest X% = 30% range are selected. All experiments are conducted on NVIDIA RTX 4080s GPUs.

[0066] Using the above zero-shot industrial anomaly classification and segmentation system based on multi-source expert scoring (MS-ExSc), anomaly score maps for each image are obtained according to steps 1 - step S3.

[0067] A comprehensive comparative analysis of MS-ExSc was conducted on the VisA and MVTec AD datasets, covering both image-level classification and pixel-level segmentation tasks. The proposed system (MS-ExSc) was compared with several state-of-the-art zero- and few-shot systems, including WinCLIP, APRIL-GAN, MuSc, RegAD, PatchCore, GraphCore, and ACR. Detailed AC / AS results on both datasets are presented in Table 1. The results demonstrate that MS-ExSc outperforms previous state-of-the-art (SOTA) algorithms, particularly on the VisA dataset, which contains many subtle anomalies that are difficult to detect. These anomalies require a strong representation of industrial product images—a capability often overlooked by previous methods. Furthermore, segmenting these anomalies is a challenging task, with existing zero- and few-shot methods performing significantly worse on VisA than on MVTec AD. The outstanding performance of MS-ExSc underscores its effectiveness in addressing these challenges. MS-ExSc leverages multiple pre-trained models to extract more comprehensive image representations, employs an expert scoring module to mitigate the impact of outliers on the overall scoring process, and fuses different anomaly scoring vectors via a cascade fusion module. This enables it to surpass state-of-the-art performance on both the VisA dataset (which contains numerous subtle anomalies) and the MVTec AD dataset. This demonstrates the effectiveness and universality of the proposed method and the robustness of MS-ExSc in handling complex industrial AC / AS tasks.

[0068]

[0069] Visualization results of the industrial image anomaly score map generated by MS-ExSc, such as Figure 5 The results show the effectiveness of this approach. Compared to the previous state-of-the-art zero-shot method, MS-ExSc excels in detecting and accurately localizing subtle defects while significantly reducing false positives. This demonstrates that MS-ExSc extracts more comprehensive information, thereby enhancing detection capabilities. The expert scoring module plays a key role in this, effectively suppressing the scoring of anomalous images, thereby reducing false positives. Furthermore, the cascaded fusion module is able to cleverly integrate anomaly score maps from different sources, stages, and layers without introducing additional false positives.

[0070] Moreover, in the MS-ExSc provided by the present invention, the time required for the model to process a single image under different numbers is also provided. The results are as follows: Figure 6 The experimental results show that although MS-ExSc introduces more pre-trained models, which requires more time to extract image features and perform anomaly scoring, the speed of MS-ExSc in processing images does not decrease significantly, which further proves the effectiveness of the present invention.

[0071] Those skilled in the art will appreciate that the embodiments herein are intended to help readers understand the principles of the present invention, and it should be understood that the scope of protection of the present invention is not limited to such specific descriptions and embodiments. Those skilled in the art may make various other specific variations and combinations based on the technical teachings disclosed herein without departing from the essence of the present invention, and such variations and combinations are still within the scope of protection of the present invention.

Claims

1. A zero-shot industrial anomaly classification and segmentation system based on multi-source expert scoring, characterized by: include: A multi-source feature extractor is used to extract features from the image to be detected using two or more pre-trained feature extraction models, respectively, to obtain two or more image source features; each image source feature includes several patch block feature vectors corresponding to the image to be detected; The expert scoring module obtains the abnormality scoring vector in each image source feature based on the expert committee; The cascade fusion module is used to perform cascade fusion processing on the anomaly score vectors of each image source feature to obtain a fused anomaly score vector, and generate the final anomaly score map based on the fused anomaly score vector, including anomaly classification score and anomaly segmentation output.

2. The zero-shot industrial anomaly classification and segmentation system based on multi-source expert scoring according to claim 1 is characterized in that: The feature extraction model is a CLIP, DINOV2 or DINO pre-trained image encoder.

3. The zero-sample industrial anomaly classification and segmentation system based on multi-source expert scoring according to claim 1 or 2 is characterized in that: The layers of the feature extraction module are divided into several stages, and the outputs of the several stages are respectively subjected to neighborhood aggregation of different scales to obtain the patch block feature vectors at different stages and scales of each patch block.

4. The zero-sample industrial anomaly classification and segmentation system based on multi-source expert scoring according to claim 3 is characterized in that: Taking 1×1, 3×3, and 5×5 as the local neighborhood aggregation range respectively, the patch block features output by each stage of the feature extraction module are aggregated at different scales to obtain the patch block features at the corresponding scales; the patch block features at different stages and scales of the same patch block constitute the corresponding patch block feature vector.

5. The zero-sample industrial anomaly classification and segmentation system based on multi-source expert scoring according to claim 3 is characterized in that: The method for constructing the expert committee is as follows: for several samples of the same category as the image to be detected, feature extraction is performed through any feature extraction model in the multi-source feature extractor to obtain several sample image source features; the extracted features of each sample at any stage and any scale are scored against each other to obtain the initial anomaly score of all sample extracted features, and the samples are sorted in ascending order according to the initial anomaly score, and the samples whose initial anomaly score meets the set threshold range are selected as normal samples to form the expert committee.

6. The zero-sample industrial anomaly classification and segmentation system based on multi-source expert scoring according to claim 5 is characterized in that: The expert scoring module calculates the minimum distance between each patch block feature in the same image source feature and each expert committee member through the minimum distance function, and obtains the anomaly score of each patch block feature in each image source feature through interval averaging operation. The anomaly scores of all patch block features constitute the anomaly score vector of the corresponding image source feature.

7. The zero-sample industrial anomaly classification and segmentation system based on multi-source expert scoring according to claim 6 is characterized in that: The minimum distance function is Euclidean distance, Mahalanobis distance or cosine similarity distance; the interval averaging operation is to take the average value of the minimum distance between the patch block feature and the expert committee that meets the minimum distance interval requirement.

8. The zero-sample industrial anomaly classification and segmentation system based on multi-source expert scoring according to claim 6 is characterized in that: The cascade fusion module includes a multi-source fusion expert, a multi-stage fusion expert and a multi-scale fusion expert; the multi-source fusion expert fuses the anomaly scores of the same patch block features from different image source features through a first single-class support vector machine to obtain a first fused anomaly score of each patch block feature; the multi-stage fusion expert averages the first fused anomaly scores of the patch block features of the same patch block at different stages at the same scale to obtain a second fused anomaly score of the same patch block at different scales; the multi-scale fusion expert uses a second single-class support vector machine to obtain a third fused anomaly score of the corresponding patch block based on the second fused anomaly scores of the same patch block at different scales; the third fused anomaly scores of all patch blocks constitute a fused anomaly score vector.

9. The zero-shot industrial anomaly classification and segmentation system based on multi-source expert scoring according to claim 8 is characterized in that: The cascade fusion module also includes an anomaly output module for reshaping the fused anomaly score vector and then upsampling the reshaped vector to match the resolution of the original image to obtain an anomaly segmentation output; at the same time, the maximum anomaly score is obtained from the fused anomaly score vector as the image-level anomaly classification score of the image to generate a final anomaly score map.

10. A zero-shot industrial anomaly classification and segmentation method based on multi-source expert scoring, characterized by: Using the system according to any one of claims 1 to 9 is carried out according to the following steps: Step 1: Perform multi-source feature extraction on the image to be detected to obtain two or more image source features; each image source feature includes several patch block feature vectors corresponding to the image to be detected; Step 2: Obtain anomaly score vectors in each image source feature based on the expert committee; Step 3: Perform cascade fusion processing on the anomaly score vectors of each image source feature to obtain a fused anomaly score vector, and generate the final anomaly score map based on the fused anomaly score vector, including anomaly classification score and anomaly segmentation output.

11. The zero-sample industrial anomaly classification and segmentation method based on multi-source expert scoring according to claim 10 is characterized in that: Step 3 The following steps are included: Step 31, fusing the anomaly scores of the same patch block feature from different image source features through a first single-class support vector machine to obtain a first fused anomaly score of each patch block feature; For the mth patch extracted by the multi-source feature extractor, the anomaly score vector from K different image source features is expressed as , is the anomaly score at the lth stage and the rth scale aggregation level; the multi-source fusion expert D S The first fused anomaly score of the patch is obtained by fusing the anomaly score vectors from K non-image source features through the first single-class support vector machine. : ; At the same scale, the first fusion anomaly score of the patch features of the same patch at different stages constitutes the anomaly score vector of the r-th scale aggregation level, which is expressed as ; Step 32: averaging the first fused anomaly scores of the patch features at different stages of the same patch at the same scale to obtain a second fused anomaly score of the same patch at different scales; Through multi-stage fusion expert D L The average operation of the anomaly score vector at the r-th scale aggregation level The fusion is performed to obtain the second fusion anomaly score of the same patch at this scale: ; The anomaly score vectors of the patch at different aggregation levels are expressed as ; Step 33: The second fused anomaly scores at different scales of the same patch are used to obtain a third fused anomaly score of the corresponding patch through a second single-class support vector machine; the third fused anomaly scores of all patches constitute a fused anomaly score vector; Multi-scale fusion expert D R The anomaly scoring vectors from different scale aggregation levels are analyzed by the second one-class support vector machine. The fusion is performed to obtain the third fusion anomaly score of the corresponding patch block: ; Repeat steps 31 to 33 to calculate the third fused anomaly scores of all patches of the image to be detected. The third fused anomaly scores of all patches constitute a fused anomaly score vector. In step 34, the fused anomaly score vector is reshaped, and then the reshaped vector is upsampled to match the resolution of the original image to obtain an anomaly segmentation output; at the same time, the maximum anomaly score is obtained from the fused anomaly score vector as the image-level anomaly classification score of the image, and the final anomaly score map is generated.

Citation Information

Cited By

  • Zero-sample multi-mode industrial defect segmentation method based on test sample relation calculation

    CN121640051A