Defect classification and segmentation method and system based on unsupervised and weak supervised combination
By combining unsupervised and weak supervision methods in defect automation detection technology, a feature memory bank is built and pseudo-tagged generation is designed, and a weak supervision network is designed, which solves the problem of existing technology dependence on labeled data, and achieves efficient and accurate industrial defect detection.
Patent Information
- Application Number
- CN202510091424.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-06-24
AI Technical Summary
The existing defect automation detection technology has a dependence on a large amount of labeled data, which leads to high labeling costs. The traditional methods perform poorly when processing complex images, making it difficult to meet the needs of efficient real-time detection.
A defect classification and segmentation method based on the combination of unsupervised and weak supervision is proposed. By constructing a feature memory bank and generating pseudo-labels, a weak-supervised network based on category activation graphs is designed, and a semantic segmentation network is trained using defect-free samples and pseudo-classification labels.
It realizes the accurate identification of various types of industrial defects without relying on large amounts of labeled data, reduces the cost of data labeling, improves the accuracy and robustness of defect detection, and has strong generalization capabilities.
Smart Images

Figure CN120198703A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of network security, and particularly relates to a defect classification and segmentation method and system based on the combination of unsupervised and weakly supervised learning. Background Art
[0002] The existing defect automatic detection technologies are mainly divided into two categories, namely defect detection based on traditional methods and defect detection based on deep learning methods. Defect detection based on traditional methods mainly relies on image processing and classical machine learning techniques, and usually includes steps such as preprocessing (such as denoising and contrast enhancement), feature extraction, feature selection and dimensionality reduction, classification (such as support vector machines and K-nearest neighbors), and post-processing of results (such as morphological operations).
[0003] Defect automatic detection based on deep learning can be further divided into defect detection based on fully supervised learning, weakly supervised learning, and unsupervised / semi-supervised learning according to the dependence on the annotation information of defect samples during the training process.
[0004] Defect detection methods based on fully supervised learning and weakly supervised learning rely on accurately labeled data and less labeled information respectively, and have high detection performance by directly learning sample features or combining inference capabilities. In order to reduce the dependence on expensive annotations during the training process of deep learning methods based on full supervision and weak supervision, defect detection methods based on unsupervised / semi-supervised learning have been proposed.
[0005] However, the defect detection methods based on traditional methods have certain limitations. First of all, these methods usually rely on manually designed features and rules, lacking generality, so they need to be customized according to specific application scenarios. Secondly, since traditional methods often rely on low-level feature extraction technologies such as edge detection and texture analysis, their performance in processing complex images is poor. Especially in the case of poor image quality or the presence of noise, the detection accuracy drops significantly. In addition, the calculation process of traditional vision algorithms is relatively cumbersome, the processing speed is slow, and as the image size and complexity increase, the performance drops sharply, making it difficult to meet the requirements of efficient real-time detection. Generally speaking, the effect of traditional machine vision methods in defect detection is not as good as that of deep learning-based methods, especially in terms of accuracy and robustness. However, defect detection methods based on fully supervised learning and weakly supervised learning usually require a large number of defect samples. This way of relying on a large amount of labeled data brings a high labeling cost, especially in the case where defect samples are difficult to obtain or the labeling process is cumbersome. For example, CN116542929A, CN119048497A, and CN119007174A require a large number of expensive bounding box annotations or pixel-level annotations to obtain better defect detection performance. In addition, there are mainly two approaches for defect detection methods based on unsupervised / semi-supervised learning, namely reconstruction-based methods such as CN117237309A, CN118014941A, CN116630696A, and representation-based methods such as CN116912173A and CN115187525A. The reconstruction-based methods have the problems of easily ignoring small defects and being difficult to reconstruct defect details well; in the representation-based methods, when constructing the feature memory bank, smaller image patches improve the defect localization accuracy, but significantly increase the storage requirement of computing resources; while larger image patches are difficult to achieve precise localization. Summary of the Invention
[0006] In view of the above problems, the present invention provides a defect classification and segmentation method and system based on the combination of unsupervised and weakly supervised learning, aiming to provide an unsupervised deep learning method and system with strong optimization, small computing resource occupancy, and excellent defect classification and defect segmentation capabilities.
[0007] According to the first aspect of the embodiments of the present disclosure, there is provided a defect classification and segmentation method based on the combination of unsupervised and weakly supervised learning, the method comprising the following steps:
[0008] Using a training data set containing only defect-free samples, extracting the feature representation of the defect-free samples through a feature extractor and storing it in a feature memory bank;
[0009] Comparing the mixed data set of unlabeled samples including defects and defect-free samples with the features in the feature memory bank to determine pseudo-classification labels;
[0010] Design a weakly supervised network based on class activation maps for feature extraction in images, a semantic segmentation network and its loss function, and train the semantic segmentation network using a training dataset of defect-free samples and pseudo-classification label data.
[0011] In some embodiments, the process of storing the feature memory bank specifically includes the following steps:
[0012] Feature extraction, extracting multi-level semantic and structural information from the input samples;
[0013] Feature metric, introducing a distance metric method to evaluate the similarity between different samples;
[0014] Memory bank sampling, selecting representative samples and storing them in the feature memory bank.
[0015] In some embodiments, the features extracted from the third and fourth layers of the Wide-ResNet50 network are used and stored in the feature memory bank. The features include local structural information and low-level patterns of the samples, as well as abstract feature information.
[0016] In some embodiments, the earth mover's distance is used as the distance metric method for feature metric, including the following steps:
[0017] Calculate the histogram for each feature to obtain its probability distribution;
[0018] Normalize the histogram to obtain the cumulative distribution function;
[0019] The earth mover's distance value EMD is the sum of the absolute values of the differences between the cumulative distribution functions of two variables X and Y:
[0020] EMD(X,Y) = ∫|F X (x) - F Y (x)|dx
[0021] where F X (x) and F Y (x) respectively represent the cumulative distribution functions of variables X and Y at point x.
[0022] In some embodiments, the core subsampling method is used to select representative samples and store them in the feature memory bank. The optimization problem of core subsampling can be expressed as:
[0023]
[0024] where X is the original dataset, containing n samples, S is the core subset, containing k samples, and k << n, and d((x i , s j ) represents the sample x iand the subset sample s j the distance between;
[0025] The greedy algorithm is used to solve the optimization problem of core subsampling, including the following steps:
[0026] Select an initial sample from the dataset as the first element of the core subset; then repeatedly select a sample from the remaining samples to minimize its distance from the current subset to maximize representativeness until the core subset reaches a predetermined size.
[0027] In some embodiments, a mixed dataset including unlabeled defective and non-defective samples is compared with the features in the feature memory bank to determine pseudo-classification labels, specifically including:
[0028] Process each unlabeled sample in the unlabeled mixed dataset to obtain a feature representation;
[0029] Perform a nearest neighbor search in the feature memory bank to calculate the anomaly score between the obtained feature representation and the nearest feature vector stored in the memory bank;
[0030] Based on the anomaly score array, calculate the precision and recall at different thresholds, and record these thresholds;
[0031] Calculate the F1-score based on the precision and recall, and the expression is: where Precision represents precision and Recall represents recall;
[0032] Traverse all thresholds to find the threshold T that maximizes the F1-score tmp , and use this threshold as the optimal threshold for determining whether a sample is abnormal;
[0033] Normalize the anomaly score using the optimal threshold to obtain the normalized anomaly score. The specific formula is as follows:
[0034]
[0035] Determine the pseudo-classification label of the sample based on the normalized anomaly score.
[0036] In some embodiments, the weakly supervised network based on the class activation map uses the pre-trained ResNeSt50 as the feature extraction network, and after the feature map extracted by ResNeSt50, a convolutional layer with a kernel size of 1×1 is added to reduce the number of channels of the feature map to the number of target classes, thereby generating the target feature map;
[0037] Apply the sigmoid activation function on the target feature map to generate a probability heat map H2, where each pixel value in the heat map H2 represents the probability of predicting a defect at that location;
[0038] Perform threshold segmentation on the probability heat map H2 to generate a preliminary binary segmentation map;
[0039] For the binary segmentation map, morphological operations are used to smooth the boundaries of the target region and remove isolated small regions.
[0040] In some embodiments, the DeepLabV3+ architecture is used as the basic framework of the segmentation network, including an encoder and a decoder; the core part of the encoder is the ResNet architecture. After the input data is processed by the backbone network, two different levels of feature embeddings are generated: a low-level feature embedding Embedding1 and a high-level feature embedding Embedding2. Among them, Embedding1 is fed into the decoder, and Embedding2 is passed to the ASPP module for multi-scale feature learning;
[0041] The ASPP module processes Embedding2 through five convolutional operations with different dilation rates to capture multi-scale context information;
[0042] In the decoder part, the output features of the ASPP module are merged with the low-level feature Embedding1 and restored to the same resolution as the original input image through step-by-step upsampling operations, thereby generating a segmentation result.
[0043] In some embodiments, the loss function of the segmentation network includes an image-level classification loss L cls and a pixel-level segmentation loss L seg , and the overall training loss is defined as: L = αL cls + βL seg , where α and β are weight factors used to balance the importance of the classification loss and the segmentation loss. The image-level classification loss L cls is used to ensure that the model can correctly distinguish image categories, and the segmentation loss L seg is used to improve the segmentation accuracy of the model at the pixel level;
[0044] The image-level classification loss L cls is defined as the cross-entropy loss between the pseudo-class label C class and the model-predicted class C predicted : L cls = CE(C class , C predicted ); The pixel-level segmentation loss L seg is defined as the cross-entropy loss between the pseudo-segmentation map S pseudo and the model-predicted segmentation map S predictedThe cross-entropy loss between: L seg = CE(S pseudo , S predicted ).
[0045] According to a second aspect of the embodiments of the present disclosure, there is provided a defect classification and segmentation system based on the combination of unsupervised and weakly supervised learning. The system includes:
[0046] A feature memory bank construction module, configured to use a training data set containing only defect-free samples, extract feature representations of the defect-free samples through a feature extractor, and store them in the feature memory bank;
[0047] A pseudo-label generation module, configured to compare a mixed data set of unlabeled samples including defective and defect-free samples with the features in the feature memory bank to determine pseudo-classification labels;
[0048] A segmentation network training module, configured to design a weakly supervised network based on a class activation map for extracting features in an image, a semantic segmentation network, and its loss function, and train the semantic segmentation network using the training data set of defect-free samples and the pseudo-classification label data.
[0049] A defect classification and segmentation method and system based on the combination of unsupervised and weakly supervised learning provided by the embodiments of the present disclosure is used for industrial defect detection. The method is called Semi-Patchcore. Innovatively, it only uses defect-free samples and a large number of unlabeled samples for training, solving the limitations of traditional supervised learning methods in the case of high label costs. Through a carefully designed two-stage training process, Semi-Patchcore realizes effective detection and localization of unseen defect features. This method not only has strong generalization ability, but also can accurately identify various types of industrial defects without relying on a large amount of labeled data. The specific advantages are as follows:
[0050] (1) The present invention discloses a defect classification and segmentation method Semi-Patchcore based on the combination of unsupervised and weakly supervised learning. This method combines the advantages of weakly supervised and unsupervised learning algorithms. By relying only on defect-free samples and a large number of unlabeled samples for training, it breaks through the dependence of traditional methods on a large amount of labeled data and realizes efficient defect detection and localization.
[0051] (2) The present invention introduces the Earth Mover's Distance (EMD), providing a more accurate method for measuring the distance between different features in industrial surface images. This method can better adapt to complex defect detection scenarios and improve the accuracy of pseudo-label generation.
[0052] (3) By training the model using only defect-free samples and a large amount of unlabeled data, the present invention significantly reduces the data annotation cost, reducing the annotation cost by approximately 53%. At the same time, it outperforms previous semi-supervised and unsupervised methods in terms of defect detection performance, demonstrating strong practical value.
[0053] (4) Experimental verification was carried out on four datasets, namely MVTecAD, BTAD, DAGM, and KSDD2. The metrics between classical unsupervised algorithms, the current state-of-the-art unsupervised algorithms, and the Semi-Patchcore algorithm proposed by the present invention were compared, and visualization pictures were given, which can prove that the Semi-Patchcore model proposed by the present invention has good generalization.
[0054] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. Brief Description of the Drawings
[0055] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention.
[0056] Figure 1 is a schematic flow chart of a defect classification and segmentation method based on the combination of unsupervised and weakly supervised in an embodiment of the present invention;
[0057] Figure 2 is an overall architecture diagram of a defect classification and segmentation method based on the combination of unsupervised and weakly supervised in an embodiment of the present invention;
[0058] Figure 3 is a schematic flow chart of a method for constructing a feature memory bank in an embodiment of the present invention;
[0059] Figure 4 is a schematic flow chart of a method for generating pseudo labels in an embodiment of the present invention;
[0060] Figure 5 is a result diagram of generating pseudo labels using two different methods in an embodiment of the present invention;
[0061] Figure 6 is an overall framework diagram of a segmentation network in an embodiment of the present invention;
[0062] Figure 7 is a schematic structural diagram of a defect classification and segmentation system based on the combination of unsupervised and weakly supervised in an embodiment of the present invention;
[0063] Figure 8 is the visualization result on the MVTecAD dataset in an embodiment of the present invention;
[0064] Figure 9It is the visualization result in the BTAD dataset in the embodiments of the present invention;
[0065] Figure 10 It is a comparison graph of only unsupervised methods and unsupervised + weakly supervised methods on different datasets in the embodiments of the present invention;
[0066] Figure 11 It is an ablation experiment using real labels and pseudo - labels in the embodiments of the present invention, and schematic diagrams of experiments using real labels and pseudo - labels respectively;
[0067] Figure 12 It is a comparison graph of using true value labels and pseudo - labels on different datasets in the embodiments of the present invention. Detailed implementation manners
[0068] The present invention will be further described in detail below with reference to the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. Additionally, it should be noted that for the convenience of description, only parts related to the present invention rather than all structures are shown in the drawings.
[0069] Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the steps as sequential processes, many of the steps can be implemented in parallel, concurrently, or simultaneously. In addition, the order of the steps can be rearranged. The process can be terminated when its operations are completed, but it can also have additional steps not included in the drawings. The process can correspond to a method, function, procedure, sub - routine, sub - program, etc.
[0070] Existing supervised defect detection methods usually rely on a large number of accurately labeled defect samples, which is difficult to meet in industrial applications, resulting in limited promotion in the application of the defect detection field. Unsupervised methods do not rely on defect samples and annotation information, and only require normal samples for training, effectively reducing the annotation cost. However, existing unsupervised / semi - supervised defect detection methods still face problems such as large computational resource storage requirements, inaccurate defect localization, and poor generalization ability. To address these problems, the present invention proposes a novel semi - supervised defect detection method. This method constructs a feature memory bank of defect - free samples and compares it with unlabeled mixed data to generate high - quality pseudo - labels. Using these pseudo - labels, this method further designs a weakly supervised segmentation network based on class activation maps, effectively improving the detection and localization ability of subtle defects and having excellent generalization performance.
[0071] The embodiments of the present invention provide the following embodiments for a defect classification and segmentation method and system based on the combination of unsupervised and weakly supervised:
[0072] As Figure 1 shown, a defect classification and segmentation method based on the combination of unsupervised and weakly supervised learning includes the following steps:
[0073] S1. Use a training data set containing only defect-free samples, extract the feature representations of the defect-free samples through a feature extractor, and store them in a feature memory bank;
[0074] S2. Compare the mixed data set of unlabeled samples (including defective and defect-free samples) with the features in the feature memory bank to determine pseudo-classification labels;
[0075] S3. Design a weakly supervised network based on class activation maps for extracting features in images, a semantic segmentation network and its loss function, and use the training data set of defect-free samples and pseudo-classification label data to train the semantic segmentation network.
[0076] As Figure 2 shown, the overall algorithm (Semi-Patchore) process of the present invention can be divided into two stages: the first stage is a label enhancement strategy based on unsupervised learning, and the second stage is an overall algorithm for a defect segmentation strategy based on weakly supervised learning. The process of the present invention is as Figure 1 shown. First, use a training data set containing only defect-free samples to construct a memory bank. These samples have low acquisition costs and wide sources because they do not contain anomalies. Then, the model compares the unlabeled mixed data set (including defective and defect-free samples) with the features in the memory bank to determine pseudo-classification labels. Finally, use the pseudo-labels to train the segmentation network. Generally speaking, the present invention consists of three core modules: constructing a feature memory bank, generating pseudo-labels, and training the segmentation network. Among them, the construction of the feature memory bank and the generation of pseudo-labels module constitute a label enhancement strategy based on unsupervised learning, and the training of the segmentation network constitutes a fine defect segmentation strategy based on weakly supervised learning.
[0077] The label enhancement strategy based on unsupervised learning is mainly divided into two steps, namely constructing a feature memory bank and generating pseudo-labels. Among them, constructing a feature memory bank means using an industrial data set without defects, extracting the feature representations of samples through a feature extractor (a pre-trained convolutional neural network), and storing these features in the feature memory bank. This memory bank can be regarded as the global statistical information of the features of defect-free samples and can be used for similarity measurement and anomaly detection of new samples.
[0078] In the step of generating pseudo-labels, the model calculates the anomaly score by comparing with the features in the memory bank and calculating the feature distance between the sample to be measured and the defect-free samples. Based on this anomaly score, pseudo-labels are further generated, and the generation of pseudo-labels can provide a supervision signal for the subsequent segmentation network.
[0079] The specific process of storing in the feature memory bank in S1 includes the following steps:
[0080] S11. Feature extraction, extracting multi-level semantic and structural information from the input samples;
[0081] S12. Feature measurement, introducing a distance measurement method to evaluate the similarity between different samples;
[0082] S13. Memory bank sampling, storing representative samples in the feature memory bank.
[0083] Specifically, as Figure 3 shown, in the process of constructing the feature memory bank based on unsupervised learning, three key strategies are invented to ensure that the memory bank can effectively capture the feature distribution of normal samples. These strategies include a feature extraction module, a feature measurement strategy, and a memory bank sampling strategy. First, the feature extraction module is responsible for extracting multi-level semantic and structural information from the input samples to ensure that the obtained features can accurately represent the surface and detail characteristics of the image. Then, in the feature measurement strategy, a distance measurement method suitable for industrial image detection is introduced to evaluate the similarity between different samples to better handle feature deviations caused by factors such as rotation and illumination. Finally, the memory bank sampling strategy is used to optimize the structure of the memory bank by effectively selecting representative samples to ensure the efficiency and robustness of the memory bank.
[0084] In S1, the features extracted from the third and fourth layers of the Wide-ResNet50 network are used and stored in the feature memory bank. The features include the local structure information and low-level patterns of the samples, as well as abstract feature information.
[0085] Specifically, the purpose of feature extraction is to extract multi-level features from defect-free samples to accurately represent the structural and texture information of the samples. Commonly used feature extraction networks include ResNet, VGG, DenseNet, etc. These network structures have their own characteristics and are suitable for different task requirements. Based on the excellent performance of ResNet, the present invention uses a pre-trained Wide-ResNet50 network as the basis for feature extraction. Wide-ResNet50 has strong feature learning ability and can capture multi-level characteristics of images, especially when representing complex industrial surface features. However, using an overly deep network may result in overly abstract extracted features, ignoring details and surface information. To balance the abstraction and detail retention of features, the present invention utilizes the features extracted from the third and fourth layers of the Wide-ResNet50 network and stores them in a memory bank. This is to better capture the surface details of the samples while maintaining the feature abstraction ability. Specifically, the features extracted from the third and fourth layers not only contain local structural information and low-level patterns of the samples, such as textures, edges, and small shapes, but also contain a little abstract feature information. These features fuse shallow and deep information, helping the model to focus on the subtle changes in the surface and deep structures, and can help the model more accurately identify defective images.
[0086] Feature measurement is one of the core steps in constructing the memory bank and directly affects the model's ability to distinguish between defect-free and defective samples. Common distance metrics include Euclidean distance, Manhattan distance, etc. However, directly using Euclidean distance may introduce errors when measuring the similarity between different features, especially when dealing with industrial surface image features. Image features are usually affected by factors such as rotation, making the Euclidean distance between the rotated defect-free image features and the original image features relatively large, thus making it difficult to accurately reflect the actual similarity when constructing the feature memory bank. To solve this problem, the present invention uses the Earth Mover's Distance (EMD) for feature measurement to improve the robustness to local and global structural changes of the images.
[0087] EMD is a distance metric method based on the optimal transport theory. It measures the similarity between samples by calculating the minimum transport cost between feature distributions, thus being able to more robustly reflect the similarity of image features even when the images have geometric changes such as rotation and scaling. Compared with Euclidean distance, EMD performs better in terms of rotation invariance, so it is more suitable for feature selection in the memory bank in the surface defect detection task.
[0088] When actually calculating the EMD of two discrete variables, based on the Cumulative Distribution Function (CDF), the distance metric method in S12 uses the Earth Mover's Distance for feature measurement, including the following steps:
[0089] Calculate the histogram for each feature to obtain its probability distribution.
[0090] Normalize the histogram so that its sum is 1 to obtain the cumulative distribution function.
[0091] The EMD value is the sum of the absolute values of the differences between the CDFs of the two variables:
[0092] EMD(X,Y)=∫|F X (x)-F Y (x)|dx (1)
[0093] where F X (x) and F Y (x) represent the CDFs of variables X and Y at point x, respectively.
[0094] Through EMD measurement, the present invention obtains a more geometrically robust similarity measurement method in the construction of the feature memory bank, thereby reducing the influence brought by geometric transformations such as rotation. In industrial defect detection tasks, image features often present complex distribution patterns. EMD can effectively capture local feature changes and global structure changes, thereby enhancing the robustness of the reasoning process and reducing misclassification caused by local differences, improving the accuracy and robustness of the feature library.
[0095] The memory bank sampling strategy focuses on how to efficiently construct a representative and compact feature memory bank. The purpose of memory bank sampling is to select the features that best represent the normal sample distribution, so that the memory bank can not only save storage space but also improve the matching efficiency.
[0096] In the process of constructing the feature memory bank, storing all the extracted features may cause the memory bank to be too large, increasing the computational complexity and significantly slowing down the reasoning process. For this reason, the present invention adopts the core subsampling method to reduce the size of the memory bank. The principle of core subsampling is based on the optimization method, and the goal is to select a small subset S={s1, s2,..., s n} from the original data set X={x1, x2,..., x k}, such that the subset S can retain the key statistical information and structural features of the original dataset. Specifically, core subsampling constructs a small subset by selecting the most representative samples, which can approximately represent the distribution of the entire dataset. To achieve this goal, sample selection is usually carried out by minimizing the distance metric between the original dataset and the selected subset, so as to ensure that the subset can represent the important information in the dataset as much as possible. Core subsampling not only needs to consider the global distribution of data points, but also pay attention to the local relationships between data points in order to better reflect the overall characteristics of the dataset. The optimization problem of core subsampling can be expressed as:
[0097]
[0098] where X is the original dataset, containing n samples. S is the core subset, containing k samples, and k << n. d(x i , s j ) represents the distance between sample x i and subset sample s j .
[0099] In S13, a greedy algorithm is used to solve the optimization problem of core subsampling, including the following steps: Select an initial sample from the dataset as the first element of the core subset; then repeatedly select a sample from the remaining samples to minimize its distance from the current subset to maximize representativeness until the core subset reaches the predetermined size.
[0100] Specifically, in practical applications, since the computational complexity of directly solving the optimization problem in Equation (2) is relatively high, approximate algorithms are usually used to find the core subset. The greedy algorithm is a common and efficient method, which constructs the core subset by gradually selecting the most representative samples. Specifically, the process of the greedy algorithm is as follows: First, select an initial sample from the dataset as the first element of the core subset; then, at each step, select one from the remaining samples to minimize its distance from the current subset to maximize representativeness. This process is repeated until the core subset reaches the predetermined size k.
[0101] Although the computational complexity of EMD is relatively high, the core subsampling method significantly reduces the size of the memory bank, ensuring computational efficiency and maintaining the inference speed. In addition, to reduce the computational complexity, this step only needs to judge whether the image is abnormal, so there is no need to store the image features of each part into the feature memory bank one by one, but the entire image features are stored. This method significantly reduces the storage requirements of the memory bank and simplifies the inference process. In this way, the present invention ensures the inference speed on large-scale datasets.
[0102] The generation of pseudo-labels in S2 involves processing the unlabeled samples in the training set to obtain feature representations through the same feature extraction network, Wide-ResNet50. Subsequently, as Figure 4 shown, perform a nearest neighbor search in the memory bank to calculate the anomaly score between the obtained feature representation and the nearest feature vector stored in the memory bank. Determining the pseudo-labels based on the anomaly score is a key consideration in this process.
[0103] In S2, first, according to the anomaly scores inferred by the model, determine a threshold that maximizes the F1-score through the optimal F1-score selection method, and then normalize this threshold to center the anomaly scores at 0.5. The specific processing steps are as follows:
[0104] (1) Anomaly score calculation: First, obtain a set of anomaly scores through model inference, where each score value represents the degree of anomaly of a specific sample. The higher the score, the greater the probability of the sample being anomalous.
[0105] (2) Calculate precision, recall, and corresponding thresholds: Based on the set of anomaly scores, calculate the precision and recall at different thresholds and record these thresholds. Precision represents the proportion of samples that are actually positive (anomalous samples) among those classified as positive. The calculation formula is as follows:
[0106]
[0107] where True Positives (TP) refers to the actual anomalous samples being correctly classified as anomalous; False Positives (FP) refers to normal samples being misclassified as anomalous. Recall represents the proportion of all actual positive (anomalous samples) that are successfully classified as positive. The formula is as follows:
[0108]
[0109] where False Negatives (FN) refers to anomalous samples being misclassified as normal.
[0110] (3) Calculate the F1-score: The F1-score calculation formula is as follows:
[0111]
[0112] Add a small constant (e.g., 1×10 -10 ) when calculating the F1-score to avoid division by zero errors.
[0113] (4) Select the optimal threshold: Traverse all thresholds to find the threshold T that maximizes the F1-score tmp,and use this threshold as the optimal threshold for determining whether a sample is abnormal.
[0114] (5) Normalization processing: Normalize the anomaly scores using the optimal threshold so that they are centered at 0.5. The specific formula is as follows:
[0115]
[0116] where max_val and min_val respectively refer to the maximum and minimum values in the anomaly score group; scores refer to each anomaly score in the anomaly score group. After obtaining the normalized anomaly scores, the present invention designs two methods to obtain pseudo-labels for subsequent training, such as Figure 4 pseudo-label 1 and pseudo-label 2.
[0117] The first method aims to use high-confidence predictions as pseudo-labels to ensure the accuracy of the training dataset. Specifically, the first method classifies data with high anomaly scores as abnormal and data with low anomaly scores as normal. The pseudo-labels generated in this way correspond to Figure 4 pseudo-label 1 in. Images a and b have relatively high anomaly scores, indicating a relatively high confidence that they are abnormal images. In contrast, the anomaly score of image f is relatively low, indicating a relatively high confidence that it is a normal image. In comparison, the anomaly scores of images c to e are neither high nor low, indicating some doubt about their classification, and these doubtful data are not used in the subsequent weakly supervised algorithm. The second method directly classifies samples with anomaly scores lower than the predefined threshold of 0.5 as normal and samples higher than this threshold as abnormal. The pseudo-labels generated by this method can be observed in Figure 5 pseudo-label 2 in. In this case, the anomaly scores of images a to e are all greater than 0.5, so they are labeled as abnormal, while the anomaly score of image f is less than 0.5, resulting in a normal label.
[0118] Through experimental verification, images with extremely high anomaly scores in the first method often exhibit obvious defect features, such as obvious cracks or missing parts. Therefore, using these highly confident pseudo-labels for training results in the subsequent network mainly learning to distinguish data with obvious defects, and it is difficult to identify and segment subtle defects. This limitation reduces the overall robustness and generalization ability of the model.
[0119] Although Method 2 introduces a certain degree of inaccuracy in the pseudo-labels, it enables the subsequent network to learn a wider range of image features, covering both obvious and subtle defects. By setting the threshold to 0.5, the present invention creates a more balanced and representative pseudo-label training dataset, thereby improving the generalization ability of the network and enabling it to detect various types of anomalies. This method alleviates the problem of the network being overly focused on identifying severe defects, thus enhancing its overall performance and reliability in practical applications.
[0120] After obtaining reliable and sufficient pseudo-labels, the present invention proposes a fine-grained defect segmentation algorithm based on weakly supervised learning as Figure 6 shown, aiming to achieve efficient and fine-grained segmentation results for industrial defects. This segmentation algorithm mainly has the following four key points, namely: training data, weakly supervised network based on class activation maps, semantic segmentation network, and design of the loss function. Among them, the training data includes the data in the original dataset that only contains normal samples, and this part of the data is also used when constructing the memory bank. In addition, the fine-grained defect segmentation algorithm based on pseudo-labels and weakly supervised learning also uses positive and negative samples with pseudo-labels for training. This part of the data is obtained based on the step of obtaining pseudo-labels, has a high confidence level, and contains subtle defects.
[0121] Fine-grained defect segmentation based on weak supervision also requires a feature extraction network to extract features in the image. In the weakly supervised network based on class activation maps, the feature extraction network used in the present invention is the ResNeSt (Residual Networks with Split-Attention) network structure. ResNeSt is an enhanced variant of ResNet, proposed in 2020, aiming to improve the performance of deep neural networks in tasks such as image classification, object detection, and segmentation. The traditional ResNet introduced a residual structure, solving the problem of gradient disappearance in deep networks and enabling the network to better expand to deeper layers, but there is still room for improvement in how to more effectively extract the interaction information between space and channels. Some previous methods (such as SENet) enhanced features in the channels but failed to fully capture the complex relationships between space and channels. The innovation of ResNeSt lies in its "split-attention mechanism", which brings greater flexibility and expressiveness to feature representation.
[0122] In S3, a pre-trained ResNeSt50 is used as the feature extraction network. After the feature map extracted by ResNeSt50, a convolutional layer with a kernel size of 1×1 is added to reduce the number of channels of the feature map to the number of target classes, thereby generating a feature map H1 ∈ R N×C×H×W, where N represents the batch size, H represents the height of the feature map, W represents the width, and C represents the number of classes. Then, the present invention applies the sigmoid activation function to H1, maps the output to the range of [0, 1], and generates the final probability heat map H2. Each pixel value in the heat map H2 represents the probability that the model predicts this position as a defect. In the post-processing stage, the present invention first performs threshold segmentation on the probability heat map H2 to generate a preliminary binary segmentation map. Subsequently, to eliminate noise and connect broken regions, the present invention uses morphological operations (dilation and erosion) to smooth the boundaries of the target region and remove isolated small regions. This post-processing method can effectively enhance the reliability of the prediction, provide a more detailed local response for defect detection, and contribute to the identification and localization of fine-grained defects.
[0123] Different from some methods that freeze the weights of pre-trained models, the present invention does not freeze the weights of ResNeSt50. This choice is based on the observation of the differences between industrial-type datasets and general datasets, that is, the weights of ResNeSt50 cannot fully adapt to these image features with significant distribution differences in the initial state. To help the weights of ResNeSt50 better adapt to the characteristics of the data in this field, the present invention introduces a GAP layer, enabling the network to extract more global feature information. After the GAP layer, the obtained features are used to predict the overall classification result of the image, and the cross-entropy loss is calculated with the pseudo-labels, thereby updating the weights of ResNeSt50. In this way, while the model gradually adapts to the specific feature distribution of the industrial dataset, it can also better capture the association between local and global features.
[0124] In S3, the classical DeepLabV3+ architecture is adopted as the basic framework of the segmentation network, which structurally includes two main modules: an encoder and a decoder. To improve the training efficiency and ensure performance, the present invention selects the ResNet architecture with excellent lightweight characteristics as the core part of the encoder. After the input data is processed by the backbone network, two different levels of feature embeddings are generated: the low-level feature embedding Embedding1 and the high-level feature embedding Embedding2. Among them, Embedding1 is sent to the decoder module, while Embedding2 is passed to the ASPP module for further processing.
[0125] The ASPP module plays a crucial role in the model, responsible for implementing multi-scale feature learning. It processes Embedding2 through five convolutional operations with different dilation rates, enabling effective capture of multi-scale context information. The use of dilated convolutions expands the receptive field of the model, allowing it to obtain more extensive image region information without significantly increasing the computational cost and the number of parameters. This design not only helps fuse information at different scales but also enhances the model's understanding of local features and the global background, making it particularly suitable for the precise segmentation of complex defects in industrial datasets.
[0126] In the decoder part, the output features of the ASPP module are merged with the low-level features Embedding1 and restored to the same resolution as the original input image through progressive upsampling operations, thereby generating accurate segmentation results. This way of fusing low-level and high-level features combines multi-scale context information, significantly improving the model's ability to detect subtle defects and ultimately achieving high-precision segmentation results.
[0127] The main processing flow of the image through the DeepLabv3+ module can be represented by the following equations:
[0128] Embedding1, Embedding2 = ResNet(image) (7)
[0129] Wherein, the image is first processed through the ResNet backbone network to extract the low-level feature embedding Embedding1 and the high-level feature embedding Embedding2. Then, the high-level feature embedding $Embedding2$ passes through the atrous spatial pyramid pooling module to capture multi-scale context information:
[0130] {P1, P2, …, P5} = ASPP(Embedding2) (8)
[0131] The ASPP module processes Embedding2 through five convolutional operations with different dilation rates respectively, and the set of feature maps {P1, P2, …, P5} obtained will be concatenated and then integrated through a convolutional layer:
[0132] Merged = conv(concat({P1, P2, …, P5})) (9)
[0133] Finally, the integrated feature map Merged is concatenated with the low-level feature embedding Embedding1 and input into the decoder module for upsampling to obtain the final segmentation result:
[0134] out = decoder(concat(Merged, Embedding1)) (10)
[0135] The process of constructing the Memory Bank and generating pseudo-labels based on unsupervised learning does not require gradient updates and does not involve any training losses. In this stage, the Memory Bank is used to store sample features for generating pseudo-labels in subsequent stages. These pseudo-labels help the segmentation network obtain a preliminary estimate of whether the image is defective. Therefore, the backpropagation process is not included in constructing the Memory Bank and generating pseudo-labels, so the design of the loss function is not involved.
[0136] In contrast, the loss function of the segmentation network in S3 includes the image-level classification loss L cls and the pixel-level segmentation loss L seg , and the overall training loss is defined as:
[0137] L = αL cls + βL seg (11)
[0138] where α and β are weight factors used to balance the importance of the classification loss and the segmentation loss. The image-level classification loss L cls is used to ensure that the model can correctly distinguish image categories, and the segmentation loss L seg is used to improve the segmentation accuracy of the model at the pixel level;
[0139] First, the image-level classification loss L cls is defined as the cross-entropy loss between the pseudo-class label C class and the model-predicted class C predicted :
[0140] L cls = CE(C class , C predicted ) (12)
[0141] Second, the pixel-level segmentation loss L seg is defined as the cross-entropy loss between the pseudo-segmentation map S pseudo and the model-predicted segmentation map S predicted :
[0142] L seg = CE(S pseudo , S predicted ) (13)
[0143] where the expression of the cross-entropy loss function CE(x, y) is as follows:
[0144]
[0145] In the above equation, (C class ) and (C predicted ) represent the pseudo-class label and the predicted class respectively; Spseudo and S predicted represent the pseudo - segmentation map and the predicted segmentation map. The use of cross - entropy loss helps to quantify the difference between the model's prediction result and the pseudo - label, thus guiding model optimization.
[0146] It is worth noting that by combining the image - level classification loss and the pixel - level segmentation loss, the complementarity of global and local information can be achieved in the model. The image - level classification loss can capture the overall defect category of the image, while the pixel - level segmentation loss focuses on the fine segmentation of the defect area. Therefore, in industrial applications, the overall loss function L = αL cls +βL seg can improve the comprehensive performance of the model in the defect detection task. By adjusting the ratio of the classification and segmentation losses through the weight factors α and β, the balance between the accuracy and robustness of the model is ensured.
[0147] Generally speaking, this training process effectively combines the image - level and pixel - level information. By dynamically adjusting the gradient update, the segmentation network can more accurately identify and segment the defect areas in industrial images.
[0148] Another embodiment is used to illustrate a defect classification and segmentation system based on the combination of unsupervised and weakly - supervised, as Figure 7 shown. The system 700 includes:
[0149] A feature memory bank construction module 710, which is used to use a training data set containing only defect - free samples, extract the feature representation of the defect - free samples through a feature extractor and store it in the feature memory bank;
[0150] A pseudo - label generation module 720, which is used to compare the mixed data set of unlabeled samples (including both defective and defect - free samples) with the features in the feature memory bank to determine the pseudo - classification labels;
[0151] A segmentation network training module 730, which is used to design a weakly - supervised network based on class activation maps for extracting features in the image, a semantic segmentation network and its loss function, and use the training data set of defect - free samples and the pseudo - classification label data to train the semantic segmentation network.
[0152] In addition to the above - mentioned modules, the system 700 may also include other components. However, since these components are not related to the content of the embodiments of the present disclosure, their illustrations and descriptions are omitted here.
[0153] For the other specific working processes of the defect classification and segmentation system 700 based on the combination of unsupervised and weakly - supervised, refer to the description of the embodiments of the defect classification and segmentation method based on the combination of unsupervised and weakly - supervised above, and will not be elaborated here.
[0154] To verify the effectiveness of the method and system of the present invention, the method proposed by the present invention is experimentally verified on four mainstream industrial product defect datasets, including the MVTecAD dataset, the BTAD dataset, the DAGM dataset, and the KSDD2 dataset. In total, the training sets of these datasets contain 6,811 images, among which 3,215 images have image-level labels, and the remaining 3,596 images are unlabeled images. Therefore, compared with traditional weakly supervised defect detection methods, the present invention reduces the labeling cost by approximately 53%.
[0155] The MVTecAD dataset contains categories from 15 industrial scenarios. The number of defective and non-defective images in the dataset is shown in Table 1. The image resolution of this dataset ranges from 700×700 to 1024×1024 pixels. To accelerate the training and inference processes in the experiment and facilitate comparison with other algorithms, all data is resized to 256×256 pixels in the experiment.
[0156] The BTAD dataset (BeanTech Anomaly Detection) is an anomaly detection dataset based on real industrial scenarios. This dataset contains a total of 2,830 images showing anomalies in three industrial products. Each category contains defective and non-defective images, and the specific data volume is shown in Table 1. The image resolution of this dataset is: 1600×1600 pixels for the images of Product 1, 600×600 pixels for Product 2, and 800×800 pixels for Product 3. To speed up the training and inference processes and facilitate comparison with other algorithms, all images are resized to 256×256 pixels in the experiment.
[0157] The DAGM2007 dataset mainly focuses on various defects on textured backgrounds. This dataset contains 10 sub-datasets. To effectively manage the data scale and focus the analysis and experiments, two of these categories (DAGM1 and DAGM2) are randomly selected for research in this experiment. The number of defective and non-defective images in these categories is shown in Table 1. The ground truth labels of this data are represented by ellipse shapes, roughly marking the defective areas, so there is a certain labeling error. In the experiment, all data is resized to 256×256 pixels.
[0158] The KSDD2 dataset is an industrial surface detection dataset, containing 356 defective images and 2,979 non-defective images, and the specific quantities are shown in Table 1. The image resolution in the dataset is approximately 230×630 pixels, but it is not exactly the same. In the experiment, all images are resized to 160×360 pixels.
[0159] Table 1 The number of normal and abnormal samples in the MVTecAD, BTAD, DAGM, and KSDD2 datasets
[0160]
[0161] In the experiments of the present invention, the dataset was divided as follows: Training set 1 only contained normal samples, while training set 2 contained both normal and abnormal samples. The test set and the validation set both contained abnormal and normal samples. For the specific composition, please refer to Table 2. The division of the dataset followed a ratio of 7:1.5:1.5 for training, validation, and testing respectively.
[0162] This division method enabled the present invention to use the dataset containing only normal samples to construct a memory bank during the training process, and use the mixed dataset (containing abnormal and normal samples) to generate pseudo-labels for training. The test set and the validation set were used to evaluate the performance of the model in actual applications, ensuring the effectiveness and reliability of the algorithm when dealing with real-world data.
[0163] All algorithm experiments were conducted on the Ubuntu 18.04 operating system. The Python programming language was used during the development process, and the deep learning model adopted the PyTorch framework (version 1.12.1). The hardware platform used in the experiments was an Nvidia Tesla V100 graphics card with 32GB of video memory, which could efficiently support the training and inference processes of large-scale models. When constructing the memory bank, Wide-ResNet50-RACM was used as the pre-trained encoder. The training batch size and the test batch size were set to 1 and 8 respectively to adapt to the training process and hardware resources of the present invention.
[0164] During the training process of the segmentation network of the present invention, the ResNeSt50 network pre-trained on the ImageNet 21 categories was used as the backbone network, and its parameters were initialized to accelerate the convergence of the model. Although the defect detection task of the present invention only involved two categories (background and defect), through this initialization method, the network could more effectively identify the background and better handle the defect category.
[0165] In addition, the Adam optimizer was selected as the optimizer, the learning rate was set to 0.001, and 50 epochs were used for training during the training process. Through the reasonable configuration of these hyperparameters, the present invention could obtain good training results in a relatively short time, ensuring the performance of the model in the actual defect detection task.
[0166] Table 2 Composition ratio of the dataset
[0167]
[0168] Comparison of experimental results
[0169] The present invention is compared with state-of-the-art algorithms in the field of unsupervised anomaly detection, including Padim, Draem, Reverse Distillation, Cflow, and Cfa, etc. These algorithms have demonstrated remarkable performance on the public leaderboard of Papers With Code and their codes have been open-sourced. The specific comparison results are shown in Table 3.
[0170] Although Diffusion-AD also performs excellently in the defect detection task, its training and inference times are too long, and the experimental time cost for a single dataset can reach several days. Therefore, Diffusion-AD is not considered in the comparison of state-of-the-art algorithms in the present invention.
[0171] Table 3 Comparison results of the algorithm of the present invention with other SOTA algorithms. The evaluation metrics include AP, Accuracy, F1-score, and MIoU
[0172]
[0173] First, the MVTecAD dataset is analyzed. To examine the generalization ability of the algorithm, the present invention conducts an average analysis of the performance of all categories in MVTecAD. The results show that among the SOTA methods, Reverse Distillation is the best-performing algorithm. However, the Semi-Patchcore method of the present invention has shown significant improvements in multiple evaluation metrics: the AP has increased by 0.36%, the Accuracy has increased by 0.17%, the F1-score has increased by 0.59%, and the MIoU has increased by 3.9%.
[0174] The visualization results of the present invention on the MVTecAD dataset are shown in Figure 8 . These results clearly demonstrate the powerful ability of the algorithm to detect defects in various categories, and through the visualization comparison with the ground truth labels, it further verifies the robustness and accuracy of the present invention in different defect categories.
[0175] By averaging the results of the three sub-datasets BTAD1, BTAD2, and BTAD3, the overall performance analysis of the BTAD dataset is obtained. The results show that among the state-of-the-art methods, Reverse Distillation and Padim perform the best. However, the algorithm of the present invention exceeds Reverse Distillation and Padim respectively in terms of the AP (Average Precision) metric, with an improvement of 0.7% and 3.2%. More significantly, in terms of the MIoU (Mean Intersection over Union) metric, the Semi-Patchcore algorithm improves by 5.6% and 9.9% compared to Reverse Distillation and Padim respectively, demonstrating significant performance improvement. In addition, Figure 9 The visualization results of the present invention on the BTAD dataset are shown.
[0176] The comparison results of the present invention with multiple SOTA methods on the DAGM and KSDD2 datasets can be seen in Table 3. It can be seen from the table that compared with other SOTA algorithms, the present invention shows significant improvement in the AP and F1-score metrics on most datasets.
[0177] On the DAGM1 and DAGM2 datasets, the algorithm of the present invention achieved scores close to 1 or 1 in terms of AP, Accuracy, and F1-score, and the MIoU metrics reached 59% and 53.1% respectively, significantly outperforming the existing SOTA algorithms. On the KSDD2 dataset, the algorithm of the present invention surpassed the SOTA method Reverse Distillation, with improvements of 2%, 4.5%, and 16% in terms of AP, Accuracy, and MIoU respectively. Especially in terms of the MIoU metric, the algorithm of the present invention showed significant improvement, fully verifying its powerful ability in industrial defect detection.
[0178] These results further highlight the effectiveness of the present invention in multiple industrial defect detection tasks, demonstrating its superior performance on different datasets and evaluation metrics, and having significant advantages compared to the existing SOTA methods.
[0179] Another important dimension that the present invention focuses on is the inference time. Table 4 shows the comparative analysis of the inference time and performance metrics of various anomaly detection methods. The experiments of the present invention were carried out on a V100 GPU with 32GB video memory, and the average inference time and evaluation metrics of the BTAD, MVTecAD, DAGM, and KSDD2 datasets were calculated.
[0180] Table 4 Comparison of inference time and performance metrics of different algorithms
[0181]
[0182] The results show that the inference time of the present invention on each sample is 0.061 seconds, which is the fastest among all methods. At the same time, it has the highest average precision (0.980) and Accuracy (0.946), demonstrating the best balance between efficiency and performance. In contrast, the inference times of Patchcore, Padim, and Cfa are similar, approximately 0.095 to 0.097 seconds. Among these methods, Patchcore performs the best, although its AP and accuracy are slightly lower than those of the method of the present invention. Draem and Reverse Distillation have slightly longer inference times but still remain competitive in terms of performance. Cflow has the slowest inference time, reaching 0.325 seconds, and its performance metrics are relatively low. Overall, the present invention has shown excellent results in both inference speed and performance metrics.
[0183] The present invention adopts a semi-supervised anomaly detection network called Semi-Patchcore, which is trained only on defect-free good samples and a large number of unlabeled samples. This network is based on the Patchcore architecture and integrates the concepts of weak supervision and unsupervised learning into a semi-supervised anomaly detection framework. Specifically, throughout the training process, unsupervised and weak supervision algorithms are jointly integrated for model learning and optimization.
[0184] To highlight the improvement effect of jointly applying unsupervised and weak supervision to the overall algorithm, we conducted ablation experiments on this process. The specific results are as Figure 10 shown. To demonstrate the generalization ability of the Semi-Patchcore algorithm, the present invention conducted a comprehensive analysis on multiple datasets. When combining unsupervised algorithms with weak supervision in joint training compared to using only unsupervised algorithms, the present invention observed improvements of 6.3%, 13.4%, 17.3%, 1.1%, 12.7%, and 8.9% respectively in key evaluation metrics.
[0185] The experimental results show that simply using unsupervised algorithms has already shown relatively excellent performance in defect detection and localization. However, introducing pseudo-labels for weak supervision training after unsupervised algorithms can further improve the detection effect. The present invention takes advantage of the ability of unsupervised learning to extract effective features from data without labeled samples and refines these features through the additional context information provided by pseudo-labels. It is also worth noting that although the integration of unsupervised and weak supervision algorithms improves performance, it also increases the overall complexity of the network, resulting in an increase in training time, which needs to be considered in practical applications.
[0186] Compared with the traditional Euclidean distance, introducing the Earth Mover's Distance (EMD) into the feature metric for defect detection significantly improves the performance. As shown in Table 5, EMD outperforms the Euclidean distance on multiple datasets, with varying degrees of improvement in the main evaluation metrics: the Average Precision (AP) increases by 1.9%, the accuracy improves by 3.1%, the F1-score rises by 2.3%, and the Mean Intersection over Union (MIoU) increases by 4.7%. This indicates that EMD can more effectively measure the distance between defect features, making the detection model more sensitive to subtle defects and thus enhancing its performance in complex scenarios.
[0187] Table 5 Comparison of Inference Time and Performance Metrics of Different Algorithms
[0188]
[0189] As can be seen from Table 5, EMD outperforms the Euclidean distance in all metrics, especially showing significant improvements in MIoU and F1-score. This improvement is particularly important in practical applications because EMD more flexibly measures the differences between unevenly distributed features, enabling the detection model to perform better when dealing with complex defect patterns. The advantage of EMD for the defect detection model lies in its precise matching ability for local features, which is more suitable for detecting various common defect forms in industrial manufacturing. Therefore, using EMD as the feature metric not only improves the detection accuracy but also enhances the robustness and adaptability of the model in real production environments.
[0190] Analysis of the Influence of Pseudo-Labels
[0191] The present invention also conducted comparative experiments using weak supervision algorithms on the same dataset to examine the influence of pseudo-labels in the algorithms. Specifically, when calculating the loss $L_{cls}$ of the training segmentation network, the pseudo-labels were attempted to be replaced with real labels for network training and inference, while keeping other structures of the network unchanged. The experimental settings are shown in Figure 11 。
[0192] By observing Figure 12 the experimental results in, the present invention found that on some datasets (such as Bottle and DAGM2), the algorithms using pseudo-labels perform better in terms of performance than those using real labels. However, in other datasets, the results of training with pseudo-labels are slightly lower than those of real labels. The reason for this phenomenon can be attributed to the inherent label inaccuracy problem in weak supervision learning. Pseudo-labels provide a rough localization of defects to a certain extent, but due to their lower precision compared to real labels, the performance is slightly lower in some more complex scenarios.
[0193] Based on the technical solutions provided in the above embodiments, a defect classification and segmentation method and system based on the combination of unsupervised and weakly supervised learning are used for industrial defect detection. The method is called Semi-Patchcore. Innovatively, it only uses defect-free samples and a large number of unlabeled samples for training, solving the limitations of traditional supervised learning methods in the case of high labeling costs. Through a carefully designed two-stage training process, Semi-Patchcore achieves effective detection and localization of unseen defect features. This method not only has strong generalization ability, but also can accurately identify various types of industrial defects without relying on a large amount of labeled data. The specific advantages are as follows:
[0194] (1) The present invention discloses a defect classification and segmentation method Semi-Patchcore based on the combination of unsupervised and weakly supervised learning. This method combines the advantages of weakly supervised and unsupervised learning algorithms. By relying only on defect-free samples and a large number of unlabeled samples for training, it breaks through the dependence on a large amount of labeled data in traditional methods and realizes efficient defect detection and localization.
[0195] (2) The present invention introduces the Earth Mover's Distance (EMD), providing a more accurate method for measuring the distance between different features in industrial surface images. This method can better adapt to complex defect detection scenarios and improve the accuracy of pseudo-label generation.
[0196] (3) By only using defect-free samples and a large number of unlabeled data for model training, the present invention significantly reduces the data labeling cost, reducing the labeling cost by about 53%. At the same time, it outperforms previous semi-supervised and unsupervised methods in terms of defect detection performance, demonstrating strong practical value.
[0197] (4) Experiments are carried out on four datasets, namely MVTec AD, BTAD, DAGM, and KSDD2. The metrics between classical unsupervised algorithms, the current state-of-the-art unsupervised algorithms, and the Semi-Patchcore algorithm proposed by the present invention are compared, and visualization pictures are given, which can prove that the Semi-Patchcore model proposed by the present invention has good generalization.
[0198] In this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a step or method including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such step or method.
[0199] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, which should all be regarded as falling within the protection scope of the present invention.
Claims
1. A defect classification and segmentation method based on the combination of unsupervised and weakly supervised methods, characterized in that: The method comprises the following steps: Using a training data set containing only defect-free samples, extracting feature representations of the defect-free samples through a feature extractor and storing them in a feature memory bank; Compare the unlabeled mixed dataset including defective and non-defective samples with the features in the feature memory to determine the pseudo classification labels; A weakly supervised network based on category activation maps is designed to extract features from images, a semantic segmentation network and its loss function, and the semantic segmentation network is trained using a training dataset of defect-free samples and pseudo-classification label data.
2. The defect classification and segmentation method based on the combination of unsupervised and weakly supervised according to claim 1, characterized in that: The feature memory storage process specifically includes the following steps: Feature extraction, extracting multi-level semantic and structural information from input samples; Feature measurement, introducing distance measurement method to evaluate the similarity between different samples; Memory library sampling, by selecting representative samples and storing them in the feature memory library.
3. The defect classification and segmentation method based on the combination of unsupervised and weakly supervised according to claim 1, characterized in that: The features extracted from the third and fourth layers of the Wide-ResNet50 network are used and stored in a feature memory, wherein the features contain local structural information and low-level patterns of the samples, as well as abstract feature information.
4. The defect classification and segmentation method based on the combination of unsupervised and weakly supervised according to claim 2, characterized in that: The distance measurement method uses the earth moving distance for feature measurement, including the following steps: Calculate the histogram for each feature to obtain its probability distribution; Normalize the histogram to obtain the cumulative distribution function; The Earth Mover Distance (EMD) is the sum of the absolute values of the differences between the cumulative distribution functions of two variables X and Y: EMD(X,Y)=∫|F X (x)-F Y (x)|dx where F X (x) and F Y (x) represents the cumulative distribution function of variables X and Y at point x respectively.
5. The defect classification and segmentation method based on the combination of unsupervised and weakly supervised according to claim 2, characterized in that: The core subsampling method is used to select representative samples and store them in the feature memory. The optimization problem of core subsampling can be expressed as: Where X is the original data set, containing n samples, S is the core subset, containing k samples, and k<<n, d(x i ,s j ) represents the sample x i and subset samples s j The distance between The greedy algorithm is used to solve the optimization problem of core subsampling, which includes the following steps: An initial sample is selected from the dataset as the first element of the core subset; then a sample is repeatedly selected from the remaining samples to minimize the distance to the current subset to maximize representativeness until the core subset reaches a predetermined size.
6. The defect classification and segmentation method based on the combination of unsupervised and weakly supervised according to claim 1, characterized in that: The unlabeled mixed dataset including defective and non-defective samples is compared with the features in the feature memory to determine the pseudo classification labels, including: Process each unlabeled sample in the unlabeled mixed data set to obtain feature representation; Perform a nearest neighbor search in the feature memory and calculate an anomaly score between the obtained feature representation and the nearest feature vector stored in the memory; Based on the anomaly score group, calculate the precision and recall under different thresholds and record these thresholds; The F1-score is calculated based on precision and recall, and the expression is: Where Precision represents accuracy and Recall represents recall rate; Traverse all thresholds and find the threshold T that maximizes the F1-score tmp , and use this threshold as the optimal threshold to determine whether the sample is abnormal; The anomaly score is normalized using the optimal threshold to obtain the normalized anomaly score. The specific formula is as follows: max_val and min_val represent the maximum and minimum values in the anomaly score array, respectively, and scores represents the anomaly score in the anomaly score array; Determine the pseudo-classification label of the sample based on the normalized anomaly score.
7. The defect classification and segmentation method based on the combination of unsupervised and weakly supervised according to claim 1, characterized in that: The weakly supervised network based on the category activation map uses the pre-trained ResNeSt50 as the feature extraction network, and adds a convolution layer with a kernel size of 1×1 after the feature map extracted by ResNeSt50 to reduce the number of channels of the feature map to the number of target categories, thereby generating the target feature map; Apply the sigmoid activation function on the target feature map to generate a probability heat map H2. Each pixel value in the heat map H2 represents the probability of predicting that the position is a defect. Perform threshold segmentation on the probability heat map H2 to generate a preliminary binary segmentation map; For the binary segmentation map, morphological operations are used to smooth the boundaries of the target area and remove isolated small areas.
8. The defect classification and segmentation method based on the combination of unsupervised and weakly supervised according to claim 1, characterized in that: The DeepLabV3+ architecture is used as the basic framework of the segmentation network, including the encoder and decoder. The core part of the encoder is the ResNet architecture. After the input data is processed by the backbone network, two different levels of feature embedding are generated: Low-level feature embedding Embedding1 and high-level feature embedding Embedding2, where Embedding1 is fed into the decoder and Embedding2 is passed to the ASPP module for multi-scale feature learning; The ASPP module processes Embedding2 through five convolution operations with different dilation rates to capture multi-scale contextual information; In the decoder part, the output features of the ASPP module are merged with the low-level features Embedding1 and restored to the same resolution as the original input image through a step-by-step upsampling operation to generate the segmentation result.
9. The defect classification and segmentation method based on the combination of unsupervised and weakly supervised according to claim 1, characterized in that: The loss function of the segmentation network includes the image-level classification loss L cls and pixel-level segmentation loss L seg , the overall training loss is defined as: L = αL cls +βL seg , where α and β are weight factors used to balance the importance of classification loss and segmentation loss, and the image-level classification loss L cls Used to ensure that the model can correctly distinguish image categories, the segmentation loss L seg Used to improve the segmentation accuracy of the model at the pixel level; Image-level classification loss L cls Defined as the pseudo class label C class And the model predicts category C predicted The cross entropy loss between: L cls =CE(C class ,C predicted ); pixel-level segmentation loss L seg Defined as a pseudo segmentation map S pseudo Segmentation map S predicted by the model predicted The cross entropy loss between: L seg =CE(S pseudo ,S predicted ).
10. A defect classification and segmentation system based on the combination of unsupervised and weakly supervised methods, characterized in that: The system comprises: Constructing a feature memory module, which is used to use a training data set containing only defect-free samples, extract feature representations of the defect-free samples through a feature extractor, and store the feature representations in a feature memory; A pseudo label generation module is used to compare the unlabeled mixed data set including defective and non-defective samples with the features in the feature memory library to determine the pseudo classification labels; The segmentation network training module is used to design a weakly supervised network based on the category activation map to extract features in the image, the semantic segmentation network and its loss function, and train the semantic segmentation network using a training dataset of defect-free samples and pseudo-classification label data.
Citation Information
Patent Citations
Knowledge distillation-based unsupervised image defect detection method and apparatus, and medium
CN115187525A
AOI defect detection method based on deep learning network DenseNet
CN116542929A
Industrial image defect detection method based on multi-head unbalanced semi-supervised network
CN116630696A
Product appearance defect detection method based on unsupervised anomaly detection
CN116912173A
Unsupervised chip defect detection method and device
CN117237309A
Cited By
Defect detection model training method, defect detection method, device and equipment
CN120510154A
Defect detection model training method, defect detection method, device and equipment
CN120510154B
Intelligent manufacturing defect automatic detection and classification method based on machine vision
CN120580226A
Automatic detection and classification method of intelligent manufacturing defects based on machine vision
CN120580226B
Method for detecting surface defects of few-sample inductance core based on model interaction
CN121504928A