Double-branch deep learning method for sealing ring defect detection
The dual-branch deep learning method enhances seal defect detection by integrating supervised and unsupervised branches to accurately identify defects in seals, overcoming environmental interference and computational inefficiencies.
Patent Information
- Application Number
- CN202510380265.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-15
AI Technical Summary
The prior art is difficult to effectively detect common and uncommon defects of sealing rings, especially in complex environments, which are prone to false detection. Traditional methods have good effect on edge defects but are difficult to detect non-edge defects. Unsupervised methods are sensitive to dust impurities, resulting in false detection.
Design a two-branch deep learning method, including supervised and unsupervised branches, leveraging regional coordinate attention and improved standardized flow modules, combined with sealing ring characteristics, improve detection accuracy and reduce false detection.
It improves the accuracy and generalization of seal ring defect detection, reduces the amount of calculation, avoids false detection, and adapts to the detection needs in complex environments.
Smart Images

Figure CN120318168A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision and industrial inspection, and specifically relates to a dual-branch deep learning method for seal ring defect detection. Background Art
[0002] At present, the defect detection of sealing rings mainly relies on traditional image processing methods, object detection-based methods, or image segmentation-based methods. Traditional image processing methods have achieved good results for defects such as burrs and incorrect sizes of sealing rings (e.g., Haibing H, Zheng X, Yin J, et al. Research on O-ring dimension measurement algorithm based on cubic spline interpolation[J]. Applied Sciences, 2021, 11(8): 3716.). However, for defects at non-edge locations on the sealing ring, traditional image processing algorithms are difficult to handle. Object detection-based methods and image segmentation-based methods can achieve good results in detecting common defects of sealing rings with the drive of a certain amount of data. However, for rare defects, especially large defects with significant variations, it is often difficult to obtain good results. Recently emerging unsupervised defect detection algorithms, such as PatchCore(
[38] Roth K, Pemula L, Zepeda J, et al. Towards total recall in industrial anomaly detection[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2022: 14318-14328.), ReverseDistillation(
[34] Deng H, Li X. Anomaly detection via reverse distillation from one-class embedding[C] / / Proceedings of the IEEE / CVF conference on computer vision and pattern recognition.), etc., although achieving good results on public datasets of various industrial products, in the complex environment of actual production and transportation, the inspected sealing rings are often easily contaminated with impurities such as dust, debris, and wool threads. These will cause the unsupervised defect detection algorithm to produce false detections, identifying the contaminated impurities as defects, resulting in a low detection accuracy. For the above reasons, the present invention designs a more suitable dual-branch defect detection method for the defect detection of sealing rings, which not only has a better detection effect on common defects but also has a certain detection effect on uncommon defects without affecting the detection accuracy. Summary of the Invention
[0003] The object of the present invention is to address the deficiencies in the prior art and propose a dual-branch deep learning method for seal ring defect detection. Firstly, the supervised branch proposed by this method improves the detection accuracy of the seal ring defect detection method based on object detection. At the same time, the unsupervised branch proposed by this method avoids the problem of false detection of pure unsupervised defect detection methods in the face of dust and impurities, and can also alleviate the problem that pure supervised defect detection methods are difficult to learn and detect large defects with large variations, improving the detection accuracy while only slightly increasing the computational complexity.
[0004] The technical solution for achieving the object of the present invention is as follows:
[0005] A dual-branch deep learning method for seal ring defect detection, comprising the following steps:
[0006] 1) Collect and preprocess seal ring images: Collect seal ring images. Since the sizes of the input seal ring images are different, first segment the seal ring from the background by threshold segmentation for the seal ring images, then eliminate dust and impurities in the background through closing operation. After obtaining the mask of the seal ring, calculate the minimum bounding rectangle of the mask area and perform cutting. Finally, scale the cut seal ring images to the same size;
[0007] 2) Use the images collected in step 1) to prepare the datasets for the dual-branch network: This dataset includes a supervised branch dataset and an unsupervised branch dataset;
[0008] 3) Construct a dual-branch network. The entire dual-branch network consists of a backbone network, an RCA-BiFPN-P2 supervised branch, and an unsupervised branch:
[0009] 3.1) First, construct the backbone network of the dual-branch defect detection network. This backbone network is a CSP network structure consistent with YOLOV8S;
[0010] 3.2) The supervised branch takes the outputs of the first 3 C2f layers and the last SPPF network layer of the backbone network as the input of the supervised branch. The supervised branch consists of a Conv layer, a C2f layer, a regional coordinate attention layer, a Merge layer, and a detection head layer; among them, the Merge layer adopts the feature fusion method of BiFPN; the detection head layer adds a small object detection head compared to YOLOV8, which greatly improves the detection accuracy for seal rings with small wire diameters;
[0011] 3.3) Construct the unsupervised branch: The unsupervised branch takes the output of the last layer of the backbone network as the input. The unsupervised branch consists of a global average pooling layer and eight identical normalizing flow modules, and sequentially connects the global average pooling layer and the normalizing flow modules before and after to process the input features;
[0012] 4) Training the supervised branch of the network: Combine the backbone network with the supervised branch constructed in step 3.2), and use the supervised branch dataset in step 2) to train, test, and validate the combined network. The training objective is to enable the supervised branch to fit the bounding box where the defect is located. After training, obtain the weights of the backbone network and the supervised branch.
[0013] 5) Training the unsupervised branch of the network:
[0014] 5.1) Freeze the weights of the entire backbone network, combine the backbone network with the unsupervised branch, and use the unsupervised dataset in step 2) to train the combined unsupervised branch. The training objective is to make the vector finally output by the unsupervised branch approach a zero-mean normal distribution. After training, assuming the vector output by the unsupervised branch is Z, for whether the input image has a defect, a defect score can be calculated. The larger the value, the more likely it is a defect. The formula is:
[0015] score = mean(Z 2 ) ;
[0016] 5.2) Use the validation set of the unsupervised dataset to calculate the defect scores of all images to select a suitable score threshold ut for determining whether an image has a defect. The unsupervised branch functions as an auxiliary for detecting large defects. Therefore, the rule for selecting the threshold is to select the threshold with a high recall rate under the condition of ensuring precision.
[0017] 6) Integrate the network structures of the backbone network, the unsupervised branch, and the supervised branch, as well as the network models and weights obtained from training in steps 4) and 5). Freeze all weights and turn off the model gradient calculation, then enter the detection phase.
[0018] 7) For the image of the to-be-detected sealing ring taken, perform the same operations as in step 1) for preprocessing. After locating the position of the sealing ring, cut the image, normalize it to between 0 and 1, and input it into the backbone network for feature extraction.
[0019] 8) The features output by the backbone network are respectively input into the supervised branch and the unsupervised branch for simultaneous calculation. As long as one branch considers it defective, it is considered defective.
[0020] 8.1) After the supervised branch network finishes the calculation, filter out unnecessary prediction boxes through non-maximum suppression and set a confidence threshold, and obtain the prediction result of the bounding box for the defect. If the prediction result is empty, it means there is no defect; if it is not empty, it means there is a defect.
[0021] 8.2) For the calculation of the unsupervised branch network, compare it with the threshold ut obtained during training. If it is greater than the threshold ut, it is considered to have an obvious large defect.
[0022] In step 2), the dataset of the supervised branch consists of defective images and defect-free images that are easily misdetected. The specific positions of the defects are marked using image annotation software, and the marking method is bounding box marking. After marking, the training, test, and validation datasets are divided in a ratio of 8:1:1. For the dataset of the unsupervised branch, it is only necessary to divide the whole image into defective and defect-free ones. The training set only includes defect-free images, and the validation set includes some defect-free images and defective images.
[0023] In step 3.2), the Region Coordiante Attention (RCA) layer has two branches. For the input feature I1 with the shape of [C, H, W], one branch performs global average pooling on the I1 feature, then performs grouped convolution with a convolution kernel size of 1 and 32 groups on the obtained feature, and finally uses the sigmoid activation function to obtain the attention matrix A1. The other branch performs adaptive average pooling operations on I1 in the H direction and W direction of the feature map respectively to obtain feature vectors with sizes of [C, H, N] and [C, N, W]. The two feature vectors are transposed to align the dimensions and concatenated to obtain [C, H+W, N], then input into a 1x1 grouped convolution, and finally the vectors with sizes of [C, H, N] and [C, N, W] are re-split. Then, two features with sizes of [C, H, W] are obtained through nearest neighbor interpolation. The sigmoid function is applied to the two features with sizes of [C, H, W] respectively, and finally the two are multiplied to obtain the attention matrix A2. Finally, (A1+A2) / 2 is used to obtain the final attention map, which is multiplied by I1 and output.
[0024] In step 3.3), global average pooling aggregates the input features into a size of 1x1 in the spatial dimension, and the channel dimension remains unchanged. The normalizing flow module consists of modules with reversible calculations. The calculation of a single normalizing flow module is expressed as:
[0025] X`=X@W1
[0026] X`1,X2`=Split(X`)
[0027] X``1=X`1
[0028] Z1,Z2=Split(MLP(X`1))
[0029]
[0030] X```2=X``2
[0031] Z3,Z4=Split(MLP(X``1))
[0032]
[0033] Y = Concat(X```1, X```2)
[0034] Where X is the input, Y is the output, X`, X`1, X``1, X```1, X`2, X``2, X```2, Z1, Z2, Z3, Z4 all represent intermediate variables, @ represents matrix multiplication, W represents learnable weights, Split represents splitting the input into two parts of the same size in the channel dimension, and MLP represents the stacking of fully connected layers, DropOut layers, and Relu activation function layers, stacked three times in total. Except that the input channel number of the first input fully connected layer is half of X, the input and output channel numbers of the remaining fully connected layers are equal to X.
[0035] First, in the supervised branch of RCA - BiFPN - P2, this technical solution uses a region coordinate attention designed for the sealing ring. Since a large characteristic of the processed sealing ring image is that it is hollow, when using general attention with global average pooling or similar operations, the generated mean value is often diluted by the middle blank area. However, the region coordinate attention operation designed in this technical solution is equivalent to dividing the entire image into regions of size NxN to do attention respectively. Without significantly increasing the computational cost, the generated attention map can better focus on the defects on the sealing ring. The introduction of the added small target detection head and the BiFPN Merge layer improves the detection effect of the sealing ring with a small wire diameter while not significantly increasing the computational cost.
[0036] Secondly, the unsupervised branch used in this technical solution is different from the existing method that uses the backbone network pre - trained on ImageNet for feature extraction. Instead, it re - uses the backbone network trained by the supervised branch of this technical solution. One is that the features it extracts are more in line with the sealing ring defect detection task, enabling the unsupervised branch to better distinguish whether it is a normal image. The other is that it avoids the repeated feature extraction process, effectively reducing the computational cost.
[0037] Finally, many current unsupervised defect detection techniques focus on pixel-level unsupervised defect detection, allowing unsupervised methods to complete all defect detection tasks independently. This is feasible in environments without complex interference factors such as laboratory environments. However, most sealing rings are very small, with a wire diameter as small as 0.3 mm. Even if very small dust adheres to them in a non-dust-free environment, it will appear large after being photographed by a high-resolution camera and be misidentified as a defect. Therefore, the unsupervised branch designed in this technical solution is transformed into modeling the sealing ring images at the image level to assist in identifying large defects. Specifically, for the features extracted by the backbone network, they are aggregated into 1x1 features through global average pooling and the characteristics of normal images are memorized through an improved normalizing flow network. Since the characteristics of the entire sealing ring are memorized, the unsupervised branch of this technical solution can better resist the interference of dust and wool. At the same time, since the defects of the sealing ring are mainly small defects, while the number of large defects is small, their area is large and the feature changes are also large, which are difficult to learn for the supervised branch. However, the image-level unsupervised modeling in this technical solution is more sensitive to large defects, which exactly makes up for this point.
[0038] This method benefits from the improvement of the proposed supervised branch, and the training convergence speed of the overall network is faster, the training time is reduced, and the network can better find the commonalities among various defects, improving the generalization of the model and thus providing the detection accuracy. At the same time, the proposed unsupervised branch performs an overall modeling of the sealing ring images with only a slight increase in computational complexity. It not only makes up for the problem that pure supervised detection methods are difficult to train and detect large defects, but also avoids the false detection problem caused by pure unsupervised methods and the redundant calculations caused by repeated feature extraction. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 Schematic diagram of regional coordinate attention;
[0040] Figure 2 Specific network structure of the supervised branch;
[0041] Figure 3 Structure of a single normalizing flow module of the unsupervised branch;
[0042] Figure 4 Regarding the detection effects of the unsupervised branch and the supervised branch for large defects, the left side of the dashed line is the defect image, the right side is the normal image, the score output by the unsupervised branch is below the image, and the square box in the defect image is the detection effect of the supervised branch;
[0043] Figure 5 Scatter plot of the defect area and the corresponding defect score of the image detected by the unsupervised branch;
[0044] Figure 6 Examples of four types of sealing rings A, B, C, and D collected in the embodiment;
[0045] Figure 7 Ablation experiment result diagram of the supervised branch compared with the original branch of YOLOV8. Specific implementation manner
[0046] The content of the present invention will be further described below in conjunction with the accompanying drawings and embodiments, but it is not a limitation of the present invention.
[0047] Embodiment:
[0048] A two-branch deep learning method for seal ring defect detection includes the following steps:
[0049] 1) Collect and preprocess seal ring images: Collect seal ring images, segment the seal ring from the background through threshold segmentation. In this example, the threshold is set to 230, and those greater than 230 are considered as the background. Then, eliminate dust and miscellaneous points in the background through closing operation. After obtaining the mask of the seal ring, calculate the minimum bounding rectangle of the mask area and perform cutting. Finally, scale the cut seal ring image to a size of 1280x1280.
[0050] 2) Prepare the dataset for the two-branch network using the images collected in step 1): This dataset includes a supervised branch dataset and an unsupervised branch dataset; for the supervised branch dataset, 2431 defective pictures are collected, which are divided into four categories A, B, C, and D according to the size of the seal ring (optional, not necessary), as Figure 6 shown. Use the Labelme software for annotation. After annotation, divide the training, test, and validation datasets according to a ratio of 8:1:1; while for the unsupervised branch, 732 defect-free pictures are collected, among which 368 are used as the training set and the others as the validation set, and 673 defective pictures.
[0051] 3) Construct a two-branch defect detection network. The entire two-branch network consists of a backbone network, an RCA-BiFPN-P2 supervised branch, and an unsupervised branch:
[0052] 3.1) First, construct the backbone network of the two-branch defect detection network. This backbone network is a CSP network structure consistent with YOLOV8S.
[0053] 3.2) The supervised branch takes the outputs of the first 3 C2f layers and the last SPPF network layer of the backbone network as the input of the supervised branch. As Figure 2 shown, the supervised branch consists of a Conv layer, a C2f layer, a regional coordinate attention layer, a Merge layer, and a detection head layer; among them, the regional coordinate attention layer has two branches, as Figure 1As shown, for the input feature I1 with the shape of [C, H, W], in one branch, global average pooling is performed on the features of A1, then grouped convolution with a kernel size of 1 and 32 groups is performed on the obtained features, and finally, the attention matrix A1 is obtained using the sigmoid activation function; in the other branch, adaptive average pooling operations are performed on I1 in the H direction and the W direction respectively to obtain feature vectors of sizes [C, H, N] and [C, N, W]. The two feature vectors are transposed to align the dimensions and concatenated to obtain [C, H + W, N], then input into a 1x1 grouped convolution. Finally, the vectors are re-split into sizes [C, H, N] and [C, N, W], and two features of size [C, H, W] are obtained through nearest neighbor interpolation. The sigmoid function is applied to each of the two features of size [C, H, W], and finally, the two are multiplied to obtain the attention matrix A2; finally, the final attention map is obtained as (A1 + A2) / 2, multiplied by I1 and output; the Merge layer adopts the feature fusion method of BiFPN; in the detection head layer, a small object detection head is additionally added compared to YOLOV8; in this example, N for the regional coordinate attention in the supervised branch is set to 5;
[0054] 3.3) Construct the unsupervised branch: The unsupervised branch takes the output of the last layer of the backbone network as the input. The unsupervised branch consists of global average pooling and eight identical normalizing flow modules. The structure of a single normalizing flow module is as Figure 3 shown, and the global average pooling layer and the normalizing flow modules are connected in sequence to process the input features;
[0055] 4) Train the supervised branch of the network: Use the collected supervised seal ring defect detection dataset to train, test, and validate the supervised branch composed of the backbone network and the supervised branch network. The training objective is to enable the supervised branch to fit the bounding box where the defect is located; after training, the weights of the backbone network and the supervised branch are obtained;
[0056] 5) Freeze the weights of the entire backbone network, combine the backbone network with the unsupervised branch, and use the unsupervised dataset to train the combined unsupervised branch. The training objective is to make the vector finally output by the unsupervised branch approach a normal distribution with a zero mean. For the convenience of showing the threshold selection, the scores of the unsupervised branch and their actual defect areas are plotted as a scatter plot here, as Figure 5 shown, Figure 5 In [reference], the maximum anomaly score stage is set to 10 for easy observation. If a high precision is to be maintained, the threshold ut can be taken as 1.9 in this example. If some misdetections can be tolerated, the threshold can be further set to a lower value;
[0057] 6) Integrate the network structures of the backbone network, the unsupervised branch, and the supervised branch, as well as the network models and weights obtained from training in steps 4) and 5), freeze all weights, and turn off the model gradient calculation, then enter the detection process;
[0058] 7) For the images of the sealing rings taken, perform the same operations as in the preprocessing of step 1). After locating the position of the sealing ring, cut the image, normalize the values of the image to between 0 and 1, and input it into the backbone network for feature extraction;
[0059] 8) The features output by the backbone network are respectively input into the supervised branch and the unsupervised branch for simultaneous calculation. As long as one branch considers it defective, it is considered defective;
[0060] 8.1) After the supervised branch network finishes the calculation, through non-maximum suppression and setting the confidence threshold to filter out unnecessary prediction boxes, the prediction results of the bounding boxes for the defects are obtained. If the prediction result is empty, it means there is no defect; if it is not empty, it means there is a defect. The improvement in the detection effect brought by the proposed supervised branch is as Figure 4 shown in the figure on the right side of the dotted line;
[0061] 8.2) For the calculation of the unsupervised branch network, compare it with the threshold ut obtained during training. If it is greater, it is considered to have obvious large defects. For example, Figure 4 the scores in the figure on the left side of the dotted line are all much greater than 1.9, so they are all considered defective samples. Those in the figure on the right side of the dotted line less than 1.9 are considered non-defective.
[0062] Figure 6 In the embodiment, the datasets of four types of sealing rings A, B, C, and D collected are used. The comparative experiments on the regional coordinate attention proposed in this example are shown as Figure 7 follows. Origin represents without adding attention, +XX represents adding a certain type of attention, Ours represents the regional coordinate attention proposed in this example, and other attention mechanisms are some relatively common attention mechanisms currently. AP50 is a commonly used indicator in the object detection task, representing the average precision of detection when the IOU threshold is 0.5. The indicators in the parentheses of AP50 represent different types of sealing rings, mAP50 is the average of each type, and mAP50-95 is the average when the threshold is set between 50 and 95. It can be seen from Figure 7 the results that the results of this example are all higher than those of other attention mechanisms, indicating that the attention mechanism of the method in this example is superior to other methods.
Claims
1. A dual-branch deep learning method for detecting seal ring defects, characterized in that, It includes the following steps: 1) Collect and preprocess the O-ring images: Collect the O-ring images. First, segment the O-ring from the background by threshold segmentation for the O-ring images, and then eliminate the dust and impurities in the background through closing operation. After obtaining the mask of the O-ring, calculate the minimum bounding rectangle of the mask area and perform cutting. Finally, scale the cut O-ring images to the same size; 2) Use the images collected in step 1) to prepare the dataset for the dual-branch network: This dataset includes a supervised branch dataset and an unsupervised branch dataset; 3) Construct a dual-branch network. The entire dual-branch network consists of a backbone network, a supervised branch of RCA-BiFPN-P2, and an unsupervised branch: 3.1) First, construct the backbone network of the dual-branch defect detection network. This backbone network is a CSP network structure consistent with YOLOV8S; 3.2) The supervised branch takes the outputs of the first 3 C2f layers and the last SPPF network layer of the backbone network as the input of the supervised branch. The supervised branch consists of a Conv layer, a C2f layer, a regional coordinate attention layer, a Merge layer, and a detection head layer; among them, the Merge layer adopts the feature fusion method of BiFPN; the detection head layer adds a small-object detection head compared to YOLOV8; 3.3) Construct the unsupervised branch: The unsupervised branch takes the output of the last layer of the backbone network as the input. The unsupervised branch consists of global average pooling and eight identical normalizing flow modules, and sequentially connects the global average pooling layer and the normalizing flow modules before and after to process the input features; 4) Train the supervised branch of the network: Combine the backbone network with the supervised branch constructed in step 3.2), and use the supervised branch dataset in step 2) to train, test, and validate the combined network. The training objective is to enable the supervised branch to fit the bounding box of the defect; after training, obtain the weights of the backbone network and the supervised branch; 5) Train the unsupervised branch of the network: 5.1) Freeze the weights of the entire backbone network, combine the backbone network with the unsupervised branch, and use the unsupervised dataset in step 2) to train the unsupervised branch of the combined network. The training objective is to make the vector finally output by the unsupervised branch approach a normal distribution with a zero mean. After training, assuming the vector output by the unsupervised branch is Z, for whether the input image has a defect, a defect score can be calculated. The larger the value, the more likely it is a defect. The formula is: score = mean(Z 2 ); 5.2) Use the validation set of the unsupervised dataset to calculate the defect scores of all images to select a suitable score threshold ut for discriminating whether the image has a defect; and the unsupervised branch mainly serves as an auxiliary branch for detecting large defects. Therefore, the rule for selecting the threshold is to select the threshold with a high recall rate under the condition of ensuring the precision; 6) Integrate the network structures of the backbone network, the unsupervised branch, and the supervised branch, as well as the network models and weights obtained from training in steps 4) and 5), freeze all weights and turn off the model gradient calculation, and enter the detection; 7) For the image of the seal ring to be detected, perform the same operations as in the preprocessing in step 1). After locating the position of the seal ring, cut the image, normalize it to between 0 and 1, and input it into the backbone network for feature extraction; 8) The features output by the backbone network are respectively input into the supervised branch and the unsupervised branch for simultaneous calculation. As long as one branch deems it defective, it is considered defective; 8.1) After the supervised branch network finishes the calculation, after filtering out unnecessary prediction boxes through non-maximum suppression and the set confidence threshold, the prediction result of the bounding box for the defect is obtained. If the prediction result is empty, it means there is no defect; if it is not empty, it indicates there is a defect; 8.2) For the calculation of the unsupervised branch network, by comparing with the score threshold ut obtained during training, if it is greater than the threshold ut, it is considered that there is an obvious large defect.
2. The dual-branch deep learning method for seal ring defect detection according to claim 1, wherein, In the dataset of the supervised branch in step 2), it consists of defective images and defect-free images that are easily misdetected. The specific positions of the defects are labeled using image annotation software, and the annotation method is bounding box annotation. After annotation, the training, test, and validation datasets are divided in a ratio of 8:1:1; while for the dataset of the unsupervised branch, only the whole image needs to be divided into defective and defect-free ones. The training set only includes defect-free images, and the validation set includes some defect-free images and defective images.
3. A dual-branch deep learning method for seal ring defect detection according to claim 1, characterized in that, In step 3.2), the regional coordinate attention layer has two branches. For the input feature I1 with the shape of [C, H, W], one branch performs global average pooling on the features of I1, then performs a grouped convolution with a kernel size of 1 and 32 groups on the obtained features, and finally uses the sigmoid activation function to obtain the attention matrix A1; the other branch performs adaptive average pooling operations on I1 in the H direction and the W direction respectively to obtain feature vectors of sizes [C, H, N] and [C, N, W]. The two feature vectors are transposed to align the dimensions and concatenated to obtain [C, H + W, N], then input into a 1x1 grouped convolution, and finally re-split into vectors of sizes [C, H, N] and [C, N, W]. Then, through nearest neighbor interpolation, two features of size [C, H, W] are obtained. The sigmoid function is applied to each of the two features of size [C, H, W], and finally the two are multiplied to obtain the attention matrix A2; finally, (A1 + A2) / 2 is used to obtain the final attention map, which is multiplied by I1 and output.
4. A dual-branch deep learning method for seal ring defect detection according to claim 1, characterized in that, In step 3.3), the global average pooling aggregates the input features into a size of 1x1 in the spatial dimension, and the channel dimension remains unchanged. The normalizing flow module consists of modules with reversible calculations. The calculation of a single normalizing flow module is expressed as: X` = X @ W1 X`1, X`2 = Split(X`) X``1 = X`1 Z1, Z2 = Split(MLP(X`1)) X```2 = X``2 Z3, Z4 = Split(MLP(X``1)) Y = Concat(X```1, X```2) Where X is the input, Y is the output, X`, X1, X``1, X```1, X`2, X``2, X```2, Z1, Z2, Z3, Z4 all represent intermediate variables, @ represents matrix multiplication, W1 represents learnable weights, Split represents splitting the input into two parts of the same size in the channel dimension, MLP represents the stacking of fully connected layers, DropOut layers, and Relu activation function layers, stacked three times in total. Except that the input channel number of the first input fully connected layer is half of X, the input and output channel numbers of the remaining fully connected layers are equal to X.
Citation Information
Cited By
Supervision-based and non-supervision-based defect detection method, model establishment method and device
CN120931630A
Supervised and unsupervised defect detection methods, model building methods and devices
CN120931630B