Two-stage weak supervision fabric flaw detection method
By combining the two-stage weak supervision methods of VGG9 network and Tiny U-Net network, the problems of high computing resources and labeling workload in fabric defect detection are solved, and high-precision pixel-level defect detection is achieved, which improves detection accuracy and efficiency.
Patent Information
- Application Number
- CN202510424819.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-08
AI Technical Summary
When facing complexity and diversity, the detection accuracy of existing fabric defect detection methods is not high, and traditional methods require a large amount of labeling work and computing resources, making it difficult to meet actual production needs.
A two-stage weak-supervised fabric defect detection method is adopted, combined with the VGG9 network and the Tiny U-Net network, and high-precision pixel-level defect detection is achieved through a small amount of pixel-level annotation. The Grad-CAM algorithm is used for defect feature extraction and preliminary positioning, and Tiny U-Net is refined.
While maintaining low manual labeling and calculation costs, the accuracy of defect positioning is significantly improved, achieving detection and positioning accuracy comparable to that of the full supervision network.
Smart Images

Figure CN120278987A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of fabric defect detection, and specifically to a two-stage weakly supervised fabric defect detection method. Background Art
[0002] During the weaving process of fabrics, due to the influence of industrial processes, equipment, and the environment, the generation of surface defects is inevitable. These defects will significantly reduce the market value of fabrics and cause economic losses to textile enterprises. Therefore, defect detection has become a key link in fabric quality control. Although certain progress has been made in the research of fabric defect detection, traditional machine vision detection methods often require predefined defect features. When facing the complexity and diversity of fabric defects, they are easily affected by uneven illumination and noise, resulting in low detection accuracy and difficulty in meeting the actual production requirements. In addition, traditional machine learning algorithms perform poorly when dealing with large-scale datasets and are difficult to be widely applied. In contrast, deep learning can automatically extract deep features from a large amount of data, reducing the dependence on feature engineering and simplifying algorithm design, thus becoming a research hotspot.
[0003] Currently, fabric defect detection methods based on deep learning are mainly divided into three categories: image classification, object detection, and image segmentation. The image classification method learns a large number of labeled fabric images, trains a model to identify and extract features, and thus accurately classifies new images. However, it has high requirements for the proportion of defects in the image and can only classify the entire image without giving the precise location of the defects. Although the object detection method extracts candidate regions in the image, classifies and locates these regions, or directly obtains the position and category of the target region in the image, thereby making up for the defects of the image classification method, the annotation workload increases significantly. The semantic segmentation method assigns specific labels to each pixel of the image to achieve more accurate detection of the fabric defect location. However, its annotation process requires a large amount of pixel-level annotation work, resulting in a huge manual workload and also requiring high computing resources and memory support. For example, in the existing patent application CN117132541A, "Weakly Supervised Fabric Surface Defect Recognition Method Based on VGG-9 Simplified Network", the used VGG-9 simplified classification network has a simple structure, a small number of parameters, and excellent detection performance, which is suitable for online detection at the fabric production site and effectively avoids the heavy manual annotation problem in the existing object detection algorithms. However, the existing method can only achieve box-level annotation and requires a small amount of pixel-level annotation to achieve higher-precision pixel-level fabric defect detection.
[0004] In summary, in order to improve the fabric detection accuracy while avoiding excessive total annotation work and computer resource occupancy rate, it is of great academic significance and practical value to design an efficient and higher-precision detection new framework. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a two-stage weakly supervised fabric defect detection method, which is used to fuse the two-stage weakly supervised target detection framework of VGG9-Grad CAM and Tiny U-Net for lightweighting, and achieve higher-precision pixel-level fabric defect detection through a small amount of pixel-level annotation.
[0006] In order to solve the above technical problems, the present invention provides a two-stage weakly supervised fabric defect detection method, the process is: after preprocessing the collected original image, the defect features are extracted through the VGG9 network and the image is classified, and the Grad-CAM algorithm is applied to the image after the defect features are extracted to obtain the class discrimination positioning map The class discrimination localization map After normalization, image stitching, threshold processing and upsampling operations are performed in sequence, a class discriminant mask matrix B with the same size as the original image is obtained. The pixel value of each coordinate in the class discriminant mask matrix B represents the identified category of each pixel. The class discriminant mask matrix B is further refined by the offline trained Tiny U-net network to generate defect prediction candidate masks, and a fabric defect detection picture with defect locations and types is output.
[0007] The Tiny U-Net network retains the "U"-shaped structure composed of down-convolution and up-sampling of the U-Net network, including at least two convolutional layers and a maximum pooling layer for down-sampling, and at least two convolutional layers and an up-sampling layer for restoring resolution, and deletes the skip connection structure. The offline training of the Tiny U-Net network uses a class discrimination positioning map. Constructed mask dataset.
[0008] As an improvement of the two-stage weakly supervised semantic segmentation fabric defect detection method of the present invention, the normalization is:
[0009]
[0010] in: It is the class discrimination positioning map of the normalized cut image.
[0011] As a further improvement of the two-stage weakly supervised semantic segmentation fabric defect detection method of the present invention, the threshold processing:
[0012] Each element in the original image score matrix y formed by image splicing is thresholded, and elements below the 0.4 threshold are assigned a value of 0, and elements above the 0.4 threshold are assigned a value of 1 to obtain the elements of the two-dimensional score matrix y'.
[0013] As a further improvement of the two-stage weakly supervised semantic segmentation fabric defect detection method of the present invention, the construction process of the mask dataset is as follows:
[0014] Collect the original images of the fabric. After cutting them with a 96×96 window, use the VGG9 network to extract defect features and classify them. Apply the Grad-CAM algorithm to the images after extracting defect features to obtain the class discriminant localization map. Construct a mask dataset. The samples in the mask dataset include three types of labels: "vertical defect", "horizontal defect", and "no defect".
[0015] As a further improvement of the two-stage weakly supervised semantic segmentation fabric defect detection method of the present invention, the offline training process of the Tiny U-net network is as follows:
[0016] The samples in the mask dataset are randomly divided into a training set and a test set after being cut with a 96×96 window. Among them, the training set includes four datasets with different specifications, and the test set is one.
[0017] During training, there is no need to call pre-trained weights. The optimizer is selected as the Adam optimizer, and the learning rate decay method is selected as cosine annealing decay. Save the model parameters of the epoch with the minimum loss on the validation set as the optimal model network weights. The evaluation metrics used are mPA and mIoU.
[0018] The beneficial effects of the present invention are mainly reflected in:
[0019] 1. Combining a convolutional neural network and gradient-weighted class activation mapping can effectively extract defect features and achieve preliminary localization of defects.
[0020] 2. Introducing Tiny U-Net to optimize the defect mask can achieve precise localization, significantly improving the accuracy of localization while maintaining low manual annotation and computational costs.
[0021] 3. The test results on industrial datasets show that the proposed algorithm achieves detection and localization accuracy comparable to that of fully supervised networks with only a small amount of manual annotation. Description of the Drawings
[0022] The following further details the specific implementation manners of the present invention in conjunction with the drawings.
[0023] Figure 1 It is a flowchart of a two-stage weakly supervised fabric defect detection method of the present invention;
[0024] Figure 2 It is the structure diagram of the VGG9 network built by the present invention;
[0025] Figure 3It is a comparison diagram of the optimization effect of using the Tiny U-Net mask. Detailed implementation manners
[0026] The present invention will be further described below in conjunction with specific embodiments, but the protection scope of the present invention is not limited thereto:
[0027] Embodiment 1. A two-stage weakly supervised fabric defect detection method, the process is as Figure 1 shown, and the specific process is as follows:
[0028] Preprocess the collected original images and cut them into multiple small images of 96×96; retain the simplified VGG9 network in the existing invention and perform offline training to extract defect features and classify them; then based on Grad-CAM and image processing methods, realize the extraction and visualization of the spatial information of defects in the defect images, that is, realize the preliminary pixel-level positioning of fabric defects to obtain a rough mask, and then use the Tiny U-Net network to optimize the rough mask obtained by image processing, and improve the intersection over union of the generated mask and the actual defect image label mask.
[0029] 1. Image preprocessing
[0030] In the experiment of the present invention, the Windows10 operating system and the Pytorch deep learning development framework are used, and Python is used as the development language. The CPU of the computer used in the experiment is an Intel(R) Core(TM) i5-12490F CPU@3.00GHz, the GPU is an NVIDIA GeForce RTX 3060, the running memory is 16G, and the software running environment is Pycharm.
[0031] (1) The image set used for the training of the VGG9 network in the present invention is collected and constructed from a cooperative enterprise. It consists of an array of 6 high-resolution CCD industrial cameras, each equipped with a 5-megapixel lens, and synchronous sampling is performed in a horizontal configuration at an interval of 0.5 meters. Three 1.2-meter LED lights are placed directly above the cameras and arranged in a group to ensure uniform illumination. The distance between the camera and the fabric is 0.65 meters, and the maximum fabric width of 4 meters can be accommodated. According to the requirements of the fabric production environment and detection standards, each camera captures images at a resolution of 3072×96 pixels. Since the collected images have extreme aspect ratios, in order to highlight the features of horizontal and vertical defects at the same time, first perform image cutting preprocessing on the collected fabric images, that is, based on the original image size, after multiple experiments, select a window size of 96×96 pixels as the optimal size after cutting. This size can well highlight the fabric defect features and the different features of horizontal and vertical defects, and also solve the problem of extreme aspect ratios of horizontal and vertical defects. The image samples after cutting contain five types of defects (broken warp, double warp, broken weft, loose pick, tight pick).
[0032] The image set for VGG9 network training contains 1,173 original images with a resolution of 3072×96, which are divided into 5 defect classes and 1 defect-free class, a total of 6 different classes. After window cutting with a 96×96 window and appropriate augmentation, a small image classification dataset consisting of 49,465 images with a size of 96*96 can be obtained for the training of the VGG9 network, as shown in Table 1.
[0033] Table 1 Training dataset of VGG9 network
[0034]
[0035] To avoid a large gap in the number of defect images of each type in the classification dataset, when performing window cutting on the original dataset images of broken diameter and double diameter, 2 - 6 defect images are obtained by cutting around the cutting window of the defect class with a step size of 20 pixels.
[0036] Then, the dataset in Table 1 is randomly divided into a training set and a test set in a ratio of 9:1 for the training of the VGG9 network.
[0037] (2) Dataset for Tiny U-Net network training
[0038] The 6 types of defects (broken warp, double warp, broken weft, loose weft, tight weft, and defect-free class) in the training dataset of the VGG9 network are reclassified, merged, and labeled as 3 classes: "vertical defect", "horizontal defect", and "defect-free". Then, the training dataset of the VGG9 network after classification and merging passes through the pre-trained VGG9 network offline and the class discriminant localization map output by applying the Grad CAM algorithm to construct a mask dataset, creating a semantic segmentation dataset for the training of the Tiny U-Net network. The mask resolution of this mask dataset is 3072×96 pixels, the same as the original image. To balance the proportion of defect types in the dataset, all vertical defect and some horizontal defect images are included in the training dataset of the Tiny U-Net.
[0039] Samples in the mask dataset are preprocessed by cutting through a 96×96 window and then constructed into four datasets of different scales: TUNet dataset_L, TUNet dataset_M, TUNet dataset_S, and TUNet dataset_T, which are respectively used to train the Tiny U-Net model. The sizes of the datasets are represented by "L", "M", "S", and "T" respectively. The training set sizes of TUNet dataset_M, TUNet dataset_S, and TUNet dataset_T are 1 / 2, 1 / 5, and 1 / 10 of TUNet dataset_L respectively, and the smaller datasets are subsets of the larger datasets. The test sets of TUNet dataset_L, TUNet dataset_M, TUNet dataset_S, and TUNet dataset_T are the same. The training datasets of different scales are designed to evaluate the impact of dataset size on model performance, while the stability of the test set ensures the reliability of the evaluation. The specific distribution information is shown in Table 2.
[0040] Table 2 Training Datasets of Tiny U-Net Network
[0041]
[0042] 2. Build the VGG9 backbone network
[0043] The present invention uses the VGG9 network simplified from the VGG11 network as the feature extraction network, aiming to improve its detection performance for all types of defects, as Figure 2 shown. It includes 8 convolutional layers, 3 pooling layers, 1 fully connected layer, and 1 softmax layer. In the convolutional layers, the number of convolutional layers is kept unchanged to retain the robust feature extraction ability of the convolutional layers. The stride of the 3×3 convolutional filter is 1, and the padding is set to 1 to ensure that the size of the convolutional image remains unchanged. The pooling layers are streamlined from 5 layers to 3 layers to ensure that when the input size is 96×96×3, the output feature map size of the pooling layer increases from 3×3×512 of VGG11 to 12×12×512. The size of the pooling layer is 2×2, and the stride is 2, reducing the image size by half. The number of fully connected layers has been changed from three layers to one layer to reduce the model parameters (measured by Params) and computational complexity (measured by FLOPs). At the same time, by restricting the decision-making ability of the fully connected layer, the convolutional layers are encouraged to extract more accurate feature maps. The fully connected layer contains 1000 channels.
[0044] 3. Realize pixel-level defect localization based on Grad-CAM
[0045] Based on the classification results of the VGG9 model, Grad-CAM is used to extract and visualize the spatial information of defects in defect images, that is, to achieve preliminary pixel-level localization of fabric defects. The working process of this algorithm is as follows: First, Grad-CAM is used to generate a low-resolution class activation heatmap for the fabric image; then, through image processing methods such as normalization, threshold processing, and upsampling, the resolution of the class activation heatmap is upsampled to the size of the original image and the accuracy of the defect mask is improved; finally, Tiny U-Net is used to optimize the preliminary mask obtained in the previous step to generate a fine mask localization.
[0046] 3.1 Preliminary Localization of Defects Based on Grad-CAM
[0047] Class activation mapping (CAM) is an algorithm that provides a clear and intuitive explanation of the neural network decision-making process. It achieves this by visualizing the feature map, that is, by performing a weighted sum on the activation map to represent the importance of different regions of the input image for a given class.
[0048] Suppose there are k feature maps A (k) ∈R (u′v) , representing the output feature map of the last max pooling layer in VGG9, where each has a width of u, a height of v, and c represents the corresponding target class. To obtain the class discriminant localization map L as the input for the subsequent normalization step, first, calculate the gradient information obtained by backpropagation Subsequently, obtain the importance weights by applying global average pooling in the width and height dimensions The calculation of these weights is shown in Equation (1).
[0049]
[0050] where: y c is the score corresponding to class c obtained by forward propagation, is the data at the coordinates (i, j) on the k-th channel of the convolutional feature map matrix A output by the last max pooling layer in VGG9, i is the width, j is the height, and Z is the product of the width i and the height j.
[0051] Then, calculate the product of the convolutional feature map matrix A k and the importance weights , then perform a weighted sum, and output after activation by the ReLU function. The mathematical expression of this process is shown in Equation (2).
[0052]
[0053] where: Ak is the convolutional feature map matrix, is the importance weight, is the class discriminant localization map.
[0054] To obtain a heat map with the same size as the original image, image processing methods such as normalization, thresholding, and upsampling need to be applied to the class discriminant localization map to upsample the resolution of the class activation heat map to the size of the original image and improve the accuracy of the defect mask. It should be noted that since the mask refinement step based on Tiny U-Net in the present invention achieves the same function as the image dilation and image erosion operations but with better effects, the image dilation and image erosion operations of the original application are not adopted, simplifying the steps and computational resources. The detailed steps are as follows:
[0055] (1) Normalization: The classification output of the VGG9 model is c. If the recognized class c corresponds to a non-defective class, no additional processing is required. In the case where the class is not non-defective, the class discriminant localization map will be subjected to min-max normalization. The mathematical expression of min-max normalization is as shown in Equation (3):
[0056]
[0057] Where: is the class discriminant localization map, is the class discriminant localization map of the normalized cropped small image.
[0058] (2) Image stitching: The class discriminant localization maps of the normalized cropped small images are stitched together in sequence to form the class discriminant localization map of the original image. In the subsequent steps, it is referred to as the original image score matrix y.
[0059] (3) Thresholding: At this stage, each element in the original image score matrix y is thresholded to obtain a two-dimensional score matrix y'. Its mathematical expression is as follows:
[0060]
[0061] Where the elements in the original image score matrix y that are lower than the threshold of 0.4 are assigned a value of 0, and the elements that exceed the threshold are assigned a value of 1 to obtain the elements of the two-dimensional score matrix y'.
[0062] Thresholding is a key step in achieving pixel-level classification of the original image. Specifically, in this step, the elements in the original image score matrix that are lower than the threshold of 0.4 are set to zero, effectively classifying the pixels with low defect confidence as the background. On the contrary, the elements that exceed the threshold of 0.4 are assigned a value of 1, representing the defect class, and are subsequently assigned an appropriate "class_id" after upsampling, that is, the specific defect class.
[0063] (4) Upsampling: Use bilinear interpolation to upsample the two-dimensional score matrix y' obtained in the previous step. Then multiply each element in the enlarged matrix by its corresponding "class_id" to achieve segmentation mapping. It should be noted that the output result of this process is a class discriminant mask matrix B with the same width and height as the original fabric image, and the pixel value at each coordinate in the class discriminant mask matrix B represents the recognized class of each pixel.
[0064] 3.2 Mask Optimization Based on Tiny U-Net
[0065] The Tiny U-Net proposed in this invention is essentially a convolutional network model for image segmentation that mimics U-Net. By deleting the skip connection structure, reducing the number of network layers, and reducing the number of filters in each layer, Tiny U-Net greatly reduces the computational complexity and memory occupancy of the model, making it suitable for edge computing and resource-constrained image segmentation tasks in fabric detection scenarios. The specific structure comparison of Tiny U-Net is shown in Table 3.
[0066] Table 3 Structure Comparison of Tiny U-Net
[0067]
[0068]
[0069] Tiny U-Net is a simplified image segmentation network that sacrifices some of the complexity and potential accuracy of the standard U-Net in exchange for model lightweight and computational efficiency. It retains the "U" - shaped structure composed of down - convolution and up - sampling in U-Net, reduces the number of convolutional layers and pooling layers, and deletes the skip connection structure that combines the high - resolution features obtained by down - convolution and the up - sampling output. The input of the model is the class discriminant mask matrix B. Table 3 shows three Tiny U-Net variants with different complexities: U5 - Net (5 layers), U7 - Net (7 layers), and U12 - Net (12 layers). All variants contain convolutional layers (conv3, 3x3 convolutional kernel), max - pooling layers (maxpool) for down - sampling, and up - sampling layers (Upsampling) for resolution recovery. As the number of layers increases (from U5 to U12), the depth of the network and the number of convolutional layers in each stage also increase correspondingly, and the number of channels (out_channels) also becomes larger in the middle of the network. The final output layer converts the feature map into a segmentation mask (mask) with the specified number of classes (num_classes), and the size is the same as the input image (96x96). Specifically, Tiny U-Net extracts the pixel information of each sliding window in the rough mask of the original image with a 96×96 - pixel sliding window and a 96 - pixel stride as the input of Tiny U-Net, and outputs an optimized window mask covering the pixel information of the original sliding window.
[0070] 4. Offline Training
[0071] (1) Offline Training of VGG9 Network
[0072] During training, there is no need to call pre - trained weights. The training parameters are as follows: the epoch is set to 400, the Batch_Size is set to 128, the initial learning rate of the model is set to 0.0006, the minimum learning rate is set to 0.000006, the optimizer is selected as the SGD optimizer, the optimizer momentum parameter is set to 0.9, the weight_decay is set to 0.0005, and the learning rate decay method is selected as cosine annealing decay. Save the model parameters of the epoch with the minimum loss on the validation set as the optimal model network weights.
[0073] The evaluation metrics used are Accuarcy, Recall, and Precision, and Accuarcy = 98.95%, thus obtaining a well - trained VGG9 network for offline training.
[0074] (2) Offline Training of Tiny U - Net Network
[0075] During training, there is no need to call the pre-trained weights. The training parameters are as follows: the number of epochs is set to 400, the Batch_Size is set to 16, the initial learning rate of the model is set to 0.000001, the minimum learning rate is set to 0.00000001, the optimizer is selected as the Adam optimizer, the momentum parameter of the optimizer is set to 0.9, the weight_decay is set to 0, and the learning rate decay method is selected as cosine annealing decay. The model parameters of the epoch with the minimum loss on the validation set are saved as the optimal model network weights. The evaluation metrics used are mPA and mIoU.
[0076] 5. Online use of the two-stage weakly supervised object detection framework integrating the trained VGG9-Grad CAM and Tiny U-Net
[0077] After applying the window method for cutting and preprocessing the original image with defects, it is input into the newly developed weakly supervised detection framework integrating VGG9-Grad CAM and Tiny U-Net obtained in steps 2 and 3 to generate a candidate mask for defect prediction, thereby outputting a fabric defect detection image with the location and type of the defect.
[0078] Experiments
[0079] 1. The experimental dataset uniformly uses the training dataset constructed in Example 1.
[0080] 2. Evaluation metrics
[0081] In the experiment, Precision, Recall, Accuracy, Mean Intersection over Union (MIoU), and MPA are used to evaluate the model performance, and the calculation formulas are as follows:
[0082]
[0083] In the formula, Precision is the precision rate, Recall is the recall rate, Accuracy is the proportion of the number of correctly classified samples to the total number of samples, TP is the number of correctly classified positive samples in the prediction result, FP is the number of misclassified positive samples, FN is the number of misclassified negative samples, TN is the number of correctly classified negative samples, i is the actual value, j represents the predicted value, p ij represents the number of pixels where i is predicted as j, p ji represents the number of pixels where j is predicted as i, p ii represents the number of pixels where i is correctly predicted as i.
[0084] 3. Comparative experiment of backbone networks
[0085] To verify the detection effect of the VGG9 model used in the present invention, VGG11 and VGG16 are introduced for comparison. The network structures are shown in Table 4 for details.
[0086] Table 4 Comparison of VGG Network Structures
[0087]
[0088]
[0089] Table 5 Comparison of VGG Network Performances
[0090]
[0091] The following conclusions can be drawn from Table 4 and Table 5: (1) Compared with VGG11 and VGG16, the VGG9 model only retains one fully connected layer and has fewer parameters. Nevertheless, the gap between it and the other two models in terms of accuracy and recall is very small, and even slightly improved in terms of precision. (2) In order to obtain a larger output feature map, the VGG9 model only retains two maxpool layers, which is to retain more defect location information before applying the Grad-CAM algorithm.
[0092] 4. Comparative Experiment on Tiny U-Net Optimization
[0093] After applying the window method cutting preprocessing to the original defective images, the preprocessed small images are sequentially input into the VGG9 model to obtain the class_id array and the convolutional feature map set. Apply the Grad-CAM algorithm to the feature map output by VGG9 to obtain the pixel-level classification result of the cut 96×96 image, that is, the class discriminant mask matrix B. In this comparative experiment, the present invention uses Tiny U-Net with different numbers of layers as the mask post-processing method to compare the improvement degree of the intersection over union between the masks of Tiny U-Net with different depths and the actual segmentation labels; use datasets containing different numbers of images to compare the influence of the dataset size on the mask quality generated by Tiny U-Net with different depths.
[0094] Table 6 shows the comparison of the mask optimization effects obtained by training Tiny U-Net with different numbers of layers using datasets of different sizes.
[0095] Table 6 Experimental Results under Different Dataset Sizes
[0096]
[0097] The experimental results show that compared with the existing application of simply using VGG9-Grad CAM, the MIoU and MPA of the segmentation mask generated after the introduction of Tiny U-Net are significantly improved. Moreover, with the increase of the number of Tiny U-Net layers (from TUNet5, TUNet7 to TUNet12), mIoU and mPA also increase accordingly. It can be seen that to a certain extent, increasing the number of convolutional layers of Tiny U-Net helps to improve the optimization effect of the segmentation mask. Comparison of the mask optimization effect of Tiny U-Net with different layers trained using the "T" size dataset Figure 3 .Depend on Figure 3 The comparison results show that: no matter for vertical defects or horizontal defects, the positioning effect of the mask generated by the algorithm proposed in the present invention is more accurate.
[0098] In addition, the original application only used the VGG9 network and Grad-CAM to achieve frame-level defect detection, while the present invention aims to integrate the VGG9 network and the Grad-CAM algorithm with the Tiny U-Net structure to obtain a more accurate pixel-level defect detection effect. This change enables the present invention to obtain more than 70% MIoU and about 80% MPA value, which is much higher than the 58.29% MIoU and 80.69% MPA value obtained by the VGG9-Grad CAM network in the original application on the same data set.
[0099] 65.97% MPA. This shows that the present invention can achieve a significant improvement in accuracy based on the original application.
[0100] Finally, it should be noted that the above examples are only some specific embodiments of the present invention. Obviously, the present invention is not limited to the above embodiments, and there are many variations. All variations that can be directly derived or associated with the content disclosed by a person skilled in the art should be considered as the protection scope of the present invention.
Claims
1. A two-stage weakly supervised fabric defect detection method. After preprocessing the collected original image, the VGG9 network is used to extract defect features and classify the image. For the image after extracting defect features, the Grad-CAM algorithm is applied to obtain a class discriminant localization map It is characterized in that: The class discrimination and localization map After performing normalization, image stitching, threshold processing, and upsampling operations in sequence, a class discrimination mask matrix B with the same size as the original image is obtained. The pixel value at each coordinate in the class discrimination mask matrix B represents the recognized class of each pixel. Then, the pre-trained Tiny U-net network is used to further refine the class discrimination mask matrix B to generate a defect prediction candidate mask, and a fabric defect detection image with defect positions and types is output; The Tiny U-Net network retains the "U" - shaped structure composed of down - convolution and up - sampling in the U-Net network, includes at least two convolutional layers and max - pooling layers for down - sampling, and at least two convolutional layers and up - sampling layers for resolution recovery, and deletes the skip - connection structure. The offline training of the Tiny U-Net network uses a class discriminant localization map The constructed dataset.
2. The two-stage weakly supervised fabric defect detection method according to claim 1, wherein The normalization is as follows: Wherein: is the class discrimination positioning map of the normalized cut small image.
3. The two-stage weakly supervised fabric defect detection method according to claim 2, wherein The threshold processing is as follows: Each element in the original graph score matrix y formed by image stitching is subjected to threshold processing. Elements lower than the 0.4 threshold are assigned 0, and elements exceeding the 0.4 threshold are assigned 1 to obtain the elements of the two-dimensional score matrix y'.
4. A two-stage weakly supervised fabric defect detection method according to claim 3, characterized in that The construction process of the mask dataset is as follows: Collect the original image of the fabric. After cutting it through a 96×96 window, use the VGG9 network to extract defect features and classify them. Apply the Grad-CAM algorithm to the image after extracting defect features to obtain a class discriminant localization map. Construct a mask dataset. The samples in the mask dataset include three types of labels: "vertical defect", "horizontal defect", and "no defect".
5. A two-stage weakly supervised fabric defect detection method according to claim 4, characterized in that, The offline training process of the TinyU-net network is as follows: Samples in the mask dataset are randomly divided into a training set and a test set after being cut by a 96×96 window. Among them, the training set includes four datasets with different specifications, and the test set is one. During training, there is no need to call pre-trained weights. The optimizer is selected as the Adam optimizer, the learning rate decay method is selected as cosine annealing decay, and the model parameters of the epoch with the smallest loss on the validation set are saved as the optimal model network weights. The evaluation metrics used are mPA and mIoU.
Citation Information
Patent Citations
Weak supervision fabric surface flaw identification method based on VGG-9 simplified network
CN117132541A