Cloth defect detection method based on deep learning

By building a real-time detection converter RT-DETR model and combining wavelet convolution and fuzzy attention mechanism, the problem of low detection accuracy in cloth defect detection is solved, and efficient and accurate defect detection is achieved to meet the needs of modern textile industry.

CN120219288APending Publication Date: 2025-06-27NANTONG UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510187812.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art has problems such as low detection accuracy, high error detection rate and leakage detection rate in cloth defect detection, and is sensitive to image noise and light changes, making it difficult to meet the needs of the modern textile industry.

Method used

Using a deep learning-based method, a real-time detection transformer RT-DETR convolutional neural network model is constructed, and the detection capability of the model is enhanced through wavelet convolution and fuzzy attention mechanisms.

Benefits of technology

Automatic, efficient and accurate detection of fabric defects is realized, which improves detection efficiency and accuracy, and reduces the cost and error of manual inspection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219288A_ABST
    Figure CN120219288A_ABST
Patent Text Reader

Abstract

The invention discloses a cloth defect detection method based on deep learning, and the method comprises the steps: 1, collecting the cloth defect features of target cloth, and making a cloth quality inspection data set; 2, taking the cloth quality inspection data set as input, taking cloth defect category and position information as output, constructing an RT-DETR convolutional neural network model, and training the model; 3, using wavelet convolution WTConv to increase the capturing ability of the model to small target defects; and 4, improving the adaptability of the model to uncertain defects by using fuzzy attention. Various defects on the cloth can be automatically, efficiently and accurately detected, the cloth quality detection efficiency and precision are improved, and the manual detection cost and error are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of machinery and artificial intelligence, and particularly relates to a method for detecting fabric defects based on deep learning. Background Art

[0002] In the textile industry, the quality inspection of fabrics is a crucial link. Traditional fabric defect detection mainly relies on manual visual inspection. This method is not only inefficient but also easily affected by human factors, making it difficult to ensure the accuracy and consistency of the detection results. With the continuous expansion of production scale and the increasing improvement of quality requirements, the manual detection method can no longer meet the needs of modern textile industry.

[0003] In recent years, automatic detection technologies based on machine vision have gradually been applied to the field of fabric defect detection. However, most of the existing machine vision detection systems adopt traditional image processing algorithms, such as edge detection, threshold segmentation, etc. When dealing with complex fabric textures and diverse defect types, these methods often have problems such as low detection accuracy, high false detection rate and missed detection rate. In addition, traditional algorithms are relatively sensitive to interference factors such as image noise and illumination changes, further affecting the stability of detection. Therefore, how to automatically, efficiently and accurately detect various defects on fabrics and improve the efficiency and accuracy of fabric quality inspection has become a hot issue in this field. Summary of the Invention

[0004] Object of the Invention: The object of the present invention is to provide a method for detecting fabric defects based on deep learning, which can automatically, efficiently and accurately detect various defects on fabrics, improve the efficiency and accuracy of fabric quality inspection, and reduce the cost and error of manual detection.

[0005] Technical Solution: A method for detecting fabric defects based on deep learning according to the present invention includes the following steps:

[0006] Step 1: Collect the fabric defect features of the target fabric and make a fabric quality inspection data set;

[0007] Step 2: Use the fabric quality inspection data set as the input, and the fabric defect category and position information as the output to construct a real-time detection transformer RT-DETR (Real-Time Detection Transformer) convolutional neural network model, and train the model;

[0008] Step 3: Use wavelet convolution WTConv to increase the model's ability to capture small target defects;

[0009] Step 4: Use fuzzy attention to improve the model's adaptability to uncertain defects.

[0010] Further, step 1 specifically includes the following steps:

[0011] Step 1.1: In the textile production workshop, collect image data of several plain cloths, including defect-free samples and samples with defects.

[0012] Step 1.2: In the data preprocessing stage, screen and eliminate defective pictures with insufficient sample size; divide the data set into four categories: the first category is warp-direction significant defects, including warp-direction problems such as knots and thick warps; the second category is weft-direction significant defects, covering weft-direction defects such as thick wefts and weft contractions; the third category is obvious stain types, including pollution problems such as various stains and oil stains; the fourth category is severe defects, involving defect types such as holes and skipped picks that seriously affect the quality of the cloth.

[0013] Step 1.3: Based on the annotation information, taking the center point of the defect annotation box as the benchmark, randomly offset 104 to 312 pixels to the upper left to determine the starting coordinates of the cropping area; then, taking this coordinate as the benchmark, extend 416 pixels to the lower right to intercept a local image of 416×416 pixels; each annotation box undergoes three random samplings to generate three new training samples; during the cropping process, if the generated image area exceeds the boundary of the original image, the sample will be discarded.

[0014] Step 1.4: Perform Mosaic data augmentation on the defect images.

[0015] Step 1.5: Divide the data set into a training set, a validation set, and a test set, and adjust the model parameters through an optimization algorithm to minimize the loss function. The loss function uses a cross-entropy loss function combined with a localization loss function. The specific formula is:

[0016] L = αL cls + βL loc

[0017] where L cls is the classification loss, L loc is the localization loss, and α and β are weight coefficients used to balance the losses of classification and localization.

[0018] Further, step 1.4 is specifically: randomly crop four pictures and then splice them onto one picture as training data; assume the four pictures are A, B, C, and D respectively, the sizes of the cropped pictures are A′, B′, C′, and D′ respectively, and the size of the spliced picture is E. Then there is:

[0019]

[0020] In the above formula, the sizes and positions of A′, B′, C′, and D′ are randomly determined.

[0021] Further, step 2 specifically includes the following steps:

[0022] Step 2.1: Build and train the RT-DETR object detection framework by decomposing multi-scale feature interaction into inner-scale interaction and cross-scale fusion;

[0023] Step 2.2: Backbone network: As a basic feature extractor, it extracts multi-scale spatial feature maps from the input image; ResNet-18 is used for feature extraction, including convolutional layers, max pooling layers, and basic modules with residual structures; the convolutional layers are responsible for initially extracting image features through convolution operations, batch normalization processing, and the application of activation functions; the max pooling layers play a role in downsampling to reduce the size of the feature maps; in the deep feature extraction part, the network uses multiple Basicblocks to further extract complex image features, and the Basicblocks in different stages have different output channel numbers;

[0024] Step 2.3: Efficient hybrid encoder: The efficient hybrid encoder is used to convert the multi-scale features extracted by the backbone network into an image feature sequence suitable for Transformer processing; this encoder integrates feature information of different scales through the inner-scale feature interaction module AIFI and the cross-scale feature fusion module CCFM, and the calculation process of the encoder is shown in the following formula:

[0025] Q = K = V = Flatten(S)

[0026] F = Reshape(AIFI(Q, K, V))

[0027] In the formula, S is the output feature, Q, K, and V respectively represent different data matrix or vector sets. Q is the query, containing the information that needs to be focused on or queried currently. K is the key, providing the reference information for comparison with Q. V is the value, containing the actual information content. Flatten means flattening and rearranging the S feature to adapt to the input format of the Transformer. Reshape means restoring the processed feature map to the original spatial dimension and number of channels;

[0028] Step 2.4: Uncertainty-minimum query selection mechanism: Select a certain number of image features from the sequence output by the encoder as the initial object queries for the decoder. Specifically define the joint uncertainty of the encoder output features in classification and localization prediction, and integrate this uncertainty into the loss function to guide the optimization process of the model. This selector can select the features most relevant to the target object according to the intersection over union (IoU) perception mechanism;

[0029] Step 2.5, Decoder: The decoder with an auxiliary prediction head iteratively optimizes the object queries to generate bounding boxes and confidence scores; the decoder processes the selected features through a self-attention mechanism and a feed-forward neural network layer, and finally outputs the predicted defect categories and location information.

[0030] Further, in step 2.4, the loss function includes an object confidence loss a classification loss and a bounding box regression loss These three parts; the total loss is the weighted sum of these three parts, and the expression is:

[0031]

[0032] where the weights of the object confidence loss, classification loss, and bounding box regression loss are λ1, λ2, and λ3 respectively; L conf and L cls are calculated using the binary cross-entropy function, and the calculation formula is as follows:

[0033] L n = -(y n log(δ(x n )) + (1 - y n ) log(1 - δ(x n ))

[0034] In the formula: x n is the score for predicting the nth sample as a positive example, y n represents the label of the nth sample, δ represents the sigmoid function, and L reg uses the CloU Loss function. CloU is a method for calculating the bounding box regression loss, and the calculation formula is as follows:

[0035]

[0036]

[0037] In the formula: the parameter A represents the ground truth box, B represents the predicted box, IoU is called the intersection over union, which represents the overlapping degree of the predicted box and the ground truth box, ρ 2 (·) represents the Euclidean distance, b represents the center point coordinates of the predicted box, b gt represents the center point coordinates of the ground truth box, a represents the balance parameter, and v is the consistency of the aspect ratios of the ground truth box and the predicted box.

[0038] Further, step 3 specifically includes the following steps:

[0039] Step 3.1: The WTConv convolutional layer utilizes the characteristics of wavelet transform to expand the receptive field of the convolutional neural network (CNN). WTConv first decomposes the input image into different frequency components through wavelet transform, including the low-frequency component LL, the horizontal high-frequency component LH, the vertical high-frequency component HL, and the diagonal high-frequency component HH. When using the Haar wavelet transform, four filters are used to perform a depth convolution operation on the image to obtain these four frequencies;

[0040] Step 3.2: Perform small-size convolutions on each frequency layer to capture the multi-scale information of the image through small convolution kernel operations; recombine the convolution results of each frequency layer through inverse wavelet transform to form the final output.

[0041] Further, step 4 specifically includes the following steps:

[0042] Step 4.1: Use the fuzzy attention mechanism to reduce the uncertainty in the feature map through fuzzy entropy and the fuzzy membership function, and improve the model's attention to target features;

[0043] Step 4.2: In fabric defect detection, each pixel point in the feature map is regarded as belonging to two fuzzy sets, "defect" and "non-defect"; by defining the fuzzy membership function, the degree to which each pixel point belongs to these two sets is quantified. Assume the input feature map is represented as F∈R H×W×C , where H, W, and C represent the height, width, and number of channels of the feature map respectively. For the feature value of each pixel point (i, j) on channel c Its fuzzy membership function is expressed as:

[0044]

[0045] where α and β are learnable parameters used to adjust the shape of the fuzzy membership function, and μ defect represents the membership degree of this pixel point belonging to the defect set;

[0046] Step 4.3: In fabric defect detection, based on the fuzzy membership function, further calculate the fuzzy entropy of each pixel point to measure its uncertainty; the higher the fuzzy entropy value, the greater the uncertainty of this pixel point; to reduce the impact of this uncertainty on the detection performance, assign a weight to each pixel point The weight is inversely proportional to the fuzzy entropy:

[0047]

[0048] where, is the fuzzy entropy value, and ∈ represents a very small positive number to ensure that the denominator is not zero, thereby enhancing the stability of numerical calculations.

[0049] The present invention also discloses a computer device, comprising a memory, a processor and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the method of the present invention.

[0050] The present invention also discloses a computer-readable storage medium on which a computer program / instruction is stored. When the computer program / instruction is executed by a processor, the steps of the method of the present invention are implemented.

[0051] The present invention also discloses a computer program product, comprising a computer program / instruction, which implements the steps of the method of the present invention when executed by a processor.

[0052] Beneficial effects: Compared with the prior art, the present invention has the following significant advantages:

[0053] 1. The present invention adopts deep learning technology, uses the transformer-based target detection network RT-DETR, and uses the self-attention mechanism to capture long-range dependencies in the entire image to achieve end-to-end target detection; by decomposing multi-scale feature interactions into two-step operations of intra-scale interaction and cross-scale fusion, the consumption of computing resources is significantly reduced, making it convenient to deploy the model on the machine vision hardware platform;

[0054] 2. Replace the wavelet convolution WTConv, use the characteristics of wavelet transform to decompose the image into multi-scale frequency components, and perform convolution operations on each component, which can capture the characteristic information of defects more comprehensively, thereby expanding the receptive field of the model and enhancing the model's ability to detect small target defects;

[0055] 3. The fuzzy attention mechanism is used to calculate the fuzzy entropy of each pixel to measure its uncertainty, so that the model can focus more on the defect characteristics of the cloth surface, while suppressing noise and background interference, thereby improving the accuracy and robustness of defect detection;

[0056] 4. The present invention adopts a roller tensioning method for the conveyor belt transport device, and adjusts the tensioning force according to the real-time tightness of the conveyor belt. Compared with the traditional conveyor belt transport, it effectively solves the problems of decreased conveying efficiency and component wear caused by loose conveyor belts, and significantly prolongs the service life of the equipment. At the same time, this tensioning method can keep the conveyor belt in stable contact with each roller, reduce deviation and jitter, and improve the smoothness of cloth transportation;

[0057] 5. The present invention is provided with adjustable width limit plates at both ends of the conveyor belt. Compared with the traditional conveyor belt without limit plates, the adjustable width limit plates of the present invention can play a guiding and anti-deviation role in the conveying process, which not only improves the conveying efficiency of the cloth, but also provides better basic conditions for the subsequent detection module to accurately identify defects. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 This is the overall flowchart of the present invention.

[0059] Figure 2 This is the structure diagram of the deep learning network.

[0060] Figure 3 This is the structure diagram of the wavelet convolutional network.

[0061] Figure 4 This is the structure diagram of the fuzzy attention network. Specific implementation manners

[0062] The technical solution of the present invention will be further described below with reference to the accompanying drawings.

[0063] Example 1

[0064] As Figure 1 shown, a method for detecting fabric defects based on deep learning of the present invention includes the following steps:

[0065] S1. Collection of fabric defect features and production of a fabric quality inspection data set;

[0066] (1) In a textile production workshop, 2264 image data of plain fabrics were collected, including 1125 defect-free samples and 1139 samples with defects. Among these defect images, each may present single or multiple defect features, and a total of 20 different types of fabric defects are involved;

[0067] (2) In the data preprocessing stage, defect images with insufficient sample size were screened and removed. Based on common defect features in the textile industry, the data set was divided into four main categories: the first category is warp-direction significant defects, including warp-direction problems such as knots and thick warps; the second category is weft-direction significant defects, covering weft-direction defects such as thick wefts and weft contractions; the third category is obvious stain types, including various stains, oil stains and other pollution problems; the fourth category is serious defects, involving defect types such as holes and skipped patterns that seriously affect the quality of the fabric:

[0068] (3) Based on the annotation information, taking the center point of the defect annotation box as the reference, randomly offset 104 to 312 pixels to the upper left to determine the starting coordinates of the cropping area. Subsequently, extending 416 pixels to the lower right based on this coordinate, a local image of 416×416 pixels is intercepted. Each annotation box undergoes three random samplings to generate three new training samples. During the cropping process, if the generated image area exceeds the boundary of the original image, the sample will be discarded. Through this data augmentation method, an augmented data set containing 2597 defect samples was finally obtained;

[0069] (4) Perform Mosaic data augmentation on the defect images. The specific operation is to randomly crop four images and then splice them onto one image as training data. Let the four images be A, B, C, and D respectively, the sizes of the cropped images be A′, B′, C′, and D′ respectively, and the size of the spliced image be E. Then we have:

[0070]

[0071] In the above formula, the sizes and positions of A′, B′, C′, and D′ are randomly determined. In this way, the cloth defect quality inspection data set is greatly enriched. Especially, random scaling adds many small targets, enhancing the model's ability to detect small target defects. The data set is divided into a training set, a validation set, and a test set in the ratio of 8:1:1. The model parameters are adjusted through an optimization algorithm to minimize the loss function. The loss function uses a cross-entropy loss function combined with a localization loss function. The specific formula is

[0072] L = αL cls + βL loc

[0073] where L cls is the classification loss, L loc is the localization loss, and α and β are weight coefficients used to balance the losses of classification and localization.

[0074] S2. Build and train the RT-DETR convolutional neural network model;

[0075] (1) Build and train the RT-DETR object detection framework. RT-DETR (Real-Time Detection Transformer) is a real-time object detection model based on the Transformer architecture. It captures long-range dependencies in the entire image through the self-attention mechanism to achieve end-to-end object detection. On this basis, RT-DETR significantly reduces the consumption of computing resources by decomposing multi-scale feature interaction into two steps: inner-scale interaction and cross-scale fusion, avoiding complex post-processing steps such as candidate region generation and non-maximum suppression (NMS) in traditional convolutional neural network methods. Compared with mainstream methods such as Faster R-CNN and YOLO, it has significant improvements in both speed and accuracy and is more suitable for real-time application scenarios. Its model architecture mainly consists of the following parts;

[0076] (2) Backbone Network: As the basic feature extractor, it extracts multi-scale spatial feature maps from the input image. ResNet-18 is used for feature extraction, which includes convolutional layers, max pooling layers, and basic modules with residual structures. The convolutional layers are responsible for initially extracting image features. Through convolution operations, batch normalization processing, and the application of activation functions, the expressive ability of features is effectively enhanced. The max pooling layers play a role in downsampling, reducing the size of the feature maps. In the deep feature extraction part, the network uses multiple Basicblocks to further extract complex image features. The Basicblocks in different stages have different output channel numbers, which can capture multi-scale features of fabric defects;

[0077] (3) Efficient Hybrid Encoder: The efficient hybrid encoder is a key component in the RT-DETR model, used to convert the multi-scale features extracted by the backbone network into an image feature sequence suitable for Transformer processing. This encoder effectively integrates feature information at different scales through the Intra-scale Feature Interaction Module (AIFI) and the Cross-scale Feature Fusion Module (CCFM), improving the model's feature extraction ability and computational efficiency. The calculation process of the encoder is shown as follows:

[0078] Q = K = V = Flatten(S)

[0079] F = Reshape(AIFI(Q, K, V))

[0080] In the above formula, S is the output feature, and Q, K, and V represent different data matrix or vector sets respectively. Q (query) contains the information that needs to be focused on or queried currently, K (key) provides the reference information for comparison with Q, V (value) contains the actual information content. Flatten means flattening and rearranging the S feature to adapt to the input format of Transformer, and Reshape means restoring the processed feature map to the original spatial dimensions and number of channels;

[0081] (4) Uncertainty-Minimal Query Selection Mechanism: Select a certain number of image features from the sequence output by the encoder as the initial object queries for the decoder. The joint uncertainty of the encoder output features in classification and localization prediction is specifically defined, and this uncertainty is integrated into the loss function to guide the optimization process of the model. This selector can select the features most relevant to the target object according to the IoU (Intersection over Union) perception mechanism, improving the detection accuracy. The defined loss function includes the object confidence loss classification loss and bounding box regression loss The three parts. The total loss is the weighted sum of these three parts, and the expression is:

[0082]

[0083] Among them, the weights of the target confidence loss, classification loss, and bounding box regression loss are λ1, λ2, and λ3 respectively. L conf and L cls are calculated using the binary cross-entropy function, and the calculation formula is as follows:

[0084] L n = -(y n log(δ(x n )) + (1 - y n ) log(1 - δ(x n )))

[0085] In the formula: x n is the score for predicting the nth sample as a positive example, y n represents the label of the nth sample, and δ represents the sigmoid function. L reg uses the CloU Loss function. CloU is a method for calculating the bounding box regression loss, and the calculation formula is as follows:

[0086]

[0087] In the formula: The parameter A represents the ground truth box, B represents the predicted box, IoU is called the intersection over union, which represents the overlapping degree of the predicted box and the ground truth box, ρ 2 (·) represents the Euclidean distance, b represents the center point coordinates of the predicted box, b gt represents the center point coordinates of the ground truth box, a represents the balance parameter, and v is the consistency of the aspect ratios of the ground truth box and the predicted box.

[0088] (5) Decoder: The decoder with an auxiliary prediction head iteratively optimizes the object queries to generate the bounding boxes and confidence scores. The decoder processes the selected features through the self-attention mechanism and the feed-forward neural network layer, and finally outputs the predicted defect class and location information.

[0089] S3. As Figure 3 shown in the network structure, the wavelet convolution WTConv is used to increase the network's ability to capture small target defects;

[0090] WTConv (Wavelet Transform Convolution) is a novel convolutional layer whose core lies in using the characteristics of wavelet transform to expand the receptive field of a convolutional neural network (CNN). Specifically, WTConv first decomposes the input image into different frequency components through wavelet transform, including the low-frequency component (LL), the horizontal high-frequency component (LH), the vertical high-frequency component (HL), and the diagonal high-frequency component (HH). For example, when using the Haar wavelet transform, four filters are used to perform depth convolution operations on the image to obtain these four frequency components;

[0091] Small-size convolutions are performed on each frequency layer. Since the spatial resolution of these frequency components is relatively low, multi-scale information of the image can be captured through small convolution kernel operations without adding too many parameters. Finally, the convolution results of each frequency layer are recombined through inverse wavelet transform to form the final output. This design enables WTConv to obtain a large receptive field while maintaining a small convolution kernel, and the number of parameters only increases logarithmically, effectively avoiding the problem of over-parameterization;

[0092] The sizes and shapes of fabric defects are diverse and may be distributed at different scales. WTConv decomposes the image into multi-scale frequency components through wavelet transform and performs convolution operations on each component, which can capture the characteristic information of defects more comprehensively, thus expanding the receptive field of the model. This enables RT-DETR to detect defects of different sizes and positions more accurately, improving the detection accuracy;

[0093] Traditional convolutional layers expand the receptive field by increasing the size of the convolution kernel, which leads to a sharp increase in the number of parameters and the computational amount. While WTConv expands the receptive field, the number of parameters only increases logarithmically. Replacing the original convolution in the RT-DETR model can improve the performance of the model without significantly increasing the model parameters and computational amount, making it more suitable for real-time fabric defect detection tasks;

[0094] Low-frequency information such as the overall texture and shape of the fabric is also very important for defect detection. WTConv can perform specialized convolution processing on the low-frequency component, enhancing the model's response ability to low-frequency information. After applying WTConv in RT-DETR, the model can better utilize the low-frequency features of the fabric to assist in the localization and identification of defects, further improving the detection effect and reducing the situations of false detection and missed detection.

[0095] S4. As Figure 4 shown in the network structure, fuzzy attention is used to improve the model's adaptability to uncertain defects;

[0096] In the actual production environment, there are various complex factors on the cloth surface, such as uneven illumination, complex texture, and diverse defect morphologies. These factors increase the difficulty of defect detection. To improve the accuracy and robustness of cloth defect detection, a fuzzy attention mechanism is used to reduce the uncertainty in the feature map through fuzzy entropy and fuzzy membership function, thereby enhancing the model's attention to target features;

[0097] In cloth defect detection, each pixel point in the feature map can be regarded as belonging to two fuzzy sets: "defect" and "non - defect". By defining the fuzzy membership function, the degree to which each pixel point belongs to these two sets can be quantified. Suppose the input feature map is denoted as F∈R H×W×C , where H, W, and C represent the height, width, and number of channels of the feature map respectively. For the feature value of each pixel point (i, j) on channel c its fuzzy membership function can be expressed as:

[0098]

[0099] where α and β are learnable parameters used to adjust the shape of the fuzzy membership function to better adapt to the feature distribution of cloth defects. μ defect represents the membership degree of this pixel point belonging to the defect set.

[0100] In cloth defect detection, based on the fuzzy membership function, the fuzzy entropy of each pixel point can be further calculated to measure its uncertainty. The higher the fuzzy entropy value, the greater the uncertainty of this pixel point. To reduce the impact of this uncertainty on the detection performance, a weight is assigned to each pixel point The weight is inversely proportional to the fuzzy entropy:

[0101]

[0102] where, is the fuzzy entropy value. This weighting process enables the model to focus more on the defect features on the cloth surface while suppressing noise and background interference, thereby improving the accuracy and robustness of defect detection.

[0103] Example 2

[0104] Training and testing are carried out on the dataset, and the dataset is divided into a training set, a test set, and a validation set in the ratio of 6:2:2.

[0105] For computing resources, 16GB of memory, NVIDIA GeForce RTX 4060, and the CUDA 11.3 GPU acceleration library are used. When training the improved RT-DETR network, the training batch is set to 100 times, the training batch size is 32, the learning rate is set to 0.937, the momentum is 0.937, the optimizer is selected as ADAM, the batch size is 64, and the input feature size is set to 640×640.

[0106] In this embodiment, mAP@0.5 is selected as the model evaluation index. mAP@0.5 refers to the mAP when the IoU threshold is 50%. As shown in Table 1 below, the improved algorithm in this embodiment is higher than the RT-DETR algorithm, YOLOv5, and YOLO11 algorithms before improvement in terms of mAP@0.5 and mAP@0.5:0.95 metrics. Among them, WTConv represents wavelet convolution, and fuzzy attention represents the fuzzy attention mechanism.

[0107] Table 1

[0108]

[0109] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A cloth defect detection method based on deep learning, characterized in that: The steps include: Step 1: Collect the defect features of the target cloth and create a cloth quality inspection data set; Step 2: Using the cloth quality inspection dataset as input and the cloth defect category and location information as output, a real-time detection transformer RT-DETR convolutional neural network model is constructed and the model is trained; Step 3: Use wavelet convolution WTConv to increase the model's ability to capture small target defects; Step 4: Use fuzzy attention to improve the model's adaptability to uncertain defects.

2. The cloth defect detection method based on deep learning according to claim 1, characterized in that: Step 1 specifically includes the following steps: Step 1.1: In a textile production workshop, collect image data of several pieces of solid-color cloth, including samples without defects and samples with defects; Step 1.2: In the data preprocessing stage, the defective images with insufficient sample size are screened and eliminated; the data set is divided into four categories: the first category is significant defects in the warp direction, including warp problems such as knots and coarse warps; the second category is significant defects in the weft direction, covering weft defects such as coarse wefts and weft shrinkage; the third category is obvious stains, including pollution problems such as various stains and oil stains; the fourth category is serious defects, involving holes and jumps that seriously affect the quality of the cloth; Step 1.3: Based on the annotation information, the center point of the defect annotation box is randomly offset 104 to 312 pixels to the upper left to determine the starting coordinates of the cropping area; then, based on the coordinates, it is extended 416 pixels to the lower right to intercept a local image of 416×416 pixels; Each annotation box is randomly sampled three times to generate three new training samples; During the cropping process, if the generated image area exceeds the boundary of the original image, the sample will be discarded; Step 1.4, perform Mosaic data enhancement on the defect image; Step 1.5: Divide the data set into training set, validation set, and test set. Adjust the model parameters through the optimization algorithm to minimize the loss function. The loss function uses the cross entropy loss function combined with the positioning loss function. The specific formula is: L=αL cls +βL loc Among them, L cls is the classification loss, L loc is the positioning loss, α and β are weight coefficients used to balance the losses of classification and positioning.

3. The cloth defect detection method based on deep learning according to claim 2, characterized in that: Step 1.4 is as follows: randomly crop the four pictures and then splice them into one picture as training data; suppose the four pictures are A, B, C, and D, the sizes of the cropped pictures are A′, B′, C′, and D′, and the size of the spliced ​​pictures is E, then: In the above formula, the sizes and positions of A′, B′, C′, and D′ are determined randomly.

4. The cloth defect detection method based on deep learning according to claim 1, characterized in that: Step 2 specifically includes the following steps: Step 2.1, build and train the RT-DETR target detection framework by decomposing the multi-scale feature interaction into intra-scale interaction and cross-scale fusion; Step 2.2, backbone network: as a basic feature extractor, extract multi-scale spatial feature maps from the input image; use ResNet-18 for feature extraction, including convolutional layer, maximum pooling layer and basic module with residual structure; the convolutional layer is responsible for the preliminary extraction of image features through convolution operation, batch normalization processing and activation function application; the maximum pooling layer plays the role of downsampling to reduce the size of the feature map; in the deep feature extraction part, the network uses multiple Basicblocks to further extract complex image features, and Basicblocks at different stages have different numbers of output channels; Step 2.3, efficient hybrid encoder: The efficient hybrid encoder is used to convert the multi-scale features extracted by the backbone network into an image feature sequence suitable for Transformer processing; the encoder integrates feature information of different scales through the intra-scale feature interaction module AIFI and the cross-scale feature fusion module CCFM. The encoder calculation process is shown in the following formula: Q=K=V=Flatten(S) F = Reshape(AIFI(Q,K,V)) In the formula, S is the output feature, Q, K, and V represent different data matrices or vector sets respectively, Q is the query, which contains the information that needs to be paid attention to or queried at present, K is the key, which provides the benchmark information for comparison with Q, and V is the value, which contains the actual information content. Flatten means flattening and rearranging the S feature to adapt to the input format of Transformer, and Reshape means restoring the processed feature map to the original spatial dimension and number of channels; Step 2.4, uncertainty-minimum query selection mechanism: select a certain number of image features from the sequence output by the encoder as the initial object query of the decoder, specifically define the joint uncertainty of the encoder output features in classification and positioning prediction, and integrate this uncertainty into the loss function to guide the optimization process of the model. The selector can select the features most relevant to the target object based on the intersection-over-union (IoU) perception mechanism; Step 2.5, decoder: The decoder with auxiliary prediction head iteratively optimizes the object query to generate bounding boxes and confidence scores; The decoder processes the selected features through the self-attention mechanism and feedforward neural network layer, and finally outputs the predicted defect category and location information.

5. The method for detecting cloth defects based on deep learning according to claim 4, characterized in that: In step 2.4, the loss function includes the target confidence loss Classification Loss and bounding box regression loss Three parts; the total loss is the weighted sum of these three parts, expressed as: Among them, the weights of target confidence loss, classification loss and bounding box regression loss are λ1, λ2, and λ3 respectively; L conf and L cls The binary cross entropy function is used for calculation, and the calculation formula is as follows: L n =-(y n log(δ(x n ))+(1-y n )log(1-δ(x n ))) Where: x n To predict the score of the nth sample as a positive example, y n represents the label of the nth sample, δ represents the sigmoid function, L reg Use the CloU Loss function. CloU is a bounding box regression loss calculation method. The calculation formula is as follows: Where: parameter A represents the real box, B represents the predicted box, IoU is called intersection over union, which represents the overlap between the predicted box and the real box, ρ 2 (·) represents the Euclidean distance, b represents the coordinates of the center point of the prediction box, and b gt represents the coordinates of the center point of the real box, a represents the balance parameter, and v is the consistency of the aspect ratio of the real box and the predicted box.

6. The method for detecting cloth defects based on deep learning according to claim 1, characterized in that: Step 3 specifically includes the following steps: Step 3.1, WTConv convolution layer uses the characteristics of wavelet transform to expand the receptive field of convolutional neural network CNN; WTConv first decomposes the input image into different frequency components through wavelet transform, including low-frequency component LL, horizontal high-frequency component LH, vertical high-frequency component HL and diagonal high-frequency component HH; when using Haar wavelet transform, four filters are used to perform deep convolution operation on the image to obtain these four frequency components; Step 3.2: Perform small-size convolution on each frequency layer to capture the multi-scale information of the image through small convolution kernel operations; recombine the convolution results of each frequency layer through inverse wavelet transform to form the final output.

7. The method for detecting cloth defects based on deep learning according to claim 1, characterized in that: Step 4 specifically includes the following steps: Step 4.1: Use the fuzzy attention mechanism to reduce the uncertainty in the feature map through fuzzy entropy and fuzzy membership function, and improve the model's attention to the target features; Step 4.2: In the fabric defect detection, each pixel in the feature map is considered to belong to two fuzzy sets: "defect" and "non-defect". By defining the fuzzy membership function, the degree to which each pixel belongs to these two sets is quantified. Assume that the input feature map is represented by F∈R H×W×C , where H, W and C represent the height, width and number of channels of the feature map respectively. For each pixel (i, j) on channel c, the feature value Its fuzzy membership function is expressed as: Among them, α and β are learnable parameters used to adjust the shape of the fuzzy membership function, μ defect Indicates the degree of membership of the pixel to the defect set; Step 4.3: In the case of fabric defect detection, the fuzzy entropy of each pixel is further calculated based on the fuzzy membership function to measure its uncertainty. The higher the fuzzy entropy value, the greater the uncertainty of the pixel. To reduce the impact of this uncertainty on the detection performance, a weight is assigned to each pixel. The weight is inversely proportional to the fuzzy entropy: in, is the fuzzy entropy value, ∈ represents a very small positive number to ensure that the denominator is not zero, thereby enhancing the stability of numerical calculation.

8. A computer device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method of claim 1.

9. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to claim 1 are implemented.

10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method according to claim 1 are implemented.

Citation Information

Cited By

  • Textile defect detection method and system based on improved YOLOv13

    CN122330124A

  • A textile defect detection method and system based on improved YOLOv13

    CN122330124B