Fabric defect detection algorithm based on non-convolution step length and EMA attention mechanism
By introducing convolutional stepless and EMA attention mechanisms into the fabric defect detection algorithm, the convolution layer is optimized and the model's sensitivity to small targets is enhanced, and the problems of fine particle size loss and insufficient recognition of small targets in fabric defect detection are solved, achieving higher detection accuracy and real-time.
Patent Information
- Application Number
- CN202410069938.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-16
- Publication Date
- 2025-07-18
AI Technical Summary
The existing fabric defect detection algorithm has problems such as loss of fine particle size and poor low-resolution feature extraction during the convolution process, and is insufficiently sensitive to small target defects, which limits the accuracy and real-timeness of the model's detection of fabric defects.
The combination method of non-convolution step size and EMA attention mechanism is adopted to optimize the convolution layer through SPD-Conv, retain fine-grained information, and introduce EMA attention mechanism in the feature output layer of the backbone network to enhance the model's sensitivity to small goals.
The accuracy and real-time nature of fabric defect detection are improved, the model's feature extraction ability on low-resolution images is improved, and the ability to identify small target defects is enhanced. The detection accuracy is 2% higher than the original model, reaching mAP of 0.865.
Smart Images

Figure SMS_1 
Figure SMS_2 
Figure HDA0004669551220000011
Abstract
Description
Technical Field
[0001] The present invention relates to a deep neural network, and in particular to an algorithm that combines non-convolutional stride and EMA attention mechanism to improve the accuracy of fabric defect detection, belonging to the field of object detection. Background Art
[0002] In the field of fabric defect detection, most of the current methods rely on manual inspection. With the development of machine vision, the current fabric defect detection technology is also iteratively updated. Most of the current popular fabric defect detection algorithms have poor detection effects and weak detection capabilities for various types of fabric defects. To solve the above industry pain points, we propose a more effective fabric defect detection algorithm based on YOLOv8, which is an algorithm based on non-convolutional stride and EMA attention mechanism. To solve the problem of poor feature extraction effect of low-resolution images and loss of fine-grained information during the convolution process of the model, we use a combination of non-convolutional stride and convolution in the backbone network to improve the feature extraction ability of low-resolution images during the convolution process of the convolutional layer. This can enable the model backbone to extract more valuable detailed information and spatial information of features, helping the model effectively learn fabric defects. To solve the problem of loss of detailed information during the information transmission process between the backbone network and the neck network of the model, as well as the problem of insufficient sensitivity of the model to small targets, we add the EMA attention mechanism to the three effective feature output layers of the backbone network. The addition of the EMA attention mechanism can help the model better identify small targets and solve the problem of loss of detailed information between different levels of the network. In addition, we have effectively verified these improvements, especially by conducting experiments on the Tianchi textile dataset for YOLO series algorithms. The experimental results show that the improved YOLOv8 algorithm proposed by us performs excellently in the field of fabric defect detection.
[0003] Our research inspiration comes from the fabric defect dataset and we found that the core challenge in fabric defect detection lies in determining the accurate location of the defect, which is a prerequisite for judging the type of the defect. In addition, insufficient extraction of contour information may lead to misidentification of the defect type.
[0004] Traditional fabric defect detection mainly relies on manual inspection in front of the fabric inspection machine with the naked eye. However, such a detection method is easily interfered by various factors, such as the worker's mood, working hours, working environment, and experience. With the continuous development of science and technology, the concept of machine vision has begun to change this situation, and fabric defect detection methods based on statistics, spectrum, model, structure, and deep learning have begun to be applied in the field of fabric defect detection. Among them, the histogram method is fast but has poor classification effect; the gray-level co-occurrence matrix method has good effect but is slow; the morphological method is fast and can locate defects, but the boundaries need to be clear. Fourier transform is suitable for images with obvious frequency domain features, wavelet transform has good classification effect but poor adaptability, and Gabor has time-domain and frequency-domain locality but complex calculations. The model-based method is suitable for situations where the fabric texture is not obvious, but it cannot handle complex textures and is not suitable for the actual production process. The algorithm based on dictionary learning is complex and has a high classification accuracy. The method based on deep learning has good detection effect, but the algorithm is complex and the training takes a long time.
[0005] Although the deep learning method has some defects, these can be solved by our later algorithm optimization. Currently, it is the most commonly used method in the field of fabric defect detection based on deep learning. The deep learning method can be simply divided into single-stage algorithms and two-stage algorithms. Among them, the two-stage algorithm has a higher detection accuracy, but its detection speed has certain limitations and it is difficult to be used in actual industrial production scenarios. Compared with the two-stage detection algorithm, the single-stage detection algorithm has slightly lower detection accuracy, but its detection speed is suitable for the actual factory environment. Therefore, we choose the single-stage algorithm as our baseline model and make corresponding improvements to the defects of this model to implement the YOLOv8 model that is more suitable for fabric defect detection.
[0006] Although the current YOLOv8 model performs well in the field of object detection, in the field of fabric defect detection, the YOLOv8 detection model still has certain defects and deficiencies. First, the model will gradually downsample in the convolutional layer to extract different features of fabric defects at different resolutions. It extracts the spatial information and approximate contour features of fabric defects at high resolution, and extracts fabric detail features at low resolution. However, at low resolution, the convolution cannot extract good detail features of fabric defects, which will also cause the problem of fine-grained loss to a certain extent. Second, the correlation between the shallow and deep features of the YOLOv8 network is insufficient, which will cause some important detail features to be lost during the transmission process, affecting the detection and localization of fabric defects by the model. The model has weak sensitivity to small target defects.
[0007] In addition, the existing datasets for publicly available fabric defects are limited in number, which restricts the generalization ability of the model and the comprehensiveness of feature extraction. Therefore, the problems faced in fabric defect detection include poor feature extraction effects, loss of fine-grained information during the convolution process, poor context relevance of the network, and insensitivity of the model to small target defects. The above reasons indicate that there is still room for improvement in the field of fabric defect detection using YOLOv8. Summary of the Invention
[0008] To address the problem of low detection accuracy of the YOLOv8 model in fabric defect detection, the present invention improves the original YOLOv8 network by using the idea of combining non-convolutional strides and convolution and introducing the EMA attention mechanism. This method can effectively improve the detection accuracy of the model for fabric defects and has good real-time performance. The specific solutions of the present invention are as follows:
[0009] By fusing the algorithm of non-convolutional strides and the EMA attention mechanism, it includes the following steps:
[0010] S1. Data acquisition;
[0011] S2. Construct an improved fabric defect detection algorithm based on the YOLOv8 model;
[0012] S3. Optimize the model backbone network and network structure and train the fabric defect detection algorithm model, and save the optimal model;
[0013] S4. Use the optimal model for prediction, save the prediction results, obtain evaluation indicators, and finally conduct result comparison;
[0014] Furthermore, the experimental data in step S1 has been widely applied to industrial practical applications. The dataset we used is the 2018 Tianchi Snow Wave Manufacturing Challenge Dataset. The fabric defect data types in this dataset are more comprehensive and the annotations are more specific compared to other datasets. This database belongs to the Alibaba Tianchi dataset. It contains data on twelve typical surface defects on the cloth strip, including holes, stains, three filaments, knotted heads, hundred feet, loose warps, grains, broken warps, pulp spots, warp knots, dense files, and abrasion marks. The database collected a total of 5,913 images with fabric defects, and the original resolution of each image is 2446×1000 pixels. The intra-class defects in the database have significant differences in appearance. For example, loose warps may appear in different forms such as horizontal, vertical, or inclined. At the same time, the inter-class defects show similarities in some aspects, such as loose warps and hundred feet. In addition, affected by the shooting environment, the gray values in the images will also vary. Generally speaking, the Tianchi textile dataset presents two major challenges: the significant differences in appearance of intra-class defects and the similar characteristics of inter-class defects.
[0015] Furthermore, the content of step S2 will mainly include using the combination of no convolutional stride and ordinary convolution to optimize the convolutional layer in the backbone network. Specifically:
[0016] S21. By setting the stride of all convolutions except the first convolution in the backbone network to 1 and adding no convolutional stride after these ordinary convolutions, a new convolutional layer is formed;
[0017] S22. Propose the EMA attention mechanism to solve the problem of the loss of detailed information transmission in the network for images and improve the sensitivity of the model to small targets.
[0018] Furthermore, in step S21, the convolutional layers in the backbone network are optimized by using ordinary convolution combined with no convolutional stride. This module is called SPD-Conv, and the performance of the model has been significantly improved. SPD-Conv consists of a spatial-to-depth (SPD) layer and a no-convolutional-stride (Conv) layer. There are certain defects in the existing convolutional layer designs. The use of convolutional stride and pooling layers may cause a certain degree of loss of the network when processing fine-grained information, thereby affecting the efficiency of feature representation. This loss may lead to a decline in the performance of the network in recognition and classification tasks. Although in most scenarios, convolutional stride and pooling operations will ignore the existence of a large amount of redundant information in the convolutional layer, when performing convolution on low-resolution images, a large amount of redundant information no longer exists. At this time, using convolutional stride and pooling may cause the loss of fine-grained information, which will further lead to the model being unable to learn sufficient feature information. However, SPD-Conv can effectively solve these problems. During the downsampling process of the feature map, the SPD layer can effectively retain the information in different channel dimensions. By introducing a non-strided convolution after each SPD layer, the number of channels in the subsequent convolutional layers can be further reduced. These strategies can be used as alternatives to stride and pooling operations to achieve the downsampling of the feature map and the optimization of the number of channels. During the downsampling process of the feature map, the SPD layer not only effectively reduces the size of the feature map but also ensures that the key information in different channel dimensions is retained. To further optimize the number of channels, we introduce a non-strided convolution and place it after each SPD layer. This strategy not only reduces the number of channels in the subsequent convolutional layers but also can be used as an alternative to traditional stride and pooling operations. More importantly, this combination method not only achieves the downsampling of the feature map but also ensures that important discriminative feature information is retained.
[0019] Furthermore, in step S22, we introduce an EMA attention mechanism module. EMA is an efficient multi-scale attention mechanism designed based on the cross-space learning method and adopts a grouped structure without the need for dimensionality reduction operations. It constructs a multi-scale parallel sub-network aiming to establish short-term and long-term dependencies. Through cross-dimensional interactions, EMA retains the information on each channel. It divides the channel dimension into multiple sub-feature groups to ensure that the spatial semantic features are evenly distributed in each feature map. This feature grouping and multi-scale structure contribute to effectively establishing short-term and long-term dependencies, thereby improving the performance of the detector. In addition, EMA can effectively reduce the required number of parameters and computational overhead.
[0020] Furthermore, a fabric defect detection algorithm is constructed for model training, and then its optimal model is saved. The trained model is used to predict the test set in the dataset.
[0021] The beneficial effects of the present invention are as follows:
[0022] The fabric defect detection algorithm based on non-convolutional stride and EMA attention mechanism disclosed in the present invention. Real-time detectors of the YOLO series are favored in the industrial community because of their advantage of achieving real-time performance. In this section, we compared the YOLO series detectors of the same scale, including YOLOX(m), YOLOv5(l), YOLOv6(l), YOLOv7(l), in terms of their performance on the fabric defect detection task in the Tianchi Textile Dataset to prove the effectiveness of the proposed method in solving the fabric defect detection problem.
[0023] Generally speaking, we proposed a defect detection algorithm named SE-YOLO, which can efficiently and accurately detect fabric defects with complex and irregular shapes. By using SPD-Conv to optimize the convolutional layers in the original backbone network, the proposed algorithm can improve the feature extraction effect of the convolutional layers on low-resolution images and solve the problem of fine-grained loss. The proposed EMA module can reduce the loss of detailed information during the transmission in the backbone and neck parts of the network and improve the sensitivity of the model to small target defects. The experimental results show that the mAP of the proposed algorithm on the Tianchi Textile Dataset reaches 0.865, which is 2% higher than that of the original model and higher than the detection accuracy of the original YOLOv8 model. This indicates that the SE-YOLO model significantly improves the defect detection accuracy on the fabric defect detection dataset and meets the processing speed standard required for fabric defect detection. However, the method proposed in this study shows certain differences in the applicability of different scale models. Especially when detecting defects with irregular shapes, blurred boundaries or extreme background interference, its performance still has limitations. Future research will focus on further developing defect detection techniques, especially in refining and precisifying defect information. Image processing techniques, such as noise reduction and enhancing boundary contour features, will be widely applied. For example, using diffusion models to reduce background noise interference is expected to achieve more robust defect detection. Therefore, the focus of future research will be on expanding the network structure to cover a wider range of defect detection scenarios to meet the diverse actual application needs.
[0024] Other advantages, objectives and features of the present invention will be elaborated in detail in the following specification. In the specification, these advantages and features will be described in detail to make them more understandable to those skilled in the art and non-related personnel, or they can be obtained from the practice of the present invention. The specification will introduce in depth how to achieve and obtain the objectives and other advantages of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be described in detail preferably with reference to the accompanying drawings, where:
[0026] Figure 1Schematic diagram of the improved overall architecture of YOLOv8;
[0027] Figure 2 Schematic diagram of the SPD-Conv module;
[0028] Figure 3 Schematic diagram of the EMA module; Specific implementation manner
[0029] The following is a further detailed description through specific implementation manners.
[0030] To solve the problem of low detection accuracy in the field of fabric defect detection, the present invention uses a convolution-free stride and an EMA attention mechanism. This method can effectively improve the detection accuracy of fabric defects and has good performance in real-time detection of fabric defects. The specific solution of the present invention is as follows:
[0031] A fabric defect detection algorithm based on a convolution-free stride and an EMA attention mechanism includes the following steps:
[0032] S1. Data acquisition;
[0033] The experimental data in step S1 has been widely applied in industrial actual applications. The dataset we used is the 2018 Tianchi Xuelang Manufacturing Challenge dataset. The fabric defect data types on this dataset are more comprehensive and the annotations are more specific compared to other datasets. This database belongs to the Alibaba Tianchi dataset. It contains data on twelve typical surface defects on the cloth strip, including holes, stains, three filaments, knotted heads, hundred feet, loose warps, grains, broken warps, pulp spots, warp knots, dense files, and abrasion marks. The database contains a total of 5913 images with fabric defects, and the original resolution of each image is 2446×1000 pixels. The intra-class defects in the database have significant differences in appearance. For example, loose warps may present different forms such as horizontal and vertical. At the same time, the inter-class defects show similarities in some aspects, such as loose warps and hundred feet. In addition, affected by the shooting environment, the gray values in the images will also be different. Generally speaking, the Tianchi textile dataset presents two major challenges: the significant differences in appearance of intra-class defects and the similar characteristics of inter-class defects.
[0034] S2. Construct an improved fabric defect detection algorithm based on the YOLOv8 model;
[0035] S21. By setting the stride of all convolutions in the backbone network except the first convolution to 1 and adding non-convolutional strides after these ordinary convolutions, a new convolutional layer is formed;
[0036] In the original convolutional neural network, three layers of convolution (convolutional layer, normalization layer, activation layer) and pooling layer are usually used for feature extraction. However, this method will cause a certain degree of loss of fine-grained information and will also limit the model's ability to learn the features of fabric defects to a certain extent. Using SPD-Conv can effectively solve this problem. SPD-Conv enables the model to effectively retain learnable information during downsampling by eliminating convolutional strides and pooling layers.
[0037] The SPD component uses an original image transformation technique to perform downsampling of feature maps inside and across the entire convolutional layer. It divides a feature map of any size M×M×C1 into several sub-feature map slices. The subsequence slices are:
[0038]
[0039] The feature map X(i, j) is composed of all mapped sub-maps fx, y, where i + x and j + y are divisible. Therefore, each sub-map can downsample the original feature map X in a certain proportion. We choose the SPD operation when scale = 2. Specifically, the SPD operation divides the feature map M of size N×N×C1 into four groups of sub-feature maps. The size of each sub-feature map is N / 2×N / 2×C1, which is obtained by sampling the feature map M by reducing its width and height by half. The specific implementation is to divide the feature map M according to a 2×2 grid size, extract the feature vectors at each position of (0, 0), (1, 0), (0, 1), and (1, 1), combine them into four sub-feature maps respectively, and connect them in sequence along the channel dimension. Such an operation converts the feature map M (of size N×N×C1) into an intermediate feature map (of size N / 2×N / 2×2C1).
[0040] The intermediate feature map generated by the SPD operation is connected to a convolutional layer with C2 1×1 convolutional kernels attached, with a stride of 1, where C2 is less than 2C1. Using non-strided (stride = 1) 1×1 convolutional kernels helps to maximize the retention of detailed information in the feature map, which is consistent with the requirements of the fabric defect detection task.
[0041] S22. Propose the EMA attention mechanism to solve the problem of loss of detailed information transmission in the network for images and improve the sensitivity of the model to small targets.
[0042] First, we found that the three effective feature output layers of the model backbone network would be directly input into different parts of the neck network. We considered that such direct transmission might cause a certain degree of loss of some key information during the transmission process, resulting in insufficient feature fusion between the neck network and the backbone network. This would prevent the neck network from fully obtaining the key information transmitted by the backbone network, leading to a discount in the final effect of the model. Moreover, fabric defects not only include larger-sized defects but also small target defects such as knots and warping knots. However, the original YOLOv8 model is not sensitive enough to small target defects, making it difficult to identify these defects.
[0043] The introduction of the EMA attention mechanism can effectively solve the above problems. The concept of the attention mechanism is inspired by human vision. When we observe an object with our naked eyes, we tend to focus more on the key parts of the object. The EMA attention mechanism is a new efficient multi-scale attention mechanism based on cross-space learning. It adopts a grouped structure without the need for dimensionality reduction operations. It constructs a multi-scale parallel sub-network aiming to establish short-term and long-term dependencies. Through cross-dimensional interactions, the EMA retains the information on each channel. It divides the channel dimension into multiple sub-feature groups to ensure that spatial semantic features are evenly distributed in each feature map. This feature grouping and multi-scale structure help to effectively establish short-term and long-term dependencies, thus improving the performance of the detector. In addition, the EMA can effectively reduce the required number of parameters and computational overhead.
[0044] After we added the EMA attention mechanism after the three effective output layers of the backbone network, this position is also the part where the backbone network and the neck network perform feature fusion, belonging to the feature enhancement part of the network. We added the EMA attention mechanism to the feature enhancement part because this can enable higher-level semantic information to better guide the model to learn the defect positions and can focus more attention on the fabric defect part. Since there is less interference information in the higher-level semantic information, in the field of fabric defect detection, because some defects highly coincide with the background information of the fabric itself, adding the EMA attention mechanism to the high-level part of the network is more conducive to the model focusing on the fabric defects themselves and avoiding the interference of irrelevant information.
[0045] S3. Optimize the model backbone and network structure and train the object detection algorithm model to save the optimal model; in order to verify whether the improvement of the original YOLOv8 network is effective.
[0046] This paper conducted ablation experiments on the Tianchi textile dataset. There are a total of 3 groups of experiments, and the experimental results are shown in the table.
[0047] Table 1 Comparison of experimental effects of different modules added to the YOLOv8 model
[0048]
[0049] As can be seen from Table 1, after adding different modules, each module has improved the model to a certain extent, especially the significant improvement in mAP50.
[0050] Compared with the original YOLOv8 network model, all AP metrics of the SE-YOLO model integrated with the SPD-Conv and EMA attention mechanisms are significantly better than the baseline model, which proves that the modules we added help to improve the detection performance of the model. When the SPD-Conv module is added to the YOLOv8 network model, both mAP50 and mAP50-95 are significantly improved, indicating that using SPD-Conv to improve the convolutional layer of the network can effectively enhance the feature extraction ability of the model on low-resolution images. After adding the EMA attention mechanism, the AP50-95 of the model also has a certain improvement, and the AP value of small target defects also increases, proving that adding the EMA attention mechanism to the high-level part of the network is beneficial for the model to learn the defect part of the model and is more conducive to the model to strengthen feature extraction.
[0051] S4. Use the optimal model for prediction, save the prediction results, obtain the evaluation metrics, and finally conduct a result comparison; Through experiments, we can see that the detection effect of our optimal model is better than the baseline model, indicating that our experiment is effective.
Claims
1. Fabric defect detection algorithm based on no convolution stride and EMA attention mechanism, the method comprising: S1. Construct a newly developed object detection algorithm based on the YOLOv8 model; S2. Optimize the model backbone and network structure and train the object detection algorithm model, and save the optimal model; S3. Use the optimal model for prediction, save the prediction results, obtain evaluation metrics, and finally conduct result comparison.
2. The fabric defect detection algorithm based on the convolutional step size and EMA attention mechanism according to claim 1, characterized in that: The content of step 2 will mainly include SPD-Conv to optimize the convolutional layer of the backbone network and the application of the EMA module. Specifically: S11. By setting the stride of all convolutions except the first convolution in the backbone network to 1, add no convolution stride after these ordinary convolutions to form a new convolutional layer; S12. Propose the EMA attention mechanism to solve the problem of loss of detailed information transmission in the network in the image and improve the sensitivity of the model to small targets.
3. The fabric defect detection algorithm based on the convolution-free stride and EMA attention mechanism according to claim 1, wherein: After the convolutional layer is optimized by SPD-Conv in step S11, the problems of loss of fine-grained information in the network and poor feature extraction under low-resolution images can be solved. The SPD layer can retain information in different channel dimensions while downsampling the feature map. After that, add no-stride convolution after each SPD layer, which can reduce the number of channels in the added convolutional layer. These operations can replace the stride and pooling operations and can downsample the feature map while retaining discriminative feature information.
4. The fabric defect detection algorithm based on the convolution-free stride and EMA attention mechanism according to claim 2, characterized in that: In step S12, the EMA module constructs a multi-scale parallel sub-network aiming to establish short-term and long-term dependencies. Through cross-dimensional interactions, EMA retains the information on each channel. It divides the channel dimension into multiple sub-feature groups to ensure that spatial semantic features are evenly distributed in each feature map. This feature grouping and multi-scale structure helps to effectively establish short-term and long-term dependencies, thereby improving the performance of the detector. In addition, EMA can effectively reduce the required number of parameters and computational overhead.