Light-weight fabric defect detection algorithm based on deformable convolution and high-level screening feature pyramid
By fusing deformable convolution, GS convolution and high-level screening feature pyramids in the fabric defect detection model, the existing models have fewer detection categories, incomplete feature extraction and weak model detection capabilities, achieving higher detection accuracy and lower computing resource consumption, which is suitable for industrial real-time detection needs.
Patent Information
- Application Number
- CN202411106091.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-05-09
AI Technical Summary
The existing fabric defect detection models have problems such as few detection categories, incomplete feature extraction and weak model detection capabilities, and the model parameters and calculation volume are large, making it difficult to meet the needs of industrial real-time detection.
By fusing deformable convolution, GS convolution and high-level screening feature pyramids, the backbone network and SPPF modules are optimized, and the model's feature extraction capability and multi-scale detection capability are improved, while reducing the model's parameter quantity and calculation quantity.
The accuracy and speed of fabric defect detection have been improved. The mAP of the model on the Tianchi textile data set has been increased by 1.7%, and the parameter quantity and calculation quantity have decreased by 37.5% and 21% respectively, making it suitable for industrial real-time detection applications.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] The present invention relates to a target detection model, and in particular to an algorithm for fusing deformable convolution, GS convolution and high-level screening feature pyramid to improve the detection accuracy of cloth surface defects and reduce the number of model parameters, belonging to the field of defect detection. Background Art
[0002] Fabric defect detection is the key to quality control in the fabric manufacturing process and is crucial to improving quality. However, the multi-scale characteristics of defects and their high integration with the fabric background make detection extremely challenging. Fabric defect detection models still have problems such as few detection types, low detection accuracy, and high model parameter and computational complexity. This paper proposes a lightweight fabric defect detection model YOLO-DGH. First, the proposed model incorporates deformable convolution into the backbone network, and its dynamic receptive field characteristics make the model more accurate in extracting fabric defect features and effectively improve network performance. Secondly, the proposed model uses the GS-SPPF module, which optimizes SPPF by using the large kernel characteristics of GSConv, so that GS-SPPF can capture a larger global receptive field and more effectively extract the features of fabric defects while reducing the number of parameters and computational complexity. Finally, the proposed model introduces the idea of lightweight high-level screening feature pyramids, which achieves multi-scale feature fusion while maintaining a low number of parameters, enhancing the model's ability to detect multi-scale fabric defects. Experimental results show that the mAP of the YOLO-DGH model on the Tianchi textile dataset reached 0.848, an increase of 1.7% compared to the original model. In addition, the number of parameters and the amount of calculation of the model decreased by 37.5% and 21% respectively, which shows the effectiveness of the model in fabric defect detection.
[0003] The inspiration for our research comes from the core challenge of detecting fabric defects in practice, which is how to accurately detect defects and have a faster detection speed, and the model should not have large differences in the detection effect of various defects. The main reasons for the poor detection effect, slow detection speed and high false detection rate of the model are the poor feature extraction ability of the model, the high number of model parameters and calculations, and the similarities between defect classes. How to solve these problems is the goal of algorithm improvement.
[0004] Fabric defect detection technology occupies an important position in the quality monitoring system of the textile industry. It can effectively identify defects on fabrics, thereby ensuring the stability of textile quality. At present, fabric defect detection technology is mainly divided into two categories: traditional detection methods and deep learning-based detection methods. Traditional fabric defect detection methods mainly include manual detection and machine vision detection. Manual detection mainly relies on the naked eye to examine the fabrics one by one. However, this method has significant limitations, such as low detection efficiency, difficulty in identifying small defects, and high experience requirements for detection personnel. Machine vision detection analyzes the texture features and defect characteristics of the fabric and combines image processing technology to detect and locate defects. However, machine vision detection methods are still insufficient in detection categories and generalization capabilities. .
[0005] With the continuous development of deep learning technology, fabric defect detection methods based on deep learning have gradually emerged. These methods can be further divided into one-stage algorithms and two-stage algorithms. Two-stage algorithms usually first extract candidate target areas through selective search, and then classify these areas to determine the precise target location. However, common two-stage detectors such as RCNN, Fast R-CNN and Faster R-CNN often face the problem of slow detection speed when dealing with fabric defect detection, which is difficult to meet the needs of industrial real-time detection.
[0006] In contrast, the one-stage algorithm highly integrates feature extraction, position information regression and classification prediction, so it has a higher detection speed and is more suitable for real-time detection requirements in industrial sites. Among the one-stage algorithms, the YOLO series has gradually become the mainstream model in the field of fabric defect detection with its excellent performance. However, although YOLO and its improved algorithms continue to improve their performance in iterative updates, there are still some problems. At present, most networks can only detect a limited number of fabric defects, and the model parameters are large, and there are significant differences in the detection capabilities of different defects. In addition, due to the complex shape rules of fabric defects, morphological differences between categories, and the scarcity of data sets for some defects, the imbalance of the model's detection capabilities for different types of defects is further aggravated. Therefore, the field of fabric defect detection still faces problems such as a small number of model detection categories, incomplete feature extraction, and weak model detection capabilities, making it difficult to effectively apply to industrial real-time detection scenarios. .
[0007] In addition, the existing public fabric defect datasets are limited in number and variety, which limits the generalization ability of the model and the comprehensiveness of feature extraction. Therefore, the problems faced by fabric defect detection include insufficient model generalization ability and incomplete feature extraction. These factors lead to low accuracy of fabric defect detection and make it difficult to effectively apply it to real-time industrial detection scenarios. Summary of the invention
[0008] In order to solve the problem of low accuracy in real-time detection of fabric defects in target detection, the present invention integrates deformable convolution, GS convolution and high-level screening feature pyramid. The method can effectively solve the problems of small number of model detection categories, incomplete feature extraction and weak model detection ability, and at the same time reduce the number of model parameters and calculation amount, so that it has good performance in real-time detection of fabric defects. The specific scheme of the present invention is as follows:
[0009] By integrating deformable convolution, GS convolution and high-level screening feature pyramid, the fabric defect detection algorithm is shown in the overall structure of the model. Figure 1 As shown, the specific implementation scheme includes the following steps:
[0010] S1, data acquisition;
[0011] S2. Build a newly developed fabric defect detection algorithm based on the YOLOv8m model;
[0012] S3, optimizing the model backbone and network structure, training the fabric defect detection model, and saving the optimal model;
[0013] S4, use the optimal model to make predictions, save the prediction results, obtain evaluation indicators, and finally compare the results;
[0014] Furthermore, the experimental data in step S1 comes from the Alibaba Cloud Tianchi Guangdong Industrial Manufacturing Competition, which contains 5096 defective images with a resolution of 2446×1000. Due to the high-resolution characteristics of the data images, it is particularly difficult for the model to extract features of small target defects under high resolution. Therefore, we use the sliding window technology to crop the original data set, and the cropping step is set to 600 pixels. The size of the cropped images is unified to 640×640 pixels. After cropping, the data set contains a total of 14269 images. To ensure the comprehensiveness and accuracy of the model training, we carefully divide the data set into training set, test set and validation set in a ratio of 8:1:1. This strategy not only improves the model's ability to extract features of small target defects, but also lays a solid foundation for subsequent model training and verification work.
[0015] Furthermore, step S2 will mainly include the use of deformable convolution, the design of GS-SPPF structure, and the design of high-level screening feature pyramid network. Specifically:
[0016] S21. Use deformable convolution to optimize the convolution layer. Since the shape of fabric defects is irregular, ordinary convolution is difficult to cover the entire defect for feature extraction, resulting in poor model feature extraction capability. Deformable convolution can better fit fabric defects, thereby better extracting fabric defect features.
[0017] S22. Construct the GS-SPPF module. The auxiliary convolution branch (deep convolution) in GSConv uses a larger convolution kernel (usually 5×5) to directly make the shallow network have a larger global receptive field, while the lightweight model often has fewer layers, resulting in a smaller receptive field. In conventional models, GSConv can be used to reduce computational complexity while maintaining or optimizing accuracy;
[0018] S23. The high-level screening feature pyramid idea is used to optimize FPN. Different from the FPN used in the YOLO series, the high-level screening feature pyramid takes into account the multi-scale characteristics of fabric defects, embeds multiple CA modules and feature transposition modules in the neck part for deep feature screening and fusion, and discards unnecessary modules in the original FPN to achieve lightweight detection model, which significantly improves the model's feature expression ability on fabric defect datasets.
[0019] Furthermore, in step S21, we improve the performance of the model by optimizing the convolution layer using deformable convolution. In deformable convolution, the sampling process is no longer restricted to a fixed grid layout. Figure 2 As shown. The deformable convolution preprocesses the input features through the convolution operation to generate offsets and modulations for adjusting the sampling position. Subsequently, these offsets are used to convert the originally regularly distributed 3×3 pixel points into irregularly and arbitrarily distributed sampling points. These newly generated sampling points can more accurately reflect the shape and size of the defects in the input image. Next, the input image is sampled using these sampling points to obtain the sampled feature map. Finally, the sampled feature map is element-wise multiplied by the convolution kernel and summed to obtain the final convolution result. This process enables the deformable convolution to more accurately extract feature information when dealing with complex and changeable defect morphologies, thereby improving the accuracy and robustness of feature representation.
[0020] Furthermore, in the application scenario of lightweight models in step S22, GSConv shows better performance than standard convolution. This is mainly due to the use of the auxiliary convolution branch in GSConv - deep convolution. Deep convolution uses a larger convolution kernel, usually 5×5. This design enables the network to have a larger global receptive field at the shallow stage. Since lightweight models usually have fewer layers, their receptive field is relatively limited. Therefore, GSConv effectively makes up for the shortcomings of lightweight models in feature extraction by expanding the receptive field. GSConv is a lightweight and efficient convolution operation. Its core is to maintain the implicit connection between channels in high-level networks as much as possible while reducing time complexity, so as to maintain the integrity of features and reduce the loss of semantic information. Compared with traditional standard convolution techniques, GSConv maintains the ability to capture global information while utilizing the characteristics of global sparsity to perform convolution calculations specifically for sparse areas. This method significantly reduces the computational cost and the burden of model parameters. As Figure 3 As shown in , GSConv uses shuffle operation to permeate the features extracted from standard convolution into each part of the features extracted by depthwise separable convolution, and allows the information from SC to be fully mixed into the output of DSC by evenly exchanging local feature information on different channels without adding additional computational burden. In the SPPF structure, we integrate GSConv into the SPPF module and propose a new GS-SPPF module. Specifically, we replace the 1×1 standard convolution feature generation layer in SPPF with GSConv, as shown in Figure 4 As shown in the figure, the original 1×1 standard convolution is replaced by GSConv, which is a combination of 3×3 standard convolution and 5×5 deep convolution. GSConv combines the advantages of channel-dense convolution and depth-separable convolution, retains the hidden connections between channels to the maximum extent, reduces the loss of semantic information in the feature conversion process, and effectively improves the model's ability to extract subtle defect features.
[0021] Further, in step S23, we introduce a high-level screening feature pyramid network to achieve multi-scale feature fusion. In the task of fabric defect detection, fabric defects have a multi-scale problem, which affects the model's ability to identify fabric defects. This effect occurs because different types of defects usually have differences in size, and the sizes of the same type of defects may also be different. The introduction of a high-level screening feature pyramid enables the model to effectively capture a more comprehensive array of feature information of fabric defects. The high-level screening feature pyramid mainly consists of two parts: a feature selection module and a feature fusion module. Initially, feature maps of different scales undergo a screening process in the feature selection module, which extracts representative information of each channel of the feature map through the CA attention mechanism while minimizing information loss. Subsequently, the filtered feature map is generated by multiplying the weight information and the corresponding scale feature map, and then the feature map of each scale is processed and input into the subsequent fusion module using a 1×1 convolution. The feature selection module strategically fuses features by using high-level features as weights to filter the basic semantic information embedded in low-level features. The feature selection module expands the high-level features through transposed convolution, and then uses bilinear interpolation to upsample or downsample the high-level features. Then, the CA module is used to convert the high-level features into corresponding attention weights to filter low-level features to obtain features with consistent dimensions. Finally, the filtered low-level features are fused with high-level features to enhance the feature representation of the model.
[0022] Furthermore, a fabric defect detection algorithm is constructed and the data set is loaded for training. Then, the optimal model is saved and the trained model is used to predict the test set of the data.
[0023] The beneficial effects of the present invention are:
[0024] The fabric defect detection algorithm based on deformable convolution, GS convolution and high-level screening feature pyramid disclosed in the present invention. In order to prove the effectiveness of the proposed method in fabric defect detection. In this section, we compare the YOLO series detectors of the same scale on the same dataset, including YOLOX(m), YOLOv5(m), YOLOv6(l), YOLOv7(t) and other mainstream fabric defect detection algorithms.
[0025] In summary, we proposed a fabric defect detection algorithm named YOLO-DGH, which can efficiently and accurately detect fabric defects with complex and irregular shapes and can effectively reduce the number of model parameters. The proposed model incorporates deformable convolution into the backbone network, and its dynamic receptive field characteristics make the model more accurate in extracting fabric defect features and effectively improve network performance. Secondly, the proposed model uses the GS-SPPF module, which optimizes SPPF by using the large kernel characteristics of GSConv, so that GS-SPPF can capture a larger global receptive field and more effectively extract the features of fabric defects while reducing the number of parameters and calculations. Finally, the proposed model introduces the idea of lightweight high-level screening feature pyramid, which realizes multi-scale feature fusion while maintaining a low number of parameters, and enhances the model's ability to detect multi-scale fabric defects. Experimental results show that the mAP of the YOLO-DGH model on the Tianchi textile dataset reaches 0.848, which is 1.7% higher than the original model. In addition, the number of parameters and the amount of computation of the model decreased by 37.5% and 21% respectively, which demonstrates the effectiveness of the model in fabric defect detection.
[0026] Other advantages, objectives and features of the present invention will be described in the following description to some extent, and to some extent, will be obvious to those skilled in the art based on the following examination and study, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below in conjunction with the accompanying drawings, wherein:
[0028] Figure 1 Schematic diagram of the improved YOLOv8m overall architecture;
[0029] Figure 2 Schematic diagram of the deformable convolution module;
[0030] Figure 3 GSConv structure diagram
[0031] Figure 4 Schematic diagram of the GS-SPPF module. DETAILED DESCRIPTION
[0032] The following is further explained in detail through specific implementation methods.
[0033] In order to solve the problem of low accuracy and slow speed in real-time detection of fabric defects in defect detection, the present invention integrates deformable convolution, GS convolution and high-level screening feature pyramid network to improve YOLOv8. The method can effectively improve the detection accuracy of fabric defects and has good real-time detection performance. The specific scheme of the present invention is as follows:
[0034] The fabric defect detection algorithm integrating deformable convolution, GS convolution and high-level screening feature pyramid network includes the following steps:
[0035] S1, data acquisition;
[0036] The experimental data in step S1 comes from the Alibaba Cloud Tianchi Guangdong Industrial Manufacturing Competition, which contains 5096 defective images with a resolution of 2446×1000. Due to the high-resolution characteristics of the data images, it is particularly difficult to extract features of some small target defects in high-definition state. Therefore, we use the sliding window technology to crop the original data set, and the cropping step is set to 100 pixels. The size of the cropped image is unified to 640×640 pixels. After cropping, the data set contains a total of 14269 images. To ensure the comprehensiveness and accuracy of the model training, we carefully divide the data set into training set, test set and validation set in a ratio of 14:3:3. This strategy not only improves the model's ability to extract features of small target defects, but also lays a solid foundation for subsequent model training and verification work.
[0037] S2, build a fabric defect algorithm based on the YOLOv8m model;
[0038] S21. Use deformable convolution to optimize the convolution layer. Since the shape of fabric defects is irregular, ordinary convolution is difficult to cover the entire defect for feature extraction, resulting in poor model feature extraction capability. Deformable convolution can better fit fabric defects, thereby better extracting fabric defect features.
[0039] When extracting features, traditional convolution operations usually slide the convolution kernel at a fixed grid position for the input feature map, which can effectively extract features for objects with regular shapes. However, when dealing with various defects with large size differences and complex deformations, traditional convolution methods are limited by fixed sampling patterns and it is difficult to fully extract the feature information of these irregular objects.
[0040] In contrast, Deformable Convolutional Networks (DCN) introduces an offset mechanism to achieve free adjustment of the sampling position. For defects with different positions, scales, and deformation characteristics, DCN can automatically adjust the distribution and scale of the sampling points, making the sampling process closer to the actual shape and size of the object.
[0041] In DCN, the sampling process is no longer restricted to a fixed grid layout. Figure 2 As shown in the figure, the input features are preprocessed through the convolution operation to generate offsets and modulations for adjusting the sampling position. Subsequently, these offsets are used to convert the originally regularly distributed 3×3 pixels into irregularly and arbitrarily distributed sampling points. These newly generated sampling points can more accurately reflect the shape and size of the defects in the input image. Next, the input image is sampled using these sampling points to obtain the sampled feature map. Finally, the sampled feature map is element-wise multiplied by the convolution kernel and summed to obtain the final convolution result. This process enables deformable convolution to more accurately extract feature information when dealing with complex and changeable defect morphologies, thereby improving the accuracy and robustness of feature representation.
[0042] S22. Construct the GS-SPPF module. The auxiliary convolution branch (deep convolution) in GSConv uses a larger convolution kernel (usually 5×5) to directly make the shallow network have a larger global receptive field, while lightweight models often have fewer layers, resulting in a smaller receptive field. In conventional models, GSConv can be used to reduce computational complexity while maintaining or optimizing accuracy.
[0043] GSConv is a lightweight convolution operation, the core of which is to introduce the concept of global sparsity to efficiently integrate global information. Compared with traditional standard convolution technology, GSConv maintains the ability to capture global information while taking advantage of the characteristics of global sparsity to perform convolution calculations specifically for sparse areas. This method significantly reduces the computational cost and the burden of model parameters. GSConv first performs shuffle processing, and then efficiently fuses the information generated by dense convolution and depthwise separable convolution. In addition, GSConv uses sparsity masking technology to ensure precise control of the sparsity level at each position, thereby further optimizing the use of computing resources. The GSConv structure diagram is shown below. Figure 3 shown.
[0044] In the application scenarios of lightweight models, GSConv shows better performance than standard convolution. This is mainly due to the use of the auxiliary convolution branch in GSConv - deep convolution. Deep convolution uses a larger convolution kernel, usually 5×5. This design enables the network to have a larger global receptive field in the shallow stage. Since lightweight models usually have fewer layers, their receptive field is relatively limited. Therefore, GSConv effectively makes up for the shortcomings of lightweight models in feature extraction by expanding the receptive field. In the SPPF structure, we specially introduced the GSConv lightweight convolution technology. Specifically, we replaced the 1×1 standard convolution feature generation layer in SPPF. As Figure 4As shown in the figure, the original 1×1 standard convolution is replaced by GSConv with 3×3 standard convolution and 5×5 depth convolution. This change not only enables the SPPF module to directly obtain a larger local receptive field, thereby improving the detection accuracy, but also effectively reduces the amount of calculation at this stage through the sparsity characteristics of GSConv.
[0045] S23. The high-level screening feature pyramid idea is used to optimize FPN. Different from the FPN used in the YOLO series, the high-level screening feature pyramid takes into account the multi-scale characteristics of fabric defects, embeds multiple CA modules and feature transposition modules in the neck part for deep feature screening and fusion, and discards unnecessary modules in the original FPN to achieve lightweight detection model, which significantly improves the model's feature expression ability on fabric defect datasets.
[0046] In the task of fabric defect detection, fabric defects have a multi-scale problem, which affects the model's ability to identify fabric defects. This effect occurs because different types of defects usually have different sizes, and the sizes of the same type of defects may also be different. In order to solve the inherent multi-scale problem in the fabric defect dataset, we introduced a high-level screening feature pyramid screening network (high-level screening feature pyramid) to achieve multi-scale feature fusion. The introduction of the high-level screening feature pyramid enables the model to effectively capture a more comprehensive array of feature information of fabric defects. The high-level screening feature pyramid mainly consists of two parts: feature selection module and feature fusion module. Initially, feature maps of different scales undergo a screening process in the feature selection module. The feature selection module extracts representative information of each channel of the feature map through the CA attention mechanism while minimizing the loss of information. Subsequently, the filtered feature map is generated by multiplying the weight information and the corresponding scale feature map, and then the feature map of each scale is processed and input into the subsequent fusion module using 1×1 convolution. The feature selection module strategically fuses features by using high-level features as weights to filter the basic semantic information embedded in low-level features. The feature selection module expands the high-level features through transposed convolution, and then uses bilinear interpolation to upsample or downsample the high-level features. Then, the CA module is used to convert the high-level features into corresponding attention weights to filter low-level features to obtain features with consistent dimensions. Finally, the filtered low-level features are fused with high-level features to enhance the feature representation of the model.
[0047] S3, optimize the model backbone and network structure and train the target detection algorithm model, and save the optimal model;
[0048] In order to verify whether the improvement of the baseline YOLOv8m network is effective, this paper conducted an ablation experiment on the Tianchi textile dataset. There are 4 groups of experiments in total, and the experimental results are shown in the table.
[0049] Table 1 Comparison of experimental results of different modules added to the v8m model
[0050]
[0051] It can be seen from Table 1 that the improved modules have improved the detection accuracy based on the original network, especially the accuracy of AP50 has been significantly improved.
[0052] When deformable convolution is used to improve the last two convolution layers of the backbone, the mAP value of the model is improved by 0.9%. This shows that we introduced deformable convolution in the high-level part of the network to solve the problem that the original convolution module cannot well extract the features of defects with large deformation and irregular patterns. This also shows that the ability of deformable convolution to dynamically adjust the scale and receptive field is more suitable for the model to extract fabric defect features.
[0053] When combining the deformable convolution module and the GS-SPPF module, the model has a mAP improvement of 1.6% relative to the Baseline, a parameter reduction of 0.2M, and a computational cost reduction of 3.4G. Compared with YOLOv8m with the DCN module added, the mAP is improved by 0.7%, the parameter reduction of 0.4M, and the computational cost reduction of 0.2. Because GSConv often has better effects than standard convolution in lightweight models, we use GSConv with 3×3 standard convolution and 5×5 deep convolution to replace the original 1×1 standard convolution in SPPF, so that the SPPF module can directly obtain a larger local receptive field to improve detection accuracy and reduce the computational cost of this stage.
[0054] On the basis of combining the deformable convolution module and the GS-SPPF module, when the high-level screening feature pyramid with lightweight multi-scale ideas is introduced, the entire model achieves the best effect. Compared with the Baseline, the mAP value of the model is improved by 1.7%, the number of parameters is reduced by 9.7M, and the amount of calculation is reduced by 16.6G. This is because the high-level screening feature pyramid initially screens feature maps of different scales in the feature selection module. Subsequently, the high-level and low-level information in these feature maps are collaboratively integrated through the selective feature fusion (SFF) mechanism. This fusion produces features with rich semantic content, which helps to detect subtle features in fabric defect images, thereby enhancing the detection ability of the model. In addition, HSFPN abandons the original module with high computational cost and retains the lightweight enhanced feature extraction module, so that the model can still maintain good detection capabilities with low parameter and low computation.
[0055] Table 2 Comparison of experimental results of different YOLO baseline models
[0056]
[0057] The YOLO series of object detection algorithms are widely praised in the industry for their excellent real-time performance and reliability. This paper aims to demonstrate the efficacy of the YOLO-DGH model in detecting fabric defects by comparing it with different sizes of YOLO series detectors (i.e., YOLOv5, YOLOv6, YOLOv7-l, and YOLOv8-m) on the Tianchi fabric dataset. The experimental data show that the YOLO-DGH model performs well in fabric defect detection, with an mAP value of 0.851, a model parameter number of only 16.5M, an FPS of 107, and a computational cost of only 62.4G. In addition, Table 2 shows the test data results of these detectors in detail, thus confirming the significant advantages of the YOLO-DGH model in the field of fabric defect detection.
[0058] Table 3 Comparison of experimental results of SOTA target detection models
[0059]
[0060] Table 4 Comparison with SOTA defect detectors on selected categories of Tianchi fabric dataset
[0061]
[0062] To demonstrate the superior performance of YOLO-DGH in fabric defect detection, this paper conducts a comparative analysis with existing advanced methods. Table 3 shows that our YOLO-DGH model has a significant advantage when the number of detection categories is the same. Compared with the YOLO-SCD model, our model has fewer parameters and higher detection accuracy. This shows that our model reduces the consumption of computing resources while maintaining high performance, and is more suitable for practical application scenarios. In addition, as shown in Table 4, for some specific categories of defects, our YOLO-DGH model is significantly better than the YOLO-SCD model. In addition, our model shows balanced detection performance in different defect categories, so it is suitable for wide application in the industrial field.
[0063] In addition, we also compare the YOLO-DGH model with Liu's model. Table 3 shows that our model outperforms Liu's model in both detection accuracy and frames per second (FPS) when the number of categories increases significantly. This shows that our model performs well in complex scenes. Table 4 shows that our YOLO-DGH model outperforms Liu's model in both detail capture and feature extraction when detecting the same categories. Even compared with other models with a larger number of detection categories, our YOLO-DGH model can maintain good FPS while greatly exceeding the detection accuracy of the FD-YOLOv5 model. This emphasizes that our model shows good real-time performance while maintaining high detection accuracy.
[0064] In conclusion, after comparative analysis with current state-of-the-art methods, it can be concluded that the YOLO-DGH model has a clear advantage in detecting fabric defects. This is evident in its parameters, detection accuracy, and real-time performance, making it more suitable for practical industrial applications.
[0065] S4. Use the optimal model to make predictions, save the prediction results, obtain evaluation indicators, and finally compare the results. After experiments, we can see that the detection effect of our optimal model is better than some baseline models and SOTA models, indicating that our experiment is effective.
Claims
1. A lightweight fabric defect detection algorithm based on deformable convolution and high-level screening feature pyramid, the method comprising: S1, data acquisition; S2. Build a newly developed fabric defect detection algorithm based on the YOLOv8m model; S3, optimize the model backbone and network structure, train the target detection algorithm model, and save the optimal model; S4. Use the optimal model to make predictions, save the prediction results, obtain evaluation indicators, and finally compare the results.
2. The lightweight fabric defect detection algorithm based on deformable convolution and high-level screening feature pyramid according to claim 1 is characterized in that: The experimental data in step S1 comes from the Alibaba Cloud Tianchi Guangdong Industrial Manufacturing Competition, which contains 5096 defective images with a resolution of 2446×1000. Due to the high-resolution characteristics of the data images, it is particularly difficult for the model to extract features of small target defects under high resolution. Therefore, we use the sliding window technology to crop the original data set, and the cropping step is set to 600 pixels. The size of the cropped image is unified to 640×640 pixels. After cropping, the data set contains a total of 14269 images. To ensure the comprehensiveness and accuracy of the model training, we carefully divide the data set into training set, test set and validation set in a ratio of 8:1:
1. This strategy not only improves the model's ability to extract features of small target defects, but also lays a solid foundation for subsequent model training and verification work.
3. The lightweight fabric defect detection algorithm based on deformable convolution and high-level screening feature pyramid according to claim 1 is characterized in that: The content of step 2 will mainly include the addition of deformable convolution, the construction of GS-SPPF, and the high-level screening feature pyramid of the neck network. Specifically: S21. Use deformable convolution to optimize the convolution layer. Improve the model feature extraction capability. ; S22. Construct GS-SPPF module. Reduce the number of model parameters and calculation while maintaining the accuracy improvement; S23. The idea of high-level feature pyramid screening is used to optimize FPN. This improves the model's feature expression ability on fabric defect datasets and significantly reduces the number of parameters and calculations.
4. The lightweight fabric defect detection algorithm based on deformable convolution and high-level screening feature pyramid according to claim 1 is characterized in that: In step S21, since the shape of the fabric defects is irregular, ordinary convolution is difficult to cover the entire defect for feature extraction, resulting in poor model feature extraction capability. Deformable convolution can better fit the fabric defects, thereby better extracting fabric defect features.
5. The lightweight fabric defect detection algorithm based on deformable convolution and high-level screening feature pyramid according to claim 3 is characterized in that: In step S22, we specially introduced the GSConv lightweight convolution technology in the SPPF structure. Specifically, we replaced the 1×1 standard convolution feature generation layer in SPPF. The original 1×1 standard convolution was replaced by GSConv with 3×3 standard convolution and 5×5 depth convolution. This change not only enables the SPPF module to directly obtain a larger local receptive field, thereby improving the detection accuracy, but also effectively reduces the amount of calculation at this stage through the sparsity characteristics of GSConv.
6. The lightweight fabric defect detection algorithm based on deformable convolution and high-level screening feature pyramid according to claim 3 is characterized in that: In step S23, unlike the FPN used in the YOLO series, the high-level screening feature pyramid takes into account the multi-scale features of fabric defects, embeds multiple CA modules and feature transposition modules in the neck part for deep feature screening and fusion, and discards unnecessary modules in the original FPN to achieve lightweight detection model, which significantly improves the feature expression ability of the model on the fabric defect dataset.
Citation Information
Cited By
Meter identification method based on deep learning and dynamic distortion correction
CN120544211A
Surface-mounted element welding defect detection method based on transfer learning and SimAM-SASPPF-YOLOv8
CN121073957A