Fabric defect detection method based on FABCS-YOLO algorithm
By improving the YOLOv8 algorithm, the FADC_C2f module, AIFI layer and CSPStage module are introduced to optimize gradient flow and bounding box regression, solving the real-time and accuracy problems of fabric defect detection in the prior art, and achieving efficient and accurate defect detection.
Patent Information
- Application Number
- CN202510403144.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-08
AI Technical Summary
The existing fabric defect detection technology has shortcomings in real-time, ability to detect small-size defects and adaptability to complex backgrounds, resulting in frequent missed and mis-checking, making it difficult to meet the inspection needs of high-speed production lines.
The fabric defect detection method based on the FABCS-YOLO algorithm is adopted, and the FADC_C2f module is constructed by introducing the FADC layer, combining the AIFI layer and the CSPStage module, multi-scale feature extraction and global context perception are enhanced, network structure and gradient flow are optimized, and bounding box regression accuracy is improved using the PIoU2 loss function.
It improves the accuracy and real-time performance of fabric defect detection, reduces mis-checking, meets the inspection needs of high-speed production lines, and improves production efficiency and product quality.
Smart Images

Figure CN120278979A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of defect detection, and particularly relates to a fabric defect detection method based on the FABCS-YOLO algorithm. Background Art
[0002] Surface defects (such as more than 80 kinds of defects like broken warp, holes, stains, etc.) generated during the cloth production process will directly lead to a reduction in product value and even affect the brand reputation of the enterprise. Therefore, developing an intelligent detection system with multi-defect synchronous recognition and adaptive production line speed has become the core path to break through the quality control bottleneck of the textile industry and achieve cost reduction and efficiency improvement.
[0003] In recent years, with the development of deep learning technology, object detection algorithms based on convolutional neural networks (CNNs) have gradually been applied to cloth surface defect detection. These algorithms can automatically extract image features and perform defect localization and recognition, greatly improving the detection accuracy and efficiency. Among them, the YOLO algorithm has been widely used and concerned due to its real-time performance, high efficiency, and continuously optimized generalization performance. The YOLOv8 algorithm is a model that balances speed and accuracy in single-stage object detection algorithms. It mainly extracts features directly from images through a single-stage detector, with high detection speed and good real-time performance.
[0004] In the prior art, during the manual defect detection link of traditional textile production, false detection and missed detection of cloth defects are likely to occur due to problems such as human eye visual fatigue and inattentiveness. Traditional mathematical detection methods are also used for cloth detection. However, when traditional mathematical detection methods are used for detecting small target defects, when there are situations such as too dark light for image sampling at the industrial site and low contrast between cloth defects and the picture background, the detection effect is relatively poor. Although the existing deep learning-based cloth defect detection technologies have a certain ability to detect defects, they also have some significant drawbacks. Traditional object detection algorithms such as the YOLOv8 algorithm are relatively fast, but when YOLOv8n detects low-contrast small defects or composite defects composed of multiple small targets, problems such as missed detection or false detection caused by incomplete local recognition are likely to occur. Secondly, although Faster R-CNN can achieve high detection accuracy, its complex network structure leads to a large amount of computation and a slow processing speed, making it difficult to meet the requirements of real-time detection. The existing deep learning cloth defect detection algorithms often need to balance between speed and accuracy, resulting in frequent problems such as low accuracy, missed detection, and false detection in small target detection.
[0005] The objective of the present invention is to improve the efficiency and accuracy of fabric defect detection through innovative technical means. Existing detection methods have problems such as insufficient real-time performance, limited ability to detect small-sized defects, and poor adaptability to complex backgrounds. Therefore, the present invention aims to overcome these challenges and achieve rapid and accurate detection of various defects on fabric cloth in a high-speed production line, including but not limited to holes, stains, knots, loose warp, and broken warp. By improving the detection accuracy and real-time performance, it can effectively improve production efficiency and product quality, providing a more reliable quality control solution for industrial manufacturing. Summary of the Invention
[0006] Aiming at the problems of missed detection, false detection, and real-time performance in the surface defect detection of textile cloth, a fabric defect detection method based on the FABCS-YOLO algorithm is proposed. By introducing the FADC layer to construct the FADC_C2f module, and at the same time introducing the AIFI layer and the CSPStage module, it enhances multi-scale feature extraction and global context awareness, effectively improving the detection accuracy of low-contrast small defects and composite defects, reducing missed detection and false detection, while maintaining real-time performance and high efficiency.
[0007] The technical means adopted by the present invention are as follows:
[0008] A fabric defect detection method based on the FABCS-YOLO algorithm, comprising:
[0009] Obtaining fabric image data to be detected;
[0010] Inputting the fabric image data to be detected into a trained fabric defect detection network, the fabric defect detection network is obtained by improving the YOLOv8 model, and the fabric defect detection network includes a Backbone part, a Neck part, and a Head part;
[0011] The Backbone part is composed of multiple convolutional layers, multiple C2f modules, multiple FADC_C2f modules, an AIFI module, and a BAMBlock module. The AIFI module is used to achieve global context awareness feature fusion through the multi-head self-attention mechanism and position encoding, and the BAMBlock module is used to enhance the feature expression of the backbone network;
[0012] The Neck part is composed of a convolutional layer, a splicing layer, an Upsample module, and a CSCPStage module. The CSCPStage module is used to optimize the gradient flow and enhance multi-level feature capture;
[0013] The Detect module in the Head part adopts the PIoU2 loss function, and the PIoU2 loss function is used for bounding box regression to improve the positioning accuracy;
[0014] Obtain the output of the fabric defect detection network as the fabric defect detection result.
[0015] Furthermore, the FADC_C2f module is implemented in the following manner:
[0016] Replace the Bottleneck unit in the traditional C2f module with the Bottleneck_FADC unit. The expression of the Bottleneck_FADC layer is defined as follows:
[0017] X hidden = σ(W1 * X)
[0018] Y = f(FADC(X hidden ))
[0019] where X hidden is the output of the hidden layer, σ() is the activation function, W1 is the weight of the first convolutional layer, X is the input feature map, Y is the output feature map, and f() is the non-linear mapping function.
[0020] Furthermore, the working process of the AIFI module is as follows:
[0021] First, input the feature map into the AIFI module, and rearrange and flatten the input feature map;
[0022] Second, calculate the two-dimensional sine-cosine position encoding of the flattened feature map;
[0023] Then, add the calculated position encoding to the flattened feature map element-wise to generate a feature map with fused position encoding;
[0024] Next, input the feature map with fused position encoding into the Transformer encoder for processing, calculate the relationship between features and generate the processed feature map. The Transformer encoder includes multi-head self-attention, feed-forward neural network, residual connection, and layer normalization;
[0025] Finally, restore the processed feature map to its original shape and output the restored feature map.
[0026] Furthermore, the BAMBlock module includes a channel attention unit and a spatial attention unit connected in parallel. The working process of the BAMBlock module is as follows:
[0027] First, input the input feature map into the channel attention unit to obtain the channel attention weight;
[0028] Second, input the input feature map into the spatial attention unit to obtain the spatial attention weight;
[0029] Then, the channel attention weight and the spatial attention weight are weighted and merged, and after passing through the Sigmoid activation function, the final attention weight is generated;
[0030] Finally, the input feature map and the final attention weight are weighted and fused to obtain the enhanced feature map.
[0031] Furthermore, the working process of the channel attention unit is as follows:
[0032] Through the global pooling operation, the spatial information of each channel of the input feature map is compressed into a 1×1 dimension, and the feature map after compressing the dimension is passed through two fully connected layers for dimension reduction and recovery to generate the channel attention weight;
[0033] The working process of the spatial attention unit is as follows:
[0034] The input feature map is passed through the dilated convolution to increase the receptive field, and the correlation between the feature maps after increasing the receptive field is calculated by convolution to generate the spatial attention weight.
[0035] Furthermore, the working process of the CSCPStage module is as follows:
[0036] First, the input feature map is passed through two convolutional layers to split the input feature map into two parts, obtaining the first part feature map and the second part feature map;
[0037] Secondly, the second part feature map is passed through several convolutional blocks to extract higher-level feature information, and the intermediate feature map is obtained after processing;
[0038] Then, the intermediate feature map and the first part feature map are concatenated together to form the concatenated feature map;
[0039] Finally, the concatenated feature map is passed through a convolutional layer to obtain the output feature map.
[0040] Furthermore, the PIoU2 loss function is specifically defined as:
[0041]
[0042] Among them, IoU is the overlap rate between the predicted bounding box and the ground truth bounding box, B pred is the coordinate of the predicted bounding box, B gt is the coordinate of the ground truth bounding box, is the PIoU2 loss function, x is the modulation factor, x = Λe -p where Λ is the relevant hyperparameter of the PIoU2 loss function, and P is the normalized bounding box gap;
[0043] The calculation formula of the normalized bounding box gap is:
[0044]
[0045] Among them, P is the normalized boundary gap, is the horizontal difference on the left side, is the horizontal difference on the right side, is the vertical difference on the upper side, is the vertical difference on the lower side, w gt is the width of the ground truth box, h gt is the height of the ground truth box.
[0046] Furthermore, the network architecture of the Backbone part is as follows:
[0047] The Backbone part includes a first convolutional layer, a second convolutional layer, a first C2f module, a third convolutional layer, a second C2f module, a fourth convolutional layer, a first FADC_C2f module, a fifth convolutional layer, a second FADC_C2f module, an AIFI module, and a BAMBlock module connected in sequence; among them, the second C2f module, the first FADC_C2f module, and the BAMBlock module respectively output a first effective feature map, a second effective feature map, and a third effective feature map.
[0048] Furthermore, the network architecture of the Neck part is as follows:
[0049] The Neck part includes a sixth convolutional layer, a first splicing layer, a first CSCPStage module, a first Upsample module, a second splicing layer, a seventh convolutional layer, an eighth convolutional layer, a fourth CSCPStage module, a third splicing layer, a tenth convolutional layer, a second CSCPStage module, a fourth splicing layer, a third CSCPStage module, a ninth convolutional layer, a fifth splicing layer, a fifth CSCPStage module, and an eleventh convolutional layer;
[0050] The sixth convolutional layer, the first splicing layer, the first CSCPStage module, the first Upsample module, the second splicing layer, and the eighth convolutional layer are connected in sequence; the fourth CSCPStage module, the third splicing layer, the tenth convolutional layer, the second CSCPStage module, the fourth splicing layer, the third CSCPStage module, the ninth convolutional layer, the fifth splicing layer, the fifth CSCPStage module, and the eleventh convolutional layer are connected in sequence; the second splicing layer is connected to the second CSCPStage module, the first CSCPStage module is connected to the third splicing layer, the second CSCPStage module is connected to the fifth splicing layer, and the eleventh splicing layer is connected to the third splicing layer;
[0051] Among them, the inputs of the eighth convolutional layer and the fourth splicing layer are the features output by the second C2f module, the inputs of the second splicing layer and the seventh convolutional layer are the features output by the first FADC_C2f module, and the input of the sixth convolutional layer is the features output by the BAMBlock module.
[0052] Further, the Head part includes three Detect modules, and the Detect modules respectively receive the feature maps output by the third CSCPStage module, the fifth CSCPStage module, and the fourth CSCPStage module to generate the final feature map.
[0053] Compared with the prior art, the present invention has the following advantages:
[0054] 1. By integrating the CSP and SPP modules, the present invention optimizes the network structure and gradient flow, effectively suppressing the gradient vanishing problem, significantly enhancing the multi-level feature capture ability, and simultaneously enhancing the training efficiency and detection accuracy.
[0055] 2. By optimizing the forward propagation calculation efficiency and the multi-scale feature fusion mechanism, the present invention performs excellently in the defect detection task, with higher real-time performance and training efficiency. It can quickly and accurately complete target recognition in a complex industrial environment, ensuring the stability of detection. Compared with other technologies, FABCS-YOLO shows obvious advantages in both detection speed and accuracy, and can meet the requirements of high-standard production detection.
[0056] Based on the above reasons, the present invention can be widely promoted in the fields of defect detection and the like. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0058] Figure 1 It is a schematic diagram of the overall network structure of FABCS-YOLO proposed by the present invention.
[0059] Figure 2 It is the overall flow chart of the present invention.
[0060] Figure 3 It is a schematic diagram of the structure of the FADC_C2f network module proposed by the present invention.
[0061] Figure 4 It is a schematic diagram of the structure of the context-aware feature fusion AIFI proposed by the present invention.
[0062] Figure 5 Structural schematic diagram of optimizing CSPStage by integrating CSP and SPP modules proposed by the present invention
[0063] Figure 6 Schematic diagram of visual comparison of detection effects of using FABCS - YOLO and YOLOv8 in the embodiments of the present invention Specific implementation manners
[0064] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0065] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above - mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order different from those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non - exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0066] As Figure 1 and Figure 2 shown, the present invention provides a fabric defect detection method based on the FABCS - YOLO algorithm. The method is specifically as follows:
[0067] Obtain the fabric image data to be detected;
[0068] Input the fabric image data to be detected into the trained fabric defect detection network, which is improved based on the YOLOv8 model. The fabric defect detection network includes a Backbone part, a Neck part, and a Head part;
[0069] The Backbone part consists of multiple convolutional layers, multiple C2f modules, multiple FADC_C2f modules, an AIFI module, and a BAMBlock module. The AIFI module is used to achieve global context-aware feature fusion through the multi-head self-attention mechanism and positional encoding, and the BAMBlock module is used to enhance the feature expression of the backbone network;
[0070] The Neck part consists of convolutional layers, a concatenation layer, an Upsample module, and a CSCPStage module. The CSCPStage module is used to optimize the gradient flow and enhance multi-level feature capture;
[0071] The Detect module in the Head part uses the PIoU2 loss function, which is used for bounding box regression to improve the localization accuracy;
[0072] Obtain the output of the fabric defect detection network as the fabric defect detection result.
[0073] The improvements of this model over the traditional YOLOv8 are as follows:
[0074] (1) Use an adaptive dilated convolution new feature extraction to improve the original C2f module and enhance the network's ability to extract features of different scales
[0075] In object detection and other vision tasks, the ability to extract multi-scale features is crucial for detection accuracy. However, the traditional C2f module uses a convolutional kernel of a fixed size for feature extraction, making it difficult to dynamically adapt to targets of different scales, especially with limited detection ability for small targets. To solve this problem, a redesigned feature extraction module - FADC_C2f is proposed.
[0076] As Figure 3 shown, on the basis of the C2f structure, the FADC_C2f module replaces the standard Bottleneck with Bottleneck_FADC. In Bottleneck_FADC, the first 1×1 convolutional layer remains unchanged, and the second 3×3 convolutional layer is replaced by FADC, enabling it to adaptively adjust the dilation rate and enhance the ability to capture multi-scale information. That is, the FADC_C2f module maintains the core framework of the C2f module, that is, a part of the features are passed through a direct connection method, and the other part is fused with the direct connection features after a series of non-linear transformations. The difference is that the standard convolution used in the traditional Bottleneck unit is replaced by Bottleneck_FADC, which uses FADC to implement dynamic dilated convolution.
[0077] Specifically, the design of the Bottleneck_FADC layer is as follows:
[0078] Xhidden = σ(W1 * X), Y = f(FADC(X hidden ))
[0079] where X is the input feature map, Y is the output feature map, f() is a non-linear mapping function, and X hidden is the output of the hidden layer, W1 is the weight of the first convolutional layer, and σ() is the activation function.
[0080] The FADC part adaptively calculates the local information of the input features and dynamically determines the dilation rate d = Φ(X) in the convolutional operation, thereby enhancing the ability to model multi-scale features. At the same time, in this module, by using nn.ModuleList to organize multiple Bottleneck_FADC units, it not only maintains the flexible scalability of the module to adapt to network architectures of different depths, but also ensures the effective transmission and fusion of information flow.
[0081] In summary, FADC_C2f is a new feature extraction module improved based on the C2f structure to enhance the multi-scale feature extraction ability of the network. Its core idea is to dynamically adjust the dilation rate of the convolutional kernel according to the input features, thereby realizing the adaptive sampling of local features. Compared with the standard C2f, FADC_C2f dynamically adjusts the receptive field through adaptive dilated convolution, improves the ability to model small targets and complex scenes, and at the same time maintains the computational efficiency of the C2f structure.
[0082] Differences between the FADC_C2f module and the traditional Bottleneck:
[0083] ① Dynamic adjustment of the receptive field: The traditional Bottleneck uses a fixed 3×3 convolutional kernel, and its receptive field remains unchanged throughout the network, restricting the multi-scale feature extraction ability. The Bottleneck FADC Branch realizes the dynamic adjustment of the receptive field through FADC, and can automatically optimize the dilation rate according to the input features to adapt to the feature extraction requirements of different scales.
[0084] ② Multi-scale feature extraction ability: Traditional convolution has limited ability to process small objects and complex backgrounds, while FADC can more effectively extract feature information of different scales by flexibly adjusting the receptive field. It has a significant improvement especially in small object detection and detail capture.
[0085] ③ Frequency selection mechanism: FADC adopts a frequency selection mechanism to select the most suitable convolutional kernel according to the frequency components of the input features, thereby further enhancing the adaptability of the model to different features.
[0086] (2) Replace the traditional feature pyramid SPPF module with the Transformer-based AIFI module
[0087] The design of the AIFI layer replaces the traditional pooling operation through the multi - head self - attention mechanism and positional encoding (based on two - dimensional sine - cosine encoding), thus providing a more global and context - aware feature fusion method.
[0088] As Figure 4 shown, during the forward propagation of the AIFI layer, first, the input feature map is fed into the AIFI module. The input feature map is flattened and transformed into a shape suitable for the Transformer encoder format, changing from [B, C, H, W] to [B, H×W, C].
[0089] Secondly, the two - dimensional sine - cosine positional encoding of the flattened feature map is calculated to ensure position - aware self - attention calculation.
[0090] Then, the calculated positional encoding is added element - by - element to the flattened feature map to generate a feature map with fused positional encoding, ensuring position - aware self - attention calculation.
[0091] Then, the feature map with fused positional encoding is input into the Transformer encoder for processing. The relationships between features are calculated and a processed feature map is generated. The Transformer encoder includes multi - head self - attention, a feed - forward neural network, residual connections, and layer normalization. Through the self - attention mechanism in the Transformer encoder, the relationships between features are calculated and a new feature representation is generated.
[0092] Finally, the processed feature map is restored to its original shape, and the restored feature map is output. The output features will be restored to the same size as the input image, i.e., [B, C, H, W], for subsequent processing or output. The whole process ensures multi - scale feature fusion and effective capture of global context information, thus enhancing the model's perception ability of images.
[0093] Compared with the traditional SPPF layer, the AIFI layer can capture the global context information in images more accurately, especially being outstanding in multi - scale object detection in complex scenes. Since the AIFI layer processes image features through the self - attention mechanism, compared with the pooling operation, it can avoid information loss caused by inconsistent scales and effectively improve the accuracy and robustness of the model in object detection tasks.
[0094] (3) Introduce the BAM (Bottleneck Attention Module) bottleneck attention mechanism into the backbone network to enhance the feature expression of the backbone network
[0095] Specifically, the BAMBlock module includes a channel attention unit and a spatial attention unit connected in parallel. The working process of the BAMBlock module is as follows:
[0096] First, input the input feature map into the channel attention unit to obtain the channel attention weights.
[0097] Secondly, input the input feature map into the spatial attention unit to obtain the spatial attention weights.
[0098] Then, weighted combine the channel attention weights and the spatial attention weights, and through the Sigmoid activation function, generate the final attention weights.
[0099] Finally, perform weighted fusion on the input feature map and the final attention weights to obtain the enhanced feature map.
[0100] The working process of the channel attention unit is as follows:
[0101] Compress the spatial information of each channel of the input feature map into a 1×1 dimension through global pooling operation, and pass the feature map after compressing the dimension through two fully connected layers for dimension reduction and restoration to generate the channel attention weights.
[0102] Multiply the generated attention weights with each channel of the input feature map element by element, so as to weight different channels, enhance the feature expression of important channels, and suppress the influence of unimportant channels.
[0103] The working process of the spatial attention unit is as follows:
[0104] Increase the receptive field of the input feature map through dilated convolution, and utilize the convolution calculation to increase the correlation between the feature maps after increasing the receptive field to generate the spatial attention weights.
[0105] Multiply the generated spatial attention weights with the input feature map element by element, so as to weight according to the importance of each position, enhance the features of important positions, and suppress the features of unimportant positions.
[0106] The BAM (Bottleneck Attention Module) bottleneck attention mechanism enhances the expressive power of feature maps by jointly modeling channel attention (CA) and spatial attention (SA). This mechanism calculates the importance weights of the channel dimension and the spatial dimension for the input feature map respectively to improve the network's attention to key information. In the BAM module, the channel attention (CA) compresses the spatial information of each channel into a 1×1 dimension through global pooling operation, and reduces and restores the dimension through two fully connected layers, finally generating the channel attention weights. Among them, the spatial attention (SA) module increases the receptive field through dilated convolution to capture richer spatial information, and calculates the correlation between features using convolution to generate the spatial attention weights. Finally, BAM generates the final weighted feature by weighted merging of the channel and spatial attention weights. This process calculates the final attention weights through the Sigmoid activation function and adjusts the original feature map, so as to strengthen the important features while maintaining the original information features. In this way, BAM can not only strengthen important features in the channel dimension, but also adjust and optimize features in the spatial dimension, thus enhancing the robustness of feature expression.
[0107] (4) Optimize the gradient flow and multi-level feature capture of the CSPStage by fusing the CSP and SPP modules
[0108] Specifically, as Figure 5 shown, the working process of the CSCPStage module is as follows:
[0109] First, the input feature map is divided into two parts through two 1×1 convolutional layers to obtain the first part of the feature map and the second part of the feature map.
[0110] Secondly, the second part of the feature map is passed through several convolutional blocks to extract higher-level feature information. After processing, an intermediate feature map is obtained. The several convolutional blocks are multiple BasicBlock_3x3_Reverse convolutional blocks. Specifically, this step processes the second part of the feature map to extract higher-level feature information. These convolutional blocks improve the gradient flow through residual connections and convolutional operations, avoid gradient disappearance, and at the same time enhance the feature expression ability, improving the expressiveness and training stability of the network.
[0111] Then, the intermediate feature map and the first part of the feature map are concatenated together to form the concatenated feature map.
[0112] Finally, the concatenated feature map is passed through a 1×1 convolutional layer to obtain the output feature map.
[0113] The design of the CSPStage module aims to address the vanishing gradient problem in deep networks and enhance the ability to capture multi-level features by combining Cross-Stage Partial Connection (CSP) and Spatial Pyramid Pooling (SPP) modules. In traditional convolutional neural network architectures, the capture of deep features is often bottlenecked by the flow of global information, which limits the training effect and stability of the model. In contrast, the improved CSPStage architecture avoids this bottleneck problem through the segmentation and partial parallel processing of the input feature map, thereby enhancing the gradient flow and significantly improving the stability and convergence speed during training.
[0114] Specifically, the CSPStage module first divides the input feature map into two parts through two 1×1 convolutional layers: one part keeps the original features, and the other part is processed layer by layer through multiple convolutional blocks (such as BasicBlock_3x3_Reverse). This process separates different information flows of the input feature map through convolution operations, so that further feature abstraction can be carried out through different processing methods in the next stage. The specific mathematical formula is: x1 = Conv 1×1 (x), x2 = Conv 1×1 (x). This way of segmentation and processing enables the network to maintain the diversity of input features and enhances the representation ability of the network by processing the features of different parts layer by layer.
[0115] In the feature fusion stage, the CSPStage concatenates the intermediate feature map {y i} that has undergone convolution processing with the unmodified x1 to form a new feature map y concat , and then fuses the concatenated features through a 1×1 convolutional layer to finally generate the output feature map This process ensures the continuity of information flow and can effectively integrate multi-level feature representations. Its mathematical expression is:
[0116] y concat = Concat({y i},x1), y out = Conv 1×1 (y concat )
[0117] In addition, to further enhance the ability to capture multi-level features, the CSPStage architecture embeds a Spatial Pyramid Pooling (SPP) module in some stages. The SPP module extracts feature information at each level through pooling operations at different levels, and integrates this information through convolutional layers, thus effectively enhancing the network's multi-level perception ability. Especially when dealing with object detection tasks in complex scenarios, it significantly improves the accuracy and robustness of the model.
[0118] The new position selection of the CSPStage module has a clear design basis. First of all, the insertion of the CSPStage module aims to optimize the feature representation ability at each scale, enabling the network to more effectively capture objects of various scales. In addition, as the network depth increases, the size of the feature map gradually shrinks, but the degree of information abstraction gradually increases. CSPStage further enhances the feature expression ability at each scale by splitting and parallel processing the feature map. Secondly, while enhancing the feature extraction ability, the CSPStage module specifically solves the problem of gradient flow in deep networks. In deep networks, gradients are often prone to vanishing or exploding, affecting the training and stability of the network. Through the split and parallel processing design of CSPStage and combined with shortcut connections, the gradient can flow effectively, ensuring the stability of the network during object detection. Especially when dealing with multi-scale feature maps, it improves the training efficiency and convergence speed of the network. The CSPStage module also adopts a method of splitting and multi-level processing of the information flow.
[0119] Through the above design, this architecture uses the CSPStage module to optimize the gradient flow, enhance the ability to capture multi-level features, and improve the training stability and object detection performance of the network. Experimental results show that this architecture is superior to traditional convolutional neural network architectures in object detection tasks, especially in deep networks, with higher training efficiency and accuracy.
[0120] (5) Replace the original CIoU loss function with the PIoU2 loss function to optimize the gradient signal and improve the localization accuracy.
[0121] In object detection, the accuracy of bounding box regression is crucial for the overall performance. The traditional CIoU loss function mainly focuses on the overlap rate and center distance between the predicted box and the ground truth box, but the gradient is prone to saturation in the case of high overlap, thus affecting the optimization of fine-grained localization. For this reason, the PIoU2 loss function introduces a normalized quantity of the boundary position deviation and an exponential modulation mechanism, which is specifically defined as:
[0122]
[0123] where IoU is the overlap rate between the predicted box and the ground truth box, Bpred Let \(B\) be the coordinates of the predicted bounding box, gt and \(G\) be the coordinates of the ground truth bounding box, let \(L_{PIoU2}\) be the PIoU2 loss function, \(x\) be the modulation factor, \(x = \Lambda e\), -p where \(\Lambda\) is a hyperparameter related to the \(L_{PIoU2}\) loss function, and \(P\) is the normalized bounding gap.
[0124] The hyperparameter \(\Lambda\) allows \(L_{PIoU2}\) to flexibly adjust the sensitivity to bounding box deviations according to the scales and shapes of different objects. This adaptive regulation mechanism helps to achieve better performance in object detection tasks with multi-scales and multi-shapes. Secondly, by introducing the exponential modulation term it ensures that sufficient gradient information can be provided even in the case of high overlap, thus more effectively guiding the model for fine-tuning.
[0125] \(P\) is the normalized bounding gap, which is used to quantify the deviation of the predicted box from the ground truth box on each side, and its definition is:
[0126]
[0127] where \(P\) is the normalized bounding gap, \(l_x\) is the horizontal difference on the left side, \(r_x\) is the horizontal difference on the right side, \(t_y\) is the vertical difference on the upper side, \(b_y\) is the vertical difference on the lower side, \(w_G\) gt is the width of the ground truth box, \(h_G\) gt is the height of the ground truth box.
[0128] The normalized bounding gap \(P\) accurately quantifies the deviation of the predicted box from the ground truth box on each side, rather than relying solely on the overall overlap rate. This enables the loss function to capture more subtle localization errors, thus improving the regression accuracy.
[0129] Overall, by introducing the normalized bounding deviation and the exponential modulation mechanism into the loss calculation, the \(L_{PIoU2}\) loss function not only compensates for the problem of insufficient gradients in the traditional IoU loss in the case of high overlap, but also provides a precise description of the subtle deviations of each side of the predicted box, thereby achieving higher accuracy and training stability in the bounding box regression task.
[0130] FABCS-YOLO evaluation metrics:
[0131] In the field of object detection, accuracy, recall, and mean average precision (mAP) are often used as evaluation metrics for the detection effect of the model. Accuracy represents the proportion of true samples among the samples predicted as positive, that is, the proportion of true positive examples in the prediction results. Recall represents the proportion of all positive examples that are predicted correctly. The calculation formulas for accuracy Precision and recall Recall are as follows:
[0132]
[0133] Among them: TP is true positives, representing the number of correct samples predicted correctly by the model. FP is false positives, representing the number of correct samples predicted incorrectly by the model. FN represents the number of incorrect samples predicted incorrectly by the model. However, accuracy and recall alone are not comprehensive enough as indicators for evaluating the performance of the model. Therefore, the indicators AP and mAP are introduced. AP is the curve integral of accuracy and recall, and mAP is the average of the detection accuracies of various defect categories. The definition formulas of both are as shown in the formula:
[0134]
[0135] Among them, C is the number of categories, and P(R) is the precision, which can be used as a function of recall, representing the precision of the model at different recall rates.
[0136] In addition, in order to better compare FABCS-YOLO with other models, the frames per second (FPS) is introduced as an evaluation indicator, which reflects the number of pictures that the algorithm can process per second.
[0137] In the present invention, the network architecture of the Backbone part is as follows:
[0138] The Backbone part includes a first convolutional layer, a second convolutional layer, a first C2f module, a third convolutional layer, a second C2f module, a fourth convolutional layer, a first FADC_C2f module, a fifth convolutional layer, a second FADC_C2f module, an AIFI module, and a BAMBlock module connected in sequence; among them, the second C2f module, the first FADC_C2f module, and the BAMBlock module respectively output a first effective feature map, a second effective feature map, and a third effective feature map.
[0139] In the present invention, the network architecture of the Neck part is as follows:
[0140] The Neck part includes a sixth convolutional layer, a first splicing layer, a first CSCPStage module, a first Upsample module, a second splicing layer, a seventh convolutional layer, an eighth convolutional layer, a fourth CSCPStage module, a third splicing layer, a tenth convolutional layer, a second CSCPStage module, a fourth splicing layer, a third CSCPStage module, a ninth convolutional layer, a fifth splicing layer, a fifth CSCPStage module, and an eleventh convolutional layer.
[0141] The sixth convolutional layer, the first splicing layer, the first CSCPStage module, the first Upsample module, the second splicing layer, and the eighth convolutional layer are connected in sequence; the fourth CSCPStage module, the third splicing layer, the tenth convolutional layer, the second CSCPStage module, the fourth splicing layer, the third CSCPStage module, the ninth convolutional layer, the fifth splicing layer, the fifth CSCPStage module, and the eleventh convolutional layer are connected in sequence; the second splicing layer is connected to the second CSCPStage module, the first CSCPStage module is connected to the third splicing layer, the second CSCPStage module is connected to the fifth splicing layer, and the eleventh splicing layer is connected to the third splicing layer.
[0142] Among them, the inputs of the eighth convolutional layer and the fourth splicing layer are the features output by the second C2f module, the inputs of the second splicing layer and the seventh convolutional layer are the features output by the first FADC_C2f module, and the input of the sixth convolutional layer is the features output by the BAMBlock module.
[0143] In the present invention, the Head part includes three Detect modules, and the Detect modules respectively receive the feature maps output by the third CSCPStage module, the fifth CSCPStage module, and the fourth CSCPStage module to generate the final feature maps.
[0144] Embodiment
[0145] 1. Data preparation
[0146] The Alibaba Tianchi cloth defect public dataset is obtained. There are 5,913 images in total, including 34 types of defects with different sizes. One picture may contain one or more types of defects. The types of cloth defects include the following: no defect, hole, water stain, oil stain, stain, foreign fiber, knot, flower springboard, centipede, hair particle, thick warp, loose warp, broken warp, hanging warp, thick fiber, weft shrinkage, sizing stain, warping knot, star jump, skip flower, broken spandex, density difference, wavy density, color difference, abrasion mark, rolling mark, dead wrinkle, and weft yarn defect. The backgrounds of the images in this dataset are complex, and the defect targets are small and difficult to identify. There are many types of cloth defects. According to the defect morphology and formation reasons, the detection tasks are carefully divided into 20 types of defect categories, and the dataset is divided into a training set and a validation set according to 8:2 for input training.
[0147] 2. Model optimization
[0148] Based on YOLOv8, the C2f module in the original backbone is improved by the designed feature extraction module FADC_C2f, enhancing the network's ability to extract features of different scales of fabric defects; the traditional feature pyramid SPPF module is replaced by the AIFI module based on Transformer, effectively improving the accuracy and robustness of the model in object detection tasks; the BAM bottleneck attention mechanism is introduced to enhance the feature expression of the backbone network; the gradient flow and multi-level feature capture of the CSPStage are optimized by fusing the CSP and SPP modules, enhancing the training stability and object detection performance of the network, and achieving higher training efficiency and accuracy; that is, it is named FABCS (FADC_C2f - AIFI - BAM - CSPStage), and the FABCS - YOLO model is obtained.
[0149] 3. Model Training
[0150] The FABCS - YOLO model is trained using the Alibaba Tianchi fabric defect public dataset, the PIoU2 loss function, and the SGD optimizer. The input image size for model training is set to 640 * 640, the number of training iterations is 300, the batch size is 32, the initial learning rate is 0.01, and the weight decay coefficient is 0.0005. After training, the model will save the best weight file best.pt with the best training results.
[0151] 4. Model Validation
[0152] The trained FABCS - YOLO model is validated using the validation set. The evaluation metrics include accuracy, recall, mean average precision, etc., to comprehensively evaluate the performance of the model.
[0153] 5. Effect Comparison
[0154] The FABCS - YOLO model is compared and tested with other models on the same validation set, and the performance of various models on each index is analyzed to verify the advantages of the FABCS - YOLO model in fabric defect detection tasks. Table 1 shows the comparison test results of different models using the same Alibaba Tianchi fabric defect public dataset among current mainstream models. Figure 6 This is a schematic diagram of the experimental result comparison between this model and the YOLOv8 model.
[0155] Table 1 Comparison Test Results of the FABCS - YOLO Model and Mainstream Models
[0156]
[0157]
[0158] 6. Deployment and Application
[0159] Deploy the trained FABCS-YOLO model to the actual fabric production line; obtain the real-time image of the fabric surface through a high-speed camera and input it into the FABCS-YOLO model for processing; the model will automatically locate and detect the defects in the image, and according to the detection results, take corresponding measures to deal with the defects to improve production efficiency and product quality.
[0160] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A fabric defect detection method based on the FABCS-YOLO algorithm, characterized in that Including: Obtain the image data of the fabric to be detected; Input the image data of the fabric to be detected into the trained fabric defect detection network, which is obtained by improving the YOLOv8 model. The fabric defect detection network includes a Backbone part, a Neck part, and a Head part; The Backbone part is composed of multiple convolutional layers, multiple C2f modules, multiple FADC_C2f modules, an AIFI module, and a BAMBlock module. The AIFI module is used to achieve global context-aware feature fusion through the multi-head self-attention mechanism and positional encoding, and the BAMBlock module is used to enhance the feature expression of the backbone network; The Neck part is composed of a convolutional layer, a splicing layer, an Upsample module, and a CSCPStage module. The CSCPStage module is used to optimize the gradient flow and enhance multi-level feature capture; The Detect module in the Head part uses the PIoU2 loss function, and the PIoU2 loss function is used for bounding box regression to improve the positioning accuracy; Obtain the output of the fabric defect detection network as the fabric defect detection result.
2. The fabric defect detection method based on the FABCS-YOLO algorithm according to claim 1, wherein The FADC_C2f module is implemented in the following way: Replace the Bottleneck unit in the traditional C2f module with the Bottleneck_FADC unit. The expression of the Bottleneck_FADC layer is defined as follows: X hidden = σ(W1 * X) Y = f(FADC(X hidden )) Among them, X hidden is the output of the hidden layer, σ() is the activation function, W1 is the weight of the first convolutional layer, X is the input feature map, Y is the output feature map, and f() is the non-linear mapping function.
3. The fabric defect detection method based on the FABCS-YOLO algorithm according to claim 1, wherein The working process of the AIFI module is as follows: First, input the feature map into the AIFI module, and rearrange and flatten the input feature map; Secondly, calculate the two-dimensional sine-cosine positional encoding of the flattened feature map; Then, add the calculated positional encoding to the flattened feature map element by element to generate a feature map fused with the positional encoding; Then, input the feature map fused with the positional encoding into the Transformer encoder for processing, calculate the relationship between features, and generate the processed feature map. The Transformer encoder includes multi-head self-attention, a feed-forward neural network, a residual connection, and layer normalization; Finally, restore the processed feature map to its original shape and output the restored feature map.
4. The fabric defect detection method based on the FABCS-YOLO algorithm according to claim 1, wherein The BAMBlock module includes a channel attention unit and a spatial attention unit connected in parallel. The working process of the BAMBlock module is as follows: First, input the input feature map into the channel attention unit to obtain the channel attention weight; Secondly, input the input feature map into the spatial attention unit to obtain the spatial attention weight; Then, weighted combine the channel attention weight and the spatial attention weight, and generate the final attention weight through the Sigmoid activation function; Finally, perform weighted fusion on the input feature map and the final attention weight to obtain the enhanced feature map.
5. The fabric defect detection method based on the FABCS-YOLO algorithm according to claim 4, wherein The working process of the channel attention unit is as follows: The spatial information of each channel of the input feature map is compressed into a 1×1 dimension through a global pooling operation, and the feature map after dimension compression is passed through two fully connected layers for dimension reduction and restoration to generate channel attention weights. The working process of the spatial attention unit is as follows: The input feature map is passed through dilated convolution to increase the receptive field, and the convolution is used to calculate the correlation between the feature maps after increasing the receptive field to generate spatial attention weights.
6. The fabric defect detection method based on the FABCS-YOLO algorithm according to claim 1, characterized in that The working process of the CSCPStage module is as follows: First, the input feature map is passed through two convolutional layers to divide the input feature map into two parts, obtaining the first part of the feature map and the second part of the feature map. Secondly, the second part of the feature map is passed through several convolutional blocks to extract higher-level feature information, and the intermediate feature map is obtained after processing. Then, the intermediate feature map and the first part of the feature map are concatenated together to form the concatenated feature map. Finally, the concatenated feature map is passed through a convolutional layer to obtain the output feature map.
7. The fabric defect detection method based on the FABCS-YOLO algorithm according to claim 1, characterized in that, The specific definition of the PIoU2 loss function is as follows: Among them, IoU is the overlap rate between the predicted bounding box and the ground truth bounding box, B pred is the coordinate of the predicted bounding box, B gt is the coordinate of the ground truth bounding box, is the PIoU2 loss function, x is the modulation factor, x = Λe -p , Λ is the relevant hyperparameter of the PIoU2 loss function, and P is the normalized bounding gap; The calculation formula of the normalized boundary gap is as follows: Among them, P is the normalized boundary gap, is the horizontal difference on the left side, is the horizontal difference on the right side, is the vertical difference on the upper side, is the vertical difference on the lower side, w gt is the width of the ground truth box, h gt is the height of the ground truth box.
8. The fabric defect detection method based on the FABCS-YOLO algorithm according to claim 1, wherein The network architecture of the Backbone part is as follows: The Backbone part includes a first convolutional layer, a second convolutional layer, a first C2f module, a third convolutional layer, a second C2f module, a fourth convolutional layer, a first FADC_C2f module, a fifth convolutional layer, a second FADC_C2f module, an AIFI module, and a BAMBlock module connected in sequence; among them, the second C2f module, the first FADC_C2f module, and the BAMBlock module output the first effective feature map, the second effective feature map, and the third effective feature map respectively.
9. The fabric defect detection method based on the FABCS-YOLO algorithm according to claim 1, characterized in that, The network architecture of the Neck part is as follows: The Neck part includes a sixth convolutional layer, a first concatenation layer, a first CSCPStage module, a first Upsample module, a second concatenation layer, a seventh convolutional layer, an eighth convolutional layer, a fourth CSCPStage module, a third concatenation layer, a tenth convolutional layer, a second CSCPStage module, a fourth concatenation layer, a third CSCPStage module, a ninth convolutional layer, a fifth concatenation layer, a fifth CSCPStage module, and an eleventh convolutional layer; The sixth convolutional layer, the first concatenation layer, the first CSCPStage module, the first Upsample module, the second concatenation layer, and the eighth convolutional layer are connected in sequence; the fourth CSCPStage module, the third concatenation layer, the tenth convolutional layer, the second CSCPStage module, the fourth concatenation layer, the third CSCPStage module, the ninth convolutional layer, the fifth concatenation layer, the fifth CSCPStage module, and the eleventh convolutional layer are connected in sequence; the second concatenation layer is connected to the second CSCPStage module, the first CSCPStage module is connected to the third concatenation layer, the second CSCPStage module is connected to the fifth concatenation layer, and the eleventh concatenation layer is connected to the third concatenation layer; Among them, the inputs of the eighth convolutional layer and the fourth concatenation layer are the features output by the second C2f module, the inputs of the second concatenation layer and the seventh convolutional layer are the features output by the first FADC_C2f module, and the input of the sixth convolutional layer is the features output by the BAMBlock module.
10. The fabric defect detection method based on the FABCS-YOLO algorithm according to claim 1, characterized in that, The Head part includes three Detect modules, and the Detect modules respectively receive the feature maps output by the third CSCPStage module, the fifth CSCPStage module, and the fourth CSCPStage module to generate the final feature maps.
Citation Information
Cited By
Vehicle defect detection method based on fusion frequency adaptive expansion convolution
CN121074005A