Fabric defect detection method based on improved YOLO model
By improving the YOLO model, using multi-scale feature extraction, fusion and enhancement modules, combined with Hal wavelet attention downsampling and lightweight encoder layer, the problems of information loss and low computing efficiency in existing fabric defect detection technology are solved, and more accurate and efficient defect detection is achieved.
Patent Information
- Application Number
- CN202510290650.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-17
AI Technical Summary
The existing fabric defect detection technology has the problems of information loss, low computing efficiency and difficulty in dealing with multi-scale targets. Especially in industrial scenarios, the size and shape of defects vary greatly, making it difficult to achieve efficient and accurate detection.
Using the improved YOLO model, the multi-scale feature map extraction module, multi-scale feature fusion module and multi-scale feature enhancement module are used, combined with Hal wavelet attention downsampling and lightweight encoder layer, the accuracy and computing efficiency of the model's defect detection are improved.
It effectively reduces the loss of feature information, improves the model's ability to capture details, enhances the ability to identify targets of different sizes, optimizes the overall performance of the model, and achieves more accurate defect detection.
Smart Images

Figure CN120163801A_ABST
Abstract
Description
[0001] The present invention relates to the field of textiles, and in particular, to a fabric defect detection method based on an improved YOLO model. Background Art
[0002] In modern textile industry, fabric defect detection has always been one of the key issues in development. The existence of defects will affect the aesthetics and quality of fabrics. Early detection and treatment of defects are crucial for textile production.
[0003] Traditional manual detection methods are inefficient, inaccurate, and suffer from subjectivity and high labor costs. Therefore, it is crucial to research an efficient and accurate intelligent fabric defect detection technology. In the field of object detection, especially in industrial scenarios, accurately identifying and locating defects is an important but challenging task. Traditional object detection methods, such as feature-based methods or early deep learning methods, often rely on manually designed image features. These methods are limited in performance when dealing with defect detection in complex backgrounds and varying lighting conditions. With the development of deep learning technology, object detection models based on convolutional neural networks (CNNs), such as the YOLO (You Only Look Once) series, have received extensive attention due to their fast and accurate detection capabilities. However, these models still face the problem of information loss when processing high-resolution images, especially during the downsampling process, resulting in the loss of important texture features and affecting detection accuracy. In addition, traditional object detection models also face challenges in dealing with multi-scale objects. Especially in industrial scenarios, the size and shape of defects may vary greatly, which requires the model to be able to capture features at different scales. Moreover, existing models also have limitations in computational efficiency, especially in application scenarios that require real-time detection, where computational resources and time costs become key factors.
[0004] For the above technical problems, no effective solutions have been proposed yet. Summary of the Invention
[0005] The present invention provides a fabric defect detection method based on an improved YOLO model, aiming to improve the detection accuracy of the model for defects while reducing the computational amount to meet the real-time or near-real-time industrial detection requirements. The present invention is proposed to solve the limitations in the prior art and improve the application effect and efficiency of object detection technology in the industrial field.
[0006] The present invention realizes this purpose through the following technical solutions:
[0007] A fabric defect detection method based on an improved YOLO model, comprising the following steps:
[0008] Step 1: Obtain fabric images on at least one fabric inspection machine;
[0009] Step 2: Input the fabric image into the defect detection model to obtain the defect detection result of the fabric image output by the defect detection model;
[0010] Step 3: The defect detection model includes a multi-scale feature map extraction module, a multi-scale feature fusion module, a multi-scale feature enhancement module, and a detection head; the multi-scale feature map extraction module is stacked by four sub-modules. The first sub-module consists of a block embedding layer and a lightweight encoder layer, and the latter three sub-modules consist of a Haar wavelet attention downsampling layer and a lightweight encoder layer; the multi-scale feature fusion module is used to fuse the multi-scale feature maps extracted by the multi-scale feature map extraction module, and it consists of three independent sub-modules, and each sub-module fuses the feature maps of adjacent two scales; the multi-scale feature enhancement module consists of two sub-modules in series and is used to further fuse the feature maps output by the multi-scale feature fusion module; the detection head sets multiple losses, including the loss of the predicted center coordinates, the loss of the width and height of the predicted bounding box, the loss of the predicted class, and the loss of the predicted confidence.
[0011] Furthermore, the defect detection model is obtained by the following steps:
[0012] S1: Construct a fabric image defect detection dataset, preprocess the image data and divide it into a training set, a validation set, and a test set;
[0013] S2: Input the image data into the multi-scale feature map extraction module to obtain four groups of feature maps with different resolutions;
[0014] S3: Input the four groups of feature maps with different resolutions into the multi-scale feature fusion module to obtain three groups of fused feature maps with different resolutions;
[0015] S4: Input the three groups of fused feature maps with different resolutions into the multi-scale feature enhancement module to obtain two groups of enhanced feature maps with different resolutions;
[0016] S5: Input the first group of fused feature maps generated in step S3 and the two groups of enhanced feature maps generated in step S4 into the detection head to obtain the predicted results of the center coordinates, the predicted results of the width and height of the bounding box, the predicted results of the class, and the predicted results of the confidence;
[0017] S6: According to the predicted results and the manually annotated results, calculate the loss of the predicted center coordinates, the loss of the width and height of the predicted bounding box, the loss of the predicted class, and the loss of the predicted confidence, perform gradient backpropagation, update the parameters, and optimize the predicted results of the bounding box and the classification predicted results;
[0018] S7: Repeat steps S2 to S6, iterate the training until the model converges, and the training stage ends;
[0019] S8: Select the optimal hyperparameters for the model using the validation set;
[0020] S9: Evaluate the generalization performance of the model using the test set.
[0021] Furthermore, the multi-scale feature extraction module is stacked by four sub-modules. The first sub-module consists of a patch embedding layer and a lightweight encoder layer, and the latter three modules both consist of a Haar wavelet attention downsampling (HWAD) layer and a lightweight encoder layer. Each module outputs a feature map of one scale, and finally four different-scale feature maps are obtained.
[0022] Furthermore, for a given fabric image H and W respectively represent the height and width of the fabric image; after passing through the multi-scale feature extraction module, four different-scale feature maps are obtained The sizes of the feature maps from shallow to deep are as follows: C is a preset hyperparameter; among them, the high-resolution feature map in the shallow layer mainly contains the fine position features of the image, and the low-resolution feature map in the deep layer mainly contains the semantic features of the image.
[0023] Furthermore, the patch embedding layer can be formally expressed as:
[0024]
[0025] where, PE represents the patch embedding layer, and finally a feature of one scale is obtained
[0026] Furthermore, the Haar wavelet attention downsampling (HWAD) layer consists of two parts: Haar wavelet transform and channel-spatial attention. The operation of the HWAD layer can be formally expressed as:
[0027]
[0028] where, i ∈ {1, 2, 3}, HWT represents the Haar wavelet transform, CSA represents the channel-spatial attention, and finally features of three scales are obtained Compared with the traditional downsampling method, the Haar wavelet attention downsampling method can better retain information, improve the model performance, and achieve a better balance in terms of the number of parameters and computational complexity.
[0029] Furthermore, the channel-spatial attention CSA operation can be formally expressed as:
[0030]
[0031] where, It is the result of the Haar wavelet transform output. SA represents the spatial attention operation, CA represents the channel attention operation, and [,] represents the concatenation operation along the channel dimension.
[0032] Furthermore, the spatial attention operation SA and the channel attention operation CA can be formally expressed as follows:
[0033]
[0034] Among them, Conv 1*1 represents a convolution operation with a kernel size of 1*1. represents a convolution operation with a kernel size of . represents element-wise multiplication.
[0035] Furthermore, for the fabric defect detection method described above, the lightweight encoder layer follows the design of the original Transformer encoder layer, but the attention mechanism is replaced by the attention mechanism implemented using efficient convolution operations in Conv2former. The operation of the lightweight encoder layer can be formally expressed as follows:
[0036]
[0037] Among them, i ∈ {1, 2, 3, 4}, Linear represents a linear layer, and DConv 3*3 represents a Deep-Wise convolution with a kernel size of 3*3. represents element-wise multiplication.
[0038] Furthermore, the multi-scale feature fusion module (MFFM) consists of three independent sub-modules. Each module will fuse the feature maps of adjacent two scales and output the fused feature maps; finally, three different scales of fused feature maps are obtained.
[0039] Furthermore, after passing through the multi-scale feature fusion module, three different scales of fused feature maps are obtained The sizes of the fused feature maps from shallow to deep are as follows:
[0040] Furthermore, the operation of the multi-scale feature fusion module MFFM adopts a two-stage fusion method and can be formally expressed as follows:
[0041]
[0042]
[0043] Among them, i ∈ {1, 2, 3}, Conv 3*3Denotes a convolution operation with a convolution kernel size of 3*3, Conv 1*1 Denotes a convolution operation with a convolution kernel size of 1*1, and POOL denotes a pooling operation.
[0044] Furthermore, the pooling operation POOL can be formally expressed as:
[0045]
[0046] Among them, Conv 1*1 Denotes a convolution operation with a convolution kernel size of 1*1, GAP denotes a global average pooling operation, and GMP denotes a global maximum pooling.
[0047] Furthermore, the multi-scale feature enhancement module (MFEM) consists of two cascaded sub-modules, which are used to further fuse the feature maps output by the multi-scale feature fusion module and output enhanced feature maps; finally, two different-scale enhanced feature maps are obtained.
[0048] Furthermore, after passing through the multi-scale feature enhancement module, two different-scale enhanced feature maps are obtained The sizes of the enhanced feature maps from shallow to deep are as follows:
[0049] Furthermore, the operation of the multi-scale feature enhancement module MFEM can be formally expressed as:
[0050]
[0051] Among them, C2F represents the C2F module of the YOLO model, CBS represents the combination of convolution, batch normalization, and SiLU activation function, and [,] represents the concatenation operation along the channel dimension.
[0052] Furthermore, the detection head uses the detection head of the YOLO model, and its input data includes The output data includes the prediction results of the center coordinates, the prediction results of the width and height of the bounding box, the prediction results of the category, and the prediction results of the confidence level.
[0053] Compared with the prior art, the beneficial effects of the present invention are:
[0054] 1. By adopting the Haar wavelet downsampling method, the present invention effectively reduces the loss of feature information caused by traditional downsampling operations, and at the same time maximally retains important texture features, enhancing the model's ability to capture details.
[0055] 2. The present invention introduces a lightweight self-attention mechanism in the Transformer encoder layer, reducing the computational amount and improving the processing efficiency.
[0056] 3. The present invention introduces a multi-scale feature fusion module, which combines features of different resolutions through a two-stage fusion method, alleviates the problem of unrecognizability due to insufficiently rich features, enhances the model's recognition ability for targets of different sizes, and optimizes the overall performance of the model.
[0057] 4. The present invention introduces a multi-scale feature enhancement module, which increases the non-linear representation ability of the model, enables the model to more accurately capture the features of defects, and thus realizes more accurate defect detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0059] Figure 1 is a flowchart of the implementation of the present invention;
[0060] Figure 2 is a structural diagram of the fabric defect detection model designed by the present invention;
[0061] Figure 3 is a structural diagram of the Haar wavelet attention downsampling (HWAD) layer designed by the present invention;
[0062] Figure 4 is a structural diagram of the multi-scale feature fusion module (MFFM) designed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0063] The following will describe the exemplary embodiments of the present invention in more detail with reference to the drawings. Although the exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present invention can be more thoroughly understood and the scope of the present invention can be completely conveyed to those skilled in the art. It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the drawings and in conjunction with the embodiments.
[0064] The present invention discloses a fabric defect detection method based on an improved YOLO model, including the following:
[0065] Obtain at least one fabric image on a fabric inspection machine;
[0066] Input the fabric image into the defect detection model to obtain the defect detection result of the fabric image output by the defect detection model.
[0067] The defect detection model is based on the YOLO model. Its main structure includes a multi-scale feature map extraction module, a multi-scale feature fusion module, a multi-scale feature enhancement module, and a detection head. The multi-scale feature map extraction module is stacked by four sub-modules. The first sub-module consists of a block embedding layer and a lightweight encoder layer, and the latter three sub-modules consist of a Haar wavelet attention downsampling layer and a lightweight encoder layer. The multi-scale feature fusion module is used to fuse the multi-scale feature maps extracted by the multi-scale feature map extraction module. It consists of three independent sub-modules, and each sub-module fuses the feature maps of two adjacent scales. The multi-scale feature enhancement module consists of two sub-modules in series, which is used to further fuse the feature maps output by the multi-scale feature fusion module. The detection head sets multiple losses, including the loss of the predicted center coordinates, the loss of the width and height of the predicted bounding box, the loss of the predicted category, and the loss of the predicted confidence.
[0068] The overall inference process of the defect detection model is as follows: Input the image data into the multi-scale feature map extraction module to obtain four groups of feature maps with different resolutions; Input the four groups of feature maps with different resolutions into the multi-scale feature fusion module to obtain three groups of fused feature maps with different resolutions; Input the three groups of fused feature maps with different resolutions into the multi-scale feature enhancement module to obtain two groups of enhanced feature maps with different resolutions; Input the first group of fused feature maps and the two groups of enhanced feature maps into the detection head to obtain the predicted results of the center coordinates, the predicted results of the width and height of the bounding box, the predicted results of the category, and the predicted results of the confidence.
[0069] The training process of the defect detection model is as Figure 1 、 Figure 2 shown. The specific steps are as follows:
[0070] S1: Construct a fabric image defect detection data set, preprocess the image data and divide it into a training set, a validation set, and a test set.
[0071] Specifically, first, an HD camera (high-resolution camera) is installed on the fabric inspection machine to capture fabric images, and the fabric images are divided into square region images of the same size with the moving speed of the fabric inspection machine * the time interval between two adjacent captures as the side length; then, through manual annotation, the defects in each square region image are annotated in the form of a rectangular box and the defect category information is set; finally, the dataset is divided. The square region images are arranged in ascending order of the acquisition time. A certain proportion (such as 60%) of the square region images acquired earliest are used as the training set, a certain proportion (such as 20%) of the square region images acquired during the middle period are used as the validation set, and a certain proportion (such as 20%) of the square region images acquired latest are used as the test set.
[0072] S2. Input the image data into the multi-scale feature map extraction module to obtain four groups of feature maps with different resolutions.
[0073] Specifically, for a given fabric image H and W respectively represent the height and width of the fabric image; after passing through the multi-scale feature extraction module, four different scales of feature maps are obtained The sizes of the feature maps from the shallow layer to the deep layer are in turn: C is a preset hyperparameter. Among them, the high-resolution feature map in the shallow layer mainly contains the fine position features of the image, and the low-resolution feature map in the deep layer mainly contains the semantic features of the image.
[0074] It should be noted that the multi-scale feature extraction module is stacked by four sub-modules. Among them, the first sub-module consists of a patch embedding layer and a lightweight encoder layer, and the latter three modules are all composed of a Haar wavelet attention downsampling (HWAD) layer and a lightweight encoder layer. Each module outputs a feature map of one scale, and finally four different scales of feature maps are obtained.
[0075] It should be noted that the embedding layer can be formally expressed as:
[0076]
[0077] Among them, PE represents the patch embedding layer, and finally a feature of one scale is obtained
[0078] It should be noted that as Figure 3 shown, the Haar wavelet attention downsampling (HWAD) layer consists of two parts: the Haar wavelet transform and the channel-spatial attention. The operation of the HWAD layer can be formally expressed as:
[0079]
[0080] where \(i\in\{1,2,3\}\), HWT represents the Haar wavelet transform, CSA represents channel-spatial attention, and finally features at three scales are obtained. Compared with traditional downsampling methods, the Haar wavelet attention downsampling method can better preserve information, improve the model performance, and achieve a good balance in terms of the number of parameters and computational complexity.
[0081] It should be noted that the channel-spatial attention CSA operation can be formally expressed as:
[0082]
[0083] where is the result of the Haar wavelet transform output, SA represents the spatial attention operation, CA represents the channel attention operation, and \([,]\) represents the concatenation operation along the channel dimension.
[0084] It should be noted that the spatial attention operation SA and the channel attention operation CA can be formally expressed as:
[0085]
[0086] where Conv 1*1 represents a convolution operation with a kernel size of \(1\times1\), represents a convolution operation with a kernel size of and represents element-wise multiplication.
[0087] It should be noted that the lightweight encoder layer follows the design of the original Transformer encoder layer, but the attention mechanism is replaced by the attention mechanism implemented using efficient convolution operations in Conv2former. The operation of the lightweight encoder layer can be formally expressed as:
[0088]
[0089] where \(i\in\{1,2,3,4\}\), Linear represents the linear layer, DConv 3*3 represents a depth-wise convolution with a kernel size of \(3\times3\), and
[0090] For \(S3\), four groups of feature maps with different resolutions are input into the multi-scale feature fusion module to obtain three groups of fused feature maps with different resolutions.
[0091] Specifically, after passing through the multi-scale feature fusion module, three different scales of fused feature maps are obtained. The sizes of the fused feature maps from the shallow layer to the deep layer are in turn:
[0092] It should be noted that the multi-scale feature fusion module (MFFM) consists of three independent sub-modules. Each module fuses the feature maps of two adjacent scales and outputs the fused feature maps; finally, three different scales of fused feature maps are obtained.
[0093] It should be noted that as Figure 4 shown, the operation of the multi-scale feature fusion module MFFM adopts a two-stage fusion method and can be formally expressed as:
[0094]
[0095] where \(i\in\{1,2,3\}\), Conv 3*3 represents a convolution operation with a convolution kernel size of \(3\times3\), Conv 1*1 represents a convolution operation with a convolution kernel size of \(1\times1\), and POOL represents a pooling operation.
[0096] It should be noted that the pooling operation POOL can be formally expressed as:
[0097]
[0098] where Conv 1*1 represents a convolution operation with a convolution kernel size of \(1\times1\), GAP represents a global average pooling operation, and GMP represents a global maximum pooling.
[0099] Step S4, input the three groups of fused feature maps with different resolutions into the multi-scale feature enhancement module to obtain two groups of enhanced feature maps with different resolutions.
[0100] Specifically, after passing through the multi-scale feature enhancement module, two different scales of enhanced feature maps are obtained The sizes of the enhanced feature maps from shallow to deep are in turn:
[0101] It should be noted that the multi-scale feature enhancement module (MFEM) consists of two cascaded sub-modules, which are used to further fuse the fused feature maps output by the multi-scale feature fusion module and output the enhanced feature maps; finally, two different scales of enhanced feature maps are obtained.
[0102] It should be noted that the operation of the multi-scale feature enhancement module MFEM can be formally expressed as:
[0103]
[0104] Among them, C2F represents the C2F module of the YOLO model, CBS represents the combination of convolution, batch normalization, and the SiLU activation function, and [,] represents the concatenation operation along the channel dimension.
[0105] S5: Input the first set of fused feature maps generated in step S3 and the two sets of enhanced feature maps generated in step S4 into the detection head to obtain the prediction results of the center coordinates, the prediction results of the width and height of the bounding box, the prediction results of the class, and the prediction results of the confidence.
[0106] Specifically, the detection head uses the detection head of the YOLO model, and its input data includes The output data includes the prediction results of the center coordinates, the prediction results of the width and height of the bounding box, the prediction results of the class, and the prediction results of the confidence.
[0107] S6: According to the prediction results and the manually annotated results, calculate the loss of the predicted center coordinates, the loss of the width and height of the predicted bounding box, the loss of the predicted class, and the loss of the predicted confidence, perform gradient backpropagation, update the parameters, and optimize the prediction results of the bounding box and the classification prediction results.
[0108] S7: Repeat steps S2 to S6, iterate the training until the model converges, and the training phase ends.
[0109] S8: Use the validation set to select the optimal hyperparameters for the model.
[0110] Specifically, set different hyperparameters for the model to obtain different model instances; for each model instance, perform training on the training set according to steps S2 to S7 and perform inference on the validation set, and calculate the performance metrics of each model instance on the validation set; compare the performance metrics of each model instance on the validation set, and select the model instance with the optimal performance metrics as the finally used model.
[0111] S9: Evaluate the generalization performance of the model using the test set.
[0112] Specifically, for the finally used model selected through the validation set, perform model inference on the test set, and calculate the performance metrics of the model on the test set.
[0113] The present invention has been described in detail through the embodiments above. However, the above content is only an exemplary embodiment of the present invention and cannot be considered as defining the scope of implementation of the present invention. The protection scope of the present invention is defined by the claims. Any use of the technical solutions described in the present invention, or any technical solutions designed by those skilled in the art inspired by the technical solutions of the present invention, within the essence and protection scope of the present invention, which achieve the above technical effects by designing similar technical solutions, or any equivalent changes and improvements made to the application scope, shall still fall within the patent coverage protection scope of the present invention.
Claims
1. A fabric defect detection method based on an improved YOLO model, characterized in that: The following steps are involved: Acquiring a fabric image on at least one fabric inspection machine; Inputting the fabric image into a defect detection model to obtain a defect detection result of the fabric image output by the defect detection model; The defect detection model includes a multi-scale feature map extraction module, a multi-scale feature fusion module, a multi-scale feature enhancement module, and a detection head; the multi-scale feature map extraction module is composed of four stacked sub-modules, the first sub-module is composed of a block embedding layer and a lightweight encoder layer, and the latter three sub-modules are composed of a Haar wavelet attention downsampling layer and a lightweight encoder layer; the multi-scale feature fusion module is used to fuse the multi-scale feature map extracted by the multi-scale feature map extraction module, and it is composed of three independent sub-modules, each sub-module fuses the feature maps of two adjacent scales; the multi-scale feature enhancement module is composed of two serially connected sub-modules, and is used to further fuse the feature map output by the multi-scale feature fusion module; the detection head sets multiple losses, including the loss of the predicted center coordinates, the loss of the width and height of the predicted bounding box, the loss of the predicted category, and the loss of the predicted confidence.
2. The fabric defect detection method based on the improved YOLO model according to claim 1, characterized in that: The detection method is as follows: S1: Construct a fabric image defect detection dataset, preprocess the image data and divide it into training set, validation set and test set; S2: Input the image data into the multi-scale feature map extraction module to obtain four sets of feature maps with different resolutions; S3: Input the four sets of feature maps with different resolutions into the multi-scale feature fusion module to obtain three sets of fused feature maps with different resolutions; S4: Input the three sets of fused feature maps with different resolutions into the multi-scale feature enhancement module to obtain two sets of enhanced feature maps with different resolutions; S5: Input the first set of fused feature maps generated in step S3 and the two sets of enhanced feature maps generated in step S4 into the detection head to obtain the prediction results of the center coordinates, the prediction results of the width and height of the bounding box, the prediction results of the category, and the prediction results of the confidence level; S6: Based on the prediction results and manual annotation results, calculate the loss of the predicted center coordinates, the loss of the width and height of the predicted bounding box, the loss of the predicted category, and the loss of the predicted confidence, perform gradient backpropagation, update parameters, and optimize the bounding box prediction results and classification prediction results; S7: Repeat steps S2 to S6, iterate the training until the model converges, and the training phase ends; S8: Use the validation set to select the optimal hyperparameters for the model; S9: Use the test set to evaluate the generalization performance of the model.
3. The fabric defect detection method based on the improved YOLO model according to claim 2, characterized in that: The multi-scale feature extraction module is composed of four stacked sub-modules; The first sub-module consists of a block embedding layer and a lightweight encoder layer, and the last three modules are composed of a Haar wavelet attention downsampling layer and a lightweight encoder layer. Each module outputs a feature map of one scale, and finally four feature maps of different scales are obtained.
4. The fabric defect detection method based on the improved YOLO model according to claim 3 is characterized in that: For a given fabric image H and W represent the height and width of the fabric image respectively; after passing through the multi-scale feature extraction module, four feature maps of different scales are obtained: The feature map sizes from shallow to deep are: C is a pre-set hyperparameter; the shallow high-resolution feature map mainly contains the fine position features of the image, and the deep low-resolution feature map mainly contains the semantic features of the image; The block embedding layer can be formally represented as: Among them, PE represents the block embedding layer, and finally obtains a scale feature The Haar wavelet attention downsampling layer consists of two parts: Haar wavelet transform and channel-spatial attention. The operation of the HWAD layer can be formally expressed as: Among them, i∈{1,2,3}, HWT represents Haar wavelet transform, CSA represents channel-space attention, and finally the features of three scales are obtained Compared with traditional downsampling methods, the Haar wavelet attention downsampling method can better retain information and improve model performance, while achieving a good balance between the number of parameters and the amount of computation.
5. The fabric defect detection method based on the improved YOLO model according to claim 4, characterized in that: The channel-spatial attention CSA operation can be formally expressed as: in, is the result of Haar wavelet transform output, SA represents the spatial attention operation, CA represents the channel attention operation, and [,] represents the concatenation operation along the channel dimension; The spatial attention operation SA and the channel attention operation CA can be formally expressed as: Among them, Conv 1*1 Indicates a convolution operation with a convolution kernel size of 1*1. Indicates that the convolution kernel size is The convolution operation, Represents the multiplication of corresponding elements.
6. The fabric defect detection method based on the improved YOLO model according to claim 3, characterized in that: The lightweight encoder layer follows the design of the original Transformer encoder layer, but the attention mechanism is replaced by the attention mechanism implemented in Conv2former using efficient convolution operations. The operation of the lightweight encoder layer can be formally expressed as: Among them, i∈{1,2,3,4}, Linear represents the linear layer, DConv 3*3 It represents the Deep-Wise convolution with a convolution kernel size of 3*3. Represents the multiplication of corresponding elements.
7. The fabric defect detection method based on the improved YOLO model according to claim 2, characterized in that: The multi-scale feature fusion module (MFFM) consists of three independent sub-modules. Each module fuses the feature maps of two adjacent scales and outputs a fused feature map. Finally, three fused feature maps of different scales are obtained. After passing through the multi-scale feature fusion module, three fused feature maps of different scales are obtained. The sizes of fused feature maps from shallow to deep layers are: The operation of the multi-scale feature fusion module MFFM adopts a two-stage fusion method, which can be formally expressed as: Among them, i∈{1,2,3}, Conv 3*3 Indicates a convolution operation with a convolution kernel size of 3*3, Conv 1*1 Represents a convolution operation with a convolution kernel size of 1*1, and POOL represents a pooling operation; The pooling operation POOL can be formally expressed as: Among them, Conv 1*1 It represents a convolution operation with a convolution kernel size of 1*1, GAP represents a global average pooling operation, and GMP represents a global maximum pooling operation.
8. The fabric defect detection method based on the improved YOLO model according to claim 2, characterized in that: The multi-scale feature enhancement module (MFEM) is composed of two serially connected sub-modules, and is used to further fuse the fused feature map output by the multi-scale feature fusion module and output an enhanced feature map; Finally, two enhanced feature maps of different scales are obtained.
9. The fabric defect detection method based on the improved YOLO model according to claim 8, characterized in that: After the multi-scale feature enhancement module, two enhanced feature maps of different scales are obtained. The sizes of the enhanced feature maps from shallow to deep layers are: The operation of the multi-scale feature enhancement module MFEM can be formally expressed as: Among them, C2F represents the C2F module of the YOLO model, CBS represents the combination of convolution, batch normalization and SiLU activation functions, and [,] represents the concatenation operation along the channel dimension.
10. The fabric defect detection method based on the improved YOLO model according to claim 2, characterized in that: The detection head uses the detection head of the YOLO model, and its input data includes The output data includes the prediction results of the center coordinates, the prediction results of the width and height of the bounding box, the prediction results of the category, and the prediction results of the confidence.