Fabric printing and dyeing defect multi-scale lightweight detection method and system based on improved YOLO
By introducing technologies such as multi-scale feature extraction, context diffusion and Efficient Local Attention in fabric printing and dyeing defect detection, and combining channel pruning, the YOLOv5 model is optimized, which solves the problems of large size spans and feature similarity of fabric defects, and achieves efficient and accurate fabric printing and dyeing defect detection.
Patent Information
- Application Number
- CN202510214722.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-24
AI Technical Summary
The size span of fabric defects is large, which brings challenges to the feature extraction of the network. The feature similarity between different defect categories is high, increasing the difficulty of the model in the target classification process.
A lightweight improved algorithm called MCF-YOLOv5 is proposed. Through key technologies such as multi-scale feature extraction, context diffusion, Efficient Local Attention and channel pruning, the fabric printing and dyeing defect detection performance is comprehensively optimized.
Through the multi-scale context aggregation module and the multi-scale context diffusion fusion pyramid network, the detection capability of defects with large aspect ratio differences is improved, the detection accuracy is improved, and the parameter amount and calculation amount of the model are significantly reduced through channel pruning technology, achieving efficient real-time detection.
Smart Images

Figure CN120198368A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of fabric printing and dyeing defect detection, and specifically to a multi-scale lightweight detection method and system for fabric printing and dyeing defects based on improved YOLO. Background Art
[0002] Fabric defect detection is a key link in textile production to ensure fabric quality. Traditional detection methods mainly rely on manual visual inspection. Although simple and easy to implement, they have disadvantages such as low detection efficiency and being greatly affected by human factors, and it is difficult to meet the requirements of rapid production in modern industry. In recent years, object detection technology based on convolutional neural networks has been widely used in the field of fabric defect detection due to its excellent real-time detection performance and powerful non-linear feature representation ability. However, the size span of fabric defects poses challenges to the feature extraction of the network. In addition, the feature similarity between different defect categories is relatively high, further increasing the difficulty of the model in the object classification process.
[0003] To solve the above problems, the present invention proposes a lightweight improved algorithm called MCF-YOLOv5. Through key technologies such as multi-scale feature extraction, context diffusion, Efficient Local Attention, and channel pruning, the detection performance of fabric printing and dyeing defects is comprehensively optimized. The multi-scale context aggregation module accepts multi-scale inputs and uses parallel large-kernel depth convolutions to capture the global context information of multi-scale features, effectively improving the detection ability for defects with large aspect ratio differences. The multi-scale context diffusion fusion pyramid network reconstructs the feature fusion path, extends the features extracted by the multi-scale context aggregation module to the shallow and deep feature spaces, avoids the loss of feature information at each scale, and thus obtains a richer and more discriminative feature representation. The C3-Star module enhances the information expression ability of features under a large receptive field by introducing element-wise multiplication and large-kernel depth convolutions in the C3 residual module. Embedding EfficientLocal Attention in the backbone network further strengthens the model's perception ability of key region features by dynamically weighting features at different spatial positions. In addition, through the channel pruning technology, the number of model parameters is significantly reduced, the inference speed is improved, and a high detection accuracy is maintained, overcoming the deficiencies of traditional algorithms such as large computational amount, slow detection speed, and strong dependence on GPU. Finally, the multi-scale lightweight model is deployed to mobile devices, effectively reducing the equipment cost and solving the problem of expensive traditional detection equipment. The present invention is tested on a self-made fabric printing and dyeing defect dataset, and finally the model with the best performance is selected for deployment, verifying the practicability and effectiveness of the algorithm. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the present invention provides a multi-scale lightweight detection method and system for fabric printing and dyeing defects based on improved YOLO, which solves the problem that the size span of fabric defects is relatively large and it is inconvenient to extract features of the network.
[0005] To achieve the above objectives, the present invention is realized through the following technical solutions: A multi-scale lightweight detection method for fabric printing and dyeing defects based on improved YOLO, comprising the following steps:
[0006] S1. On the premise of ensuring the same production conditions, collect fabric printing and dyeing defect pictures containing different defect types to ensure the diversity and representativeness of the data.
[0007] S2. Label the collected fabric printing and dyeing defect pictures according to the corresponding defect categories and divide the data set to obtain sample image data;
[0008] S3. Perform data augmentation on the sample image data;
[0009] S4. Use a multi-scale context aggregation module, a multi-scale context diffusion fusion pyramid network, a C3-Star module, and an Efficient Local Attention to innovate the structure of the YOLOv5 model;
[0010] S5. Use the sample image data to train the improved YOLOv5s model;
[0011] S6. Perform structured pruning on the trained improved YOLO model to obtain a multi-scale lightweight model;
[0012] S7. Deploy the multi-scale lightweight model to the target mobile device for fabric printing and dyeing defect detection.
[0013] Preferably, the method for collecting fabric printing and dyeing defect pictures includes:
[0014] On the premise of ensuring the same production conditions, collect fabric printing and dyeing defect pictures containing different defect types to ensure the diversity and representativeness of the data.
[0015] Preferably, the method for labeling the collected fabric printing and dyeing defect pictures according to the corresponding defect categories and dividing the data set to obtain sample image data includes:
[0016] Use the minimum bounding rectangle in the LabelImg software to label the category and position of each fabric printing and dyeing defect. Among them, class represents the category of the defect, xmin and ymin are the coordinates of the upper left vertex of the minimum bounding rectangle, and xmax and ymax are the coordinates of the lower right vertex, so as to accurately describe the type of the defect and its specific position in the image.
[0017] Preferably, the method for performing data enhancement on sample image data includes:
[0018] One or more operations of random rotation, random scaling, random cropping, dynamic adjustment of brightness and contrast, and Mosaic enhancement are applied to the sample image data. After the operation is completed, the sample image data is associated with its corresponding defect category, so as to generate new sample image data to enrich the data set and improve the generalization ability of the model.
[0019] Preferably, the method of structurally innovating the YOLOv5 model by using a multi-scale context aggregation module, a multi-scale context diffusion fusion pyramid network, a C3-Star module, and Efficient Local Attention includes:
[0020] Use the multi-scale context aggregation module to obtain the multi-scale context information of features;
[0021] A multi-scale context diffusion fusion pyramid network is used to replace the Neck network of the YOLOv5 model, and the output of the multi-scale context aggregation module is diffused to all detection scales, so that each detection scale has rich multi-scale context information.
[0022] The C3-Star module is used to replace the C3 module of the YOLOv5 model, and the abstract semantic information of the features is extracted through high-dimensional space mapping to enhance the feature extraction capability of the model.
[0023] Use Efficient Local Attention to dynamically enhance the model in different areas, thereby focusing on the key areas where defects are located.
[0024] Preferably, the method of performing structured pruning on the trained improved YOLO model to obtain a multi-scale lightweight model includes:
[0025] Sparse regularization is introduced into the trained improved YOLO model to promote the sparsification of redundant parameters in the network;
[0026] Set a Speed_up parameter to control the pruning rate of the model, where Speed_up is defined as:
[0027]
[0028] Among them, FLOPs1 is the floating-point computing amount of the unpruned model, and FLOPs2 is the floating-point computing amount of the pruned model. By adjusting the value of Speed_up, the network slimming ratio can be flexibly controlled, thereby reducing the computational complexity of the model while maintaining its detection performance as much as possible.
[0029] Preferably, the method for obtaining a multi-scale lightweight model by performing structured pruning on the trained improved YOLO model further includes the steps of:
[0030] Training with sample image data on improved YOLOv5 models with different depths and widths, including: improved YOLOv5n, improved YOLOv5s, and improved YOLOv5m, to generate multiple trained improved YOLO models;
[0031] Performing performance evaluation and comparison on multiple trained models, and selecting the model with the best performance as the final trained improved YOLO model according to the detection accuracy, inference speed, and multi-scale lightweight effect indicators of the models.
[0032] Preferably, the method for performing performance comparison on multiple trained improved YOLO models includes:
[0033] Evaluating and comparing multiple trained improved YOLO models using common performance evaluation indicators in the field of object detection: precision (P), recall (R), mean average precision (mAP), floating point operations (FLOPs), number of parameters (Parameters), frames per second (FPS), and inference latency (Latency), to comprehensively measure the detection performance, computational complexity, and running efficiency of the models.
[0034] Preferably, the method for deploying the multi-scale lightweight model to a target mobile device for fabric printing and dyeing defect detection includes:
[0035] Replacing the YOLO model file in the pre-imported original detection software with the relevant files of the improved YOLO model; checking and validating the input and output data dimensions of the model, and adjusting the input and output dimensions if necessary to match the actual requirements, while modifying the label file to ensure consistency with the detection categories of the model;
[0036] Setting the name and icon of the target APP, generating and exporting an installation package containing the new model; importing and installing the installation package onto the target mobile device to generate the target APP;
[0037] Opening the target APP and running the multi-scale lightweight model to achieve efficient detection of fabric printing and dyeing defects.
[0038] A multi-scale lightweight detection system for fabric printing and dyeing defects based on improved YOLO, includes:
[0039] A sample acquisition module for obtaining fabric printing and dyeing defect pictures and providing the original data required for detection;
[0040] A data annotation module for annotating the defect categories of the collected fabric printing and dyeing defect pictures and dividing the data to generate sample image data;
[0041] A data augmentation module that performs various data augmentation operations on the sample image data to improve the robustness and generalization ability of the model;
[0042] A model structure optimization module that uses a multi-scale context aggregation module, a multi-scale context diffusion fusion pyramid network, a C3-Star module, and an Efficient Local Attention to innovate the structure of the YOLOv5 model and obtain an improved YOLOv5s model;
[0043] A training module that uses the augmented sample data to train the improved YOLOv5s model;
[0044] A pruning module used to perform structured pruning on the trained improved YOLO model to obtain a multi-scale lightweight model;
[0045] A deployment module used to deploy the multi-scale lightweight model to a target mobile device for fabric printing and dyeing defect detection.
[0046] The present invention provides a multi-scale lightweight detection method and system for fabric printing and dyeing defects based on improved YOLO. It has the following beneficial effects:
[0047] 1. By targeting the multi-scale characteristics of fabric printing and dyeing defects, the present invention uses a multi-scale context aggregation module to obtain the global context features of defect features, and uses a multi-scale context diffusion fusion pyramid network to spread these global context features to each detection scale, further improving the model's detection ability for multi-scale targets.
[0048] 2. By targeting the characteristics that fabric printing and dyeing defects are relatively similar to the background, the present invention uses a C3-Star module and an Efficient Local Attention to enhance the model's feature extraction ability and the positioning ability of the target area, accurately obtaining the position and contour information of the defects, thereby effectively improving the detection accuracy of the model.
[0049] 3. By introducing sparse regularization to the trained improved YOLO model and using structured pruning technology, the present invention greatly reduces the model's parameter quantity and computational complexity, realizing the slimming optimization of the model. The multi-scale lightweight model after pruning has been greatly improved in inference speed, providing the possibility for real-time detection on mobile devices.
[0050] 4. Through the innovative design of the multi-scale lightweight model, the present invention is successfully deployed on mobile devices with limited computing power, significantly reducing the hardware device cost and the difficulty of industrial deployment, and meeting the requirements of the textile printing and dyeing industry for real-time performance and efficiency. Combining with the high-performance performance of the multi-scale lightweight model, it greatly improves the industrial production efficiency, reduces the defective rate in the textile printing and dyeing process, and promotes the intelligent upgrading of the industry. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 It is a flowchart of a multi-scale lightweight detection method for fabric printing and dyeing defects of the present invention;
[0052] Figure 2 It is a schematic diagram of the overall structure of the YOLOv5s model improved by the present invention;
[0053] Figure 3 It is a schematic diagram of the module structure of the improved multi-scale context aggregation module of the present invention;
[0054] Figure 4 It is a schematic diagram of the structure of the multi-scale context diffusion fusion pyramid network of the present invention;
[0055] Figure 5 It is a schematic diagram of the working principle of C3-Star of the present invention;
[0056] Figure 6 It is a schematic diagram of the working principle of Efficient Local Attention of the present invention;
[0057] Figure 7 It is a schematic diagram of the structured pruning principle of the present invention;
[0058] Figure 8 It is a schematic diagram of the structure of a multi-scale lightweight detection system for fabric printing and dyeing defects of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0059] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0060] Embodiment 1:
[0061] Please refer to the attached Figure 1 - attached Figure 8 , the embodiment of the present invention provides a multi-scale lightweight detection method for fabric printing and dyeing defects based on the improved YOLO, including the following steps:
[0062] S1. Collect fabric printing defect pictures;
[0063] S2. Label the collected fabric printing defect pictures according to the corresponding defect categories and divide the dataset to obtain sample image data;
[0064] S3. Perform data augmentation on the sample image data;
[0065] S4. Use a multi-scale context aggregation module to obtain multi-scale context information of features;
[0066] S5. Use a multi-scale context diffusion fusion pyramid network to replace the Neck network of the YOLOv5 model, and diffuse the output of the multi-scale context aggregation module to all detection scales, so that each detection scale has rich multi-scale context information.
[0067] S6. Use the C3-Star module to replace the C3 module of the YOLOv5 model, and extract the abstract semantic information of features through high-dimensional space mapping to enhance the feature extraction ability of the model.
[0068] S7. Insert Efficient Local Attention into the backbone network of the model to achieve dynamic enhancement of different regions of the model, so as to focus on the key regions where the defects are located, and obtain an improved YOLOv5s model.
[0069] S8. Use the sample image data to train the improved YOLOv5s model;
[0070] S9. Perform structured pruning on the trained improved YOLO model to obtain a multi-scale lightweight model;
[0071] S10. Deploy the multi-scale lightweight model to the target mobile device for fabric printing defect detection.
[0072] In step S1, the method for collecting fabric printing defect pictures is as follows:
[0073] S1A: On the premise of ensuring the same production conditions, collect fabric printing defect pictures containing different defect types to ensure the diversity and representativeness of the data;
[0074] S1B: Obtain an image resolution of 1920×1080 and save it in png format.
[0075] In step S2, the method for labeling the collected fabric printing defect pictures according to the corresponding defect categories and dividing the dataset to obtain sample image data is as follows:
[0076] S2A: Use the minimum bounding rectangle in the LabelImg software to label the category and location of each fabric printing defect. Among them, the target minimum bounding rectangle is used to completely enclose the corresponding defect area, and the labeling format is (class, xmin, ymin, xmax, ymax), where class represents the category of the defect, xmin and ymin are the coordinates of the upper left vertex of the minimum bounding rectangle, and xmax and ymax are the coordinates of the lower right vertex, so as to accurately describe the type of the defect and its specific location in the image.
[0077] S2B: Divide all the data sets into a training set, a test set and a validation set according to the ratio of 7:2:1 and put them into the images folder, and put the corresponding txt files into the labels file.
[0078] In step S3, the method for data augmentation of the sample image data is as follows:
[0079] S3A: When processing fabric printing defect images, we noticed that there are defect targets of different sizes in the images. To address this issue, we proposed a data augmentation strategy, including operations such as random rotation, random scaling, random cropping, dynamic adjustment of brightness and contrast, and Mosaic augmentation. Through random rotation, we can process fabric printing defects at various angles and increase the diversity of the data; random scaling and cropping can enable the model to better adapt to defect targets of different scales; by dynamically adjusting the brightness and contrast, the robustness of the model can be increased to adapt to changes in lighting conditions; at the same time, Mosaic augmentation can splice multiple different images into a large image to increase the diversity of the training data. The improved YOLOv5 structure is as Figure 2 shown:
[0080] As Figure 3 reflects the principle of the multi-scale context aggregation module.
[0081] In step S4, the method for obtaining multi-scale context information of features using the multi-scale context aggregation module is as follows:
[0082] S4A: The multi-scale context aggregation module will receive three feature maps with different resolutions output by the backbone network
[0083]
[0084] Among them, P i represents the input feature, C i , H i , W i represent the number of channels, height and width of the feature map respectively, and i = 3, 4, 5.
[0085] S4B: To fuse multi-scale features, it is necessary to adjust the sizes of the three feature maps to be the same using upsampling or downsampling. However, in the case of a large difference in the sizes of the feature maps, consecutive upsampling or downsampling may lead to information loss or distortion. Therefore, we choose to make the sizes of P3 and P5 the same as that of P4 to avoid this phenomenon. In addition, each channel of the feature map usually represents different feature maps (such as texture, edge, color, etc.). If the number of channels is too large, there may be multiple channels representing similar information repeatedly, resulting in the model learning redundant features. Therefore, we introduce a channel scaling factor e = 0.5 to reduce the number of channels while adjusting the size of the feature map: C = C4 * e. The process of adjusting the size and number of channels of the feature map is shown as follows
[0086]
[0087] where DownSample represents the downsampling operation, which uses a convolution operation with a convolution kernel of 3, a stride of 2, and a padding of 1. UpSample represents the upsampling operation, which uses the nearest neighbor interpolation method, and Conv 1×1 represents the convolution operation with a convolution kernel of 1.
[0088] S4C: The adjusted feature map P' i is concatenated in the channel dimension, and then a set of parallel depth convolutions are used to capture the global context dependencies under different receptive fields. By combining convolution kernels of different sizes, the multi-scale context aggregation module can capture multi-level information from local to global in a single depth convolution layer, thereby improving the detection ability for targets of different sizes. The process is shown as follows
[0089] P Concat = Concat(P i ′ ), i = 3, 4, 5 #(3)
[0090] P l = DWConv m×m (P Concat ), m = 5, 7, 9, 11, l = 1, 2, 3, 4 #(4)
[0091] where Concat represents the concatenation operation, DWConv m×m represents the depth convolution with a convolution kernel size of m, and P l represents the output generated by the parallel depth convolution.
[0092] S4D: To effectively fuse the original features and the features extracted by deep convolution, residual connections are adopted in the channel dimension to add all features element-wise. Subsequently, 1×1 convolution is used to achieve information interaction between channels. However, the spatial receptive field of 1×1 convolution is relatively small, making it unable to capture and transmit local spatial details well. If the features are stacked or compressed multiple times, some key spatial details may be further lost. Therefore, we add the original features to the output of the 1×1 convolution again to ensure the integrity of the original features. This process is shown in the following formula
[0093] C = Add(Conv 1×1 (Add(P Concat , P l )), P Concat )#(5)
[0094] Among them, Add means adding the feature tensors element-wise in the channel dimension, and C is the output feature of the multi-scale context aggregation module.
[0095] As Figure 4 reflects the principle of the multi-scale context diffusion fusion pyramid network.
[0096] In step S5, the multi-scale context diffusion fusion pyramid network is used to replace the Neck network of the YOLOv5 model, and the output of the multi-scale context aggregation module is diffused to all detection scales to enable each detection scale to have rich multi-scale context information. The method is as follows
[0097] S5A: Let the outputs generated by the multi-scale context aggregation module be C4 and O4 respectively, as shown in the following formula
[0098] C4 = MCAM(P3, P4, P5)#(6)
[0099] O4 = MCAM(C3, C4, C5)#(7)
[0100] C4 and O4 are fused with shallow features and deep features in a diffused manner respectively, so that rich global context information is diffused to each detection scale. Among them, the global context information from the intermediate layer can enhance the ability to capture texture details in the shallow features and help detect smaller defects; while in the deep features, this context information provides stronger high-level semantic information, which is more effective for detecting larger-scale defects.
[0101] As Figure 5 reflects the principle of C3-Star.
[0102] In step S6, the method of using the C3-Star module to replace the C3 module of the YOLOv5 model and extracting the abstract semantic information of features through high-dimensional space mapping to enhance the feature extraction ability of the model is as follows:
[0103] S6A: The module divides the original features into two parts in the channel dimension through two 1×1 convolutions. One part is input into the StarBottleneck module for feature extraction, and the other part serves as an identity mapping.
[0104] S6B: In the StarBottleneck structure, depthwise convolutions with large kernels are used to expand the receptive field of the convolutional layer. Secondly, the relationship between the implicitly expanded feature space dimension and the channel dimension d of the input features through element-wise multiplication can be approximately expressed as Therefore, in order to fully expand the feature dimension, the features are first upsampled using two 1×1 convolutions, and then the two upsampled features are multiplied element-wise to achieve implicit high-dimensional space feature mapping. Moreover, element-wise multiplication also has a feature enhancement mechanism: when the feature values of two channels are both large at the same position, the product will be larger, thereby increasing the sensitivity to key features. After completing the high-dimensional feature mapping, 1×1 convolutions are used again to compress the feature dimension while retaining the abstract features in the high-dimensional space. Additionally, in the element-wise multiplication operation, if the element value of a certain feature map is very small, the multiplication result will cause the feature value at the corresponding position to disappear or become smaller, thereby resulting in information loss. Therefore, an identity mapping is performed on the original features to ensure the integrity of information transmission.
[0105] S6C: The identity mapping part is directly concatenated and fused with the output of the StarBottleneck module. Finally, 1×1 convolutions are used to adjust the number of channels of the feature map and enhance the information interaction between channels.
[0106] As Figure 6 reflects the principle of Efficient Local Attention.
[0107] In step S7, the method of inserting Efficient Local Attention before the SPPF module to achieve dynamic strengthening of different regions of the model, thereby focusing on the key region where the defect is located and obtaining the improved YOLOv5s model is as follows:
[0108] S7A: Efficient Local Attention uses global average pooling to extract features in the x and y directions. Assume the input feature of this module is F C×H×W, C, H, and W represent the number of channels, height, and width of the feature map respectively. 1D convolution is used to process the feature vectors in the x and y directions to enhance the interaction ability of the position information embedding and the long-distance dependence relationship. This process is shown as follows
[0109]
[0110] where x c (h, i) and x c (j, w) represent the input feature values in the x and y directions on the c channel respectively. and represent global average pooling in the x and y directions on the c channel respectively.
[0111] S7B: The output vector is normalized using group normalization, and the normalized feature vector is mapped to between [0, 1] through the sigmoid function to obtain the weight parameters output in the x and y directions after the feature passes through the Efficient Local Attention.
[0112]
[0113] where σ represents the sigmoid function and GN represents the group normalization method. and represent the weight parameters in the x and y directions on the c channel.
[0114] S7C: Multiply the input original feature map by the weight parameters in the two directions, enabling the network to precisely focus on the regions that are more crucial for the prediction task while suppressing the relatively unimportant information, thereby achieving dynamic adjustment of the importance of different regions on the feature map.
[0115]
[0116] where Out c represents the output of the Efficient Local Attention on the c channel, and x c (i, j) represents the original value of the input feature on the c channel.
[0117] S7D: Select the experimental environment configuration as shown in Table 1, set the training hyperparameters as shown in Table 2, and start training the improved model after setting the parameters.
[0118] Table 1 Experimental Environment Configuration
[0119]
[0120] Table 2 Experimental Hyperparameter Settings
[0121] Hyperparameter Information Number of iterations 300 Batch size 32 Optimizer SGD Initial learning rate 0.01 Number of GPUs 1 Weight decay coefficient 0.0005
[0122] In step S9, the trained improved YOLO model is subjected to structured pruning to obtain a multi-scale lightweight model as follows:
[0123] S9A: Structured pruning is used to slim down the model. First, sparse regularization is performed on the trained model. L1 regularization is added to the scaling factor of the batch normalization (BN) layer in the model. This will make the value of the scaling factor in BN approach 0 so that the network can distinguish the importance of channels (or neurons), because each scaling factor has a corresponding specific convolution channel (or neuron). The specific principle is as follows: Figure 7 shown.
[0124] S9B: Set a Speed_up parameter to control the pruning rate of the model, where Speed_up is defined as
[0125]
[0126] Among them, FLOPs1 is the floating-point computing amount of the unpruned model, and FLOPs2 is the floating-point computing amount of the pruned model. By adjusting the value of Speed_up, the network slimming ratio can be flexibly controlled, thereby reducing the computational complexity of the model while maintaining its detection performance as much as possible.
[0127] The multi-scale lightweight detection method for fabric printing and dyeing defects based on improved YOLO also includes the following steps:
[0128] S9'. Use sample image data to train the improved YOLOv5s model multiple times to obtain multiple trained improved YOLO models, compare the performance of multiple trained improved YOLO models, and select the best model as the final trained improved YOLO model. The specific method is as follows:
[0129] S9'A: Use the commonly used performance evaluation indicators in the field of object detection to evaluate and compare multiple trained improved YOLO models: precision (P), recall (R), average precision (mAP), floating point operations (FLOPs), parameters (Parameters), frames per second (FPS) and inference delay (Latency) to comprehensively measure the detection performance, computational complexity and operation efficiency of the model.
[0130] S9'B: The P value represents the probability that the model detects an image containing fabric printing and dyeing defects in the test set images as positive; the R value represents the ratio of fabric printing and dyeing defects detected by the model in the test set images to the real fabric printing and dyeing defects; the mAP value represents the average accuracy of the model in all categories, as shown in formulas (8)-(11)
[0131]
[0132]
[0133]
[0134]
[0135] Among them, TP (true positive) represents the number of fabric defects that can be correctly identified; FP (false positive) represents the number of fabric defects that are misidentified; FN (false negative) represents the number of fabric defects that are not detected. The average precision (AP) represents the detection precision of each fabric defect category of the model, and N represents the number of fabric defect types.
[0136] S9’C: This study also used FLOPs and Parameters values to evaluate the model complexity; FPS and Latency values to verify the inference speed of the model, and their calculation formulas are shown in (12)-(16)
[0137] FLOPs = W × H × K × K × C in × C out #(20)
[0138] Parameters = C in × C out × K × K#(21)
[0139]
[0140]
[0141]
[0142] Among them, K represents the size of the convolutional kernel; W and H represent the width and height of the input fabric defect feature map; C in × C out represent the number of input and output channels respectively; T s represents the total time of all test inferences; N t represents the total number of tests; BS represents the number of test images in each batch; ITPI represents the inference time of each image.
[0143] S9’D: First, a comparative experiment was conducted on the selection of the model size.
[0144] S9’E: Secondly, a comparative experiment was conducted on the selection of the attention mechanism to select the attention mechanism with the best effect.
[0145] S9’F: Finally, different pruning rates were set to select the optimal model.
[0146] In all the above experiments, the same experimental data need to be controlled, such as the backbone network channel size, attention mechanism, and experimental environment, etc.
[0147] In step S10, the method of deploying the multi-scale lightweight model to the target mobile device for fabric printing and dyeing defect detection is as follows:
[0148] S10A: Model quantization makes the model more adaptable to mobile device hardware, improves deployment efficiency and performance by reducing the model size, accelerating the inference speed, and reducing power consumption. Common methods include Float 16 quantization (reducing the floating-point precision from 32 bits to 16 bits), dynamic quantization (dynamically mapping weights and activation values to low precision during inference), and full integer quantization (quantizing weights and activation values to INT8, significantly compressing the model). For scenarios with limited mobile computing power, the input resolution can be reduced by setting --imgsz=320 to further reduce the computational overhead.
[0149] S10B: Replace the corresponding YOLO model file, view the model input and output data dimensions, modify the input and output dimensions and label files, set the name and icon of the mobile APP, and package it into an APK format installation package.
[0150] S10C: Download the Android Stdio software, put the mobile phone into the developer mode and turn on USB debugging. After running, the set APP will be automatically installed. After opening and agreeing to the permissions, real-time detection of fabric printing and dyeing defects on the mobile device can be achieved.
[0151] This specification also provides a multi-scale lightweight detection system for fabric printing and dyeing defects based on the improved YOLO. Please refer to the appendix Figure 8 , including:
[0152] The acquisition module 100 is used to acquire fabric printing and dyeing defect pictures;
[0153] The sample module 200 is used to label and divide the dataset of the acquired fabric printing and dyeing defect pictures according to the corresponding defect categories to obtain sample image data;
[0154] The sample enhancement module 300 is used to perform data enhancement on the sample image data;
[0155] The insertion module 400 is used to innovate the structure of the YOLOv5 model by adopting a multi-scale context aggregation module, a multi-scale context diffusion fusion pyramid network, a C3-Star module, and an Efficient Local Attention;
[0156] The training module 500 is used to train the improved YOLOv5s model using the sample image data;
[0157] The pruning module 600 is used to perform structured pruning on the trained improved YOLO model to obtain a multi-scale lightweight model;
[0158] The deployment module 700 is used to deploy the multi-scale lightweight model to a target mobile device for fabric printing and dyeing defect detection.
[0159] Beneficial effects of the first embodiment: In the first embodiment, by setting up a multi-scale context aggregation module, a multi-scale context diffusion and fusion pyramid network, and a C3-Star module, the detection accuracy and inference speed of the model are significantly improved. The Efficient Local Attention is introduced in the backbone layer to improve the focusing ability of the model on the key regions where defects are located. Further, a channel pruning method is adopted to further reduce the model complexity and accelerate the inference. Finally, through experimental verification, the proposed multi-scale lightweight model has a good detection effect and shows excellent performance in terms of detection accuracy and speed. And it is beneficial for deployment and implementation on embedded systems and mobile devices. Utilizing the existing research results, rapid detection of fabric printing and dyeing defects is realized in actual application scenarios.
[0160] The second embodiment
[0161] On the basis of the first embodiment, the S1 step is modified as follows:
[0162] S1. Collect fabric printing and dyeing defect pictures, and during the collection process, it is necessary to ensure the same production conditions, where the production conditions include printing and dyeing temperature, dye concentration, and printing and dyeing speed;
[0163] S1A. The printing and dyeing temperature is specifically: controlled within an accuracy range of ±2°C;
[0164] S1B. The dye concentration is specifically: ensuring that the deviation of the dye concentration during each printing and dyeing process does not exceed ±0.5%;
[0165] S1C. The printing and dyeing speed is specifically: stable within a fluctuation range of ±1% of the set speed;
[0166] S1D. And in terms of imaging, a polarized light illumination system is adopted, which consists of a polarized light source and a polarization filter. The polarized light emitted by the polarized light source irradiates the fabric surface, and after reflection, the polarization filter is used to filter out the reflected light irrelevant to the defect characteristics, making the defect area more clearly distinguishable in the image. At the same time, a high-resolution industrial camera is used, with a pixel count of more than 5 million, which can capture the fine defect textures on the fabric surface. The exposure time and gain parameters of the camera are dynamically adjusted according to the color and gloss of the fabric to ensure that the collected images have good contrast and brightness.
[0167] Beneficial effects of the second embodiment: In the data acquisition process of the second embodiment, the production conditions and imaging technology are optimized. The printing and dyeing temperature, dye concentration, and printing and dyeing speed are accurately controlled to reduce the interference of production condition fluctuations on defect features, enabling the model to learn more representative defect information and improving the recognition accuracy. The polarized light illumination system and high-resolution industrial camera are adopted, and the camera parameters are dynamically adjusted according to the fabric characteristics, making the defect area clearer. The obtained images have rich texture details, appropriate contrast, and brightness, providing high-quality data for image annotation and model training, helping the model learn subtle and accurate defect features, and comprehensively improving the detection accuracy.
[0168] The third embodiment
[0169] Based on the first embodiment, the S7D step is modified as follows:
[0170] S7D. Select the experimental environment configuration, set the training hyperparameters, and start training the improved model after setting the parameters. The specific training hyperparameters are set as follows:
[0171] Define the search space of the hyperparameters, including searching for the learning rate in the range of 0.001 - 0.1 and selecting the batch size between 16 and 64. Encode the hyperparameter combinations as chromosomes, and each hyperparameter corresponds to a gene of the chromosome;
[0172] The fitness function of the genetic algorithm is designed to comprehensively consider the mean average precision (mAP), recall rate (R), and inference speed (FPS) of the model on the validation set;
[0173] The specific calculation formula is: Fitness = w1 × mAP + w2 × R + w3 × FPS
[0174] Where: w1, w2, and w3 are weight coefficients, which are adjusted according to actual requirements to further improve the accuracy of the parameters;
[0175] In each generation of evolution, new hyperparameter combinations are generated through selection, crossover, and mutation operations. The selection operation adopts the roulette wheel selection method, and chromosomes are probabilistically selected according to the magnitude of the fitness value. The crossover operation adopts single-point crossover, randomly selects a crossover point, and exchanges the genes at the corresponding positions of the two chromosomes. The mutation operation randomly changes the gene values on the chromosome with a certain probability to avoid the algorithm falling into a local optimum.
[0176] According to the performance feedback on the validation set, continuously iterate and update the hyperparameter combinations until the set number of evolution generations is reached or the performance no longer improves.
[0177] Beneficial effects of Embodiment 3: Embodiment 3 uses a genetic algorithm to optimize hyperparameters. Define the search ranges for the learning rate and batch size, encode the hyperparameter combinations as chromosomes, and generate new combinations through roulette wheel selection, single-point crossover, and mutation operations. The fitness function comprehensively considers the average precision, recall rate, and inference speed, adjusts the weight coefficients according to actual requirements, and continuously iterates the hyperparameters under the performance feedback of the validation set. This method can automatically search for hyperparameters suitable for the fabric printing and dyeing defect detection task, effectively balance the detection accuracy and speed of the model, significantly improve the training effect and performance of the model, and meet the requirements of real-time and high-precision detection.
[0178] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A multi-scale lightweight detection method for fabric printing and dyeing defects based on improved YOLO, characterized in that: The following steps are involved: First, collect pictures of fabric printing and dyeing defects, annotate the pictures according to the corresponding defect categories and divide the data sets to construct sample image data; Subsequently, data enhancement is performed on the sample image data to improve the robustness and generalization ability of the model. In terms of model structure, the YOLOv5 model is structurally innovated by combining the multi-scale context aggregation module, the multi-scale context diffusion fusion pyramid network, the C3-Star module and the Efficient Local Attention to obtain an improved YOLOv5s model. The improved YOLOv5s model is trained using the sample image data. After the training is completed, the model is optimized using the structured pruning technology to further reduce the computational complexity of the model and obtain a multi-scale lightweight model. Finally, the optimized multi-scale lightweight model is deployed on the target mobile device to achieve efficient fabric printing and dyeing defect detection.
2. The multi-scale lightweight detection method and system for fabric printing and dyeing defects based on improved YOLO according to claim 1 is characterized in that: The method for collecting fabric printing and dyeing defect pictures comprises: Under the premise of ensuring the same production conditions, pictures of fabric printing and dyeing defects containing different types of defects are collected to ensure the diversity and representativeness of the data.
3. The multi-scale lightweight detection method and system for fabric printing and dyeing defects based on improved YOLO according to claim 1 is characterized in that: The method of labeling the collected fabric printing and dyeing defect images according to the corresponding defect categories and dividing the data set to obtain sample image data includes: The minimum circumscribed rectangular box in the LabelImg software is used to mark the category and location of each fabric printing and dyeing defect, where class represents the category of the defect, xmin and ymin are the coordinates of the upper left corner vertex of the minimum circumscribed rectangular box, and xmax and ymax are the coordinates of the lower right corner vertex, so as to accurately describe the type of defect and its specific location in the image.
4. The multi-scale lightweight detection method and system for fabric printing and dyeing defects based on improved YOLO according to claim 1, characterized in that: The method for performing data enhancement on sample image data comprises: One or more operations of random rotation, random scaling, random cropping, dynamic adjustment of brightness and contrast, and Mosaic enhancement are applied to the sample image data. After the operation is completed, the sample image data is associated with its corresponding defect category, so as to generate new sample image data to enrich the data set and improve the generalization ability of the model.
5. The multi-scale lightweight detection method and system for fabric printing and dyeing defects based on improved YOLO according to claim 1, characterized in that: The method of structurally innovating the YOLOv5 model by using a multi-scale context aggregation module, a multi-scale context diffusion fusion pyramid network, a C3-Star module, and Efficient Local Attention includes: Use the multi-scale context aggregation module to obtain the multi-scale context information of features; A multi-scale context diffusion fusion pyramid network is used to replace the Neck network of the YOLOv5 model, and the output of the multi-scale context aggregation module is diffused to all detection scales, so that each detection scale has rich multi-scale context information. The C3-Star module is used to replace the C3 module of the YOLOv5 model, and the abstract semantic information of the features is extracted through high-dimensional space mapping to enhance the feature extraction capability of the model. Use Efficient LocalAttention to insert it into the backbone network to enable the model to dynamically strengthen different areas, thereby focusing on the key areas where the defects are located.
6. The multi-scale lightweight detection method and system for fabric printing and dyeing defects based on improved YOLO according to claim 1, characterized in that: The method of performing structured pruning on the trained improved YOLO model to obtain a multi-scale lightweight model includes: Performing sparse regularization on the trained improved YOLO model; Set a Speed_up parameter to control the model pruning rate; The method of performing structured pruning on the trained improved YOLO model to obtain a multi-scale lightweight model includes: Sparse regularization is introduced into the trained improved YOLO model to promote the sparsification of redundant parameters in the network; Set a Speed_up parameter to control the pruning rate of the model, where Speed_up is defined as: Among them, FLOPs1 is the floating-point computing amount of the unpruned model, and FLOPs2 is the floating-point computing amount of the pruned model. By adjusting the value of Speed_up, the network slimming ratio can be flexibly controlled, thereby reducing the computational complexity of the model while maintaining its detection performance as much as possible.
7. The multi-scale lightweight detection method and system for fabric printing and dyeing defects based on improved YOLO according to claim 1, characterized in that: The following steps are also included: Improved YOLOv5 models of different depths and widths, including improved YOLOv5n, improved YOLOv5s, and improved YOLOv5m, are trained using sample image data to generate multiple trained improved YOLO models; The performance of multiple trained models is evaluated and compared. According to the model's detection accuracy, inference speed, and multi-scale lightweight effect indicators, the model with the best performance is selected as the final trained improved YOLO model.
8. The multi-scale lightweight detection method and system for fabric printing and dyeing defects based on improved YOLO according to claim 1, characterized in that: The method for comparing the performance of multiple trained improved YOLO models includes: Multiple trained improved YOLO models are evaluated and compared using commonly used performance evaluation indicators in the field of target detection: precision (P), recall (R), mean average precision (mAP), floating point operations (FLOPs), parameters, frames per second (FPS), and inference latency (Latency), in order to comprehensively measure the detection performance, computational complexity, and operating efficiency of the models.
9. The multi-scale lightweight detection method and system for fabric printing and dyeing defects based on improved YOLO according to claim 1, characterized in that: The method of deploying the multi-scale lightweight model on a target mobile device to perform fabric printing and dyeing defect detection comprises: Use the relevant files of the improved YOLO model to replace the YOLO model files in the original detection software that were pre-imported; check and verify the input and output data dimensions of the model, and adjust the input and output dimensions to match the actual needs if necessary, and modify the label file to ensure that it is consistent with the detection category of the model; Set the name and icon of the target APP, generate and export an installation package containing the new model; import and install the installation package on the target mobile device to generate the target APP; Open the target APP and run the multi-scale lightweight model to achieve efficient detection of fabric printing and dyeing defects.
10. The multi-scale lightweight detection system for fabric printing and dyeing defects based on improved YOLO according to claim 1, characterized in that: include: The sample collection module is used to obtain pictures of fabric printing and dyeing defects and provide the original data required for detection; The data annotation module annotates the defect categories of the collected fabric printing and dyeing defect images and divides the data to generate sample image data; The data enhancement module performs various data enhancement operations on sample image data to improve the robustness and generalization ability of the model; The model structure optimization module uses a multi-scale context aggregation module, a multi-scale context diffusion fusion pyramid network, a C3-Star module, and Efficient Local Attention to innovate the structure of the YOLOv5 model and obtain an improved YOLOv5s model. Training module, using enhanced sample data to train the improved YOLOv5s model; The pruning module is used to perform structured pruning on the trained improved YOLO model to obtain a multi-scale lightweight model; The deployment module is used to deploy the multi-scale lightweight model to a target mobile device to perform fabric printing and dyeing defect detection.
Citation Information
Cited By
High-reflection and high-transmittance material surface flaw detection method based on improved YOLOv11
CN120580221A
High-reflection high-transmission material surface flaw detection method based on improved YOLOv11
CN120580221B