A fabric defect detection method and device based on improved YOLOv5
By improving the YOLOv5 model, combining data enhancement, clustering and attention mechanisms, the shortcomings in accuracy and speed of existing fabric defect detection technologies are solved, and more efficient and accurate fabric defect detection is achieved, reducing the economic losses of enterprises.
Patent Information
- Application Number
- CN202310857463.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-13
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2043-07-13
AI Technical Summary
The existing fabric defect detection technology has insufficient detection accuracy and speed, resulting in inefficient and economic losses in textile enterprises in defect detection.
A fabric defect detection method based on the improved YOLOv5 is proposed. Through data preprocessing, Kmeans++ clustering, mosic-9 data enhancement, introduction of deformable convolution, attention mechanism and efficient detection head, a more efficient and accurate defect detection model is constructed.
It improves the accuracy and speed of fabric defect detection, reduces the economic losses and production costs of enterprises, and improves the quality and inspection efficiency of fabrics.
Smart Images

Figure CN116740050B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of product defect detection, and in particular to a fabric defect detection method-level device based on improved YOLOv5. Background Art
[0002] The textile industry is a traditional pillar industry with a long history in China. As the world's largest textile producer and exporter, it plays an important role in the national economy. According to statistics from the industrial site, even experienced inspection workers can only detect 70% of defects at most, and the inspection speed generally does not exceed 2m / s. Market research shows that defects on the surface of fabrics will cause their selling price to drop by about 50%, and in severe cases, the product will be unsalable, which will bring huge economic losses to textile companies.
[0003] In recent years, deep learning technology has developed rapidly and has performed well in all aspects. In terms of target detection, networks such as R-CNN, SSD, and YOLO series have emerged one after another. Among them, the YOLO series not only has the characteristics of fast detection speed and lightweight model, but also has the accuracy that is not inferior to other networks, so it has been widely used. However, for fabric defect detection, the accuracy needs to be improved urgently due to the influence of the characteristics of the defects and the existence of noise interference. In this context, it is of great research value to propose a fabric defect detection algorithm with high detection accuracy, fast speed and good practicality. Realizing the detection of fabric defects can not only improve the detection efficiency and fabric quality, but also reduce corporate losses and production costs. Summary of the invention
[0004] Purpose of the invention: The present invention relates to the technical field of product defect detection, and specifically to a fabric defect detection method and device based on improved YOLOv5, which can improve detection efficiency and fabric quality.
[0005] Technical solution: The present invention proposes a fabric defect detection method based on improved YOLOv5, which specifically includes the following steps:
[0006] (1) Collect fabric defect data sets and preprocess the data;
[0007] (2) Use the Kmeans++ clustering algorithm to cluster all defect GT frames in the fabric defect dataset and obtain K prior frames;
[0008] (3) dividing the extended fabric image dataset according to a preset ratio to obtain a training set and a validation set;
[0009] (4) Build an improved YOLOv5 model: replace the online enhancement mosic-4 in the YOLOv5 model with the mosic-9 data enhancement method; introduce the deformable convolution DCNV3 in the backbone network of feature extraction; add the convolution attention module Biformer in the neck network of feature fusion; introduce Decoupled_Detect in the detection network head;
[0010] (5) training the fabric defect training set obtained in step (3) through an improved YOLOv5 model and verifying it through a validation set to obtain an optimal weight model of the improved YOLOv5 for fabric defect detection;
[0011] (6) The preprocessed fabric defect image to be detected is input into the optimal weight model for defect detection.
[0012] Furthermore, the implementation process of step (1) is as follows:
[0013] According to the image data defect_Images and annotation information json file in the dataset, it is converted into the corresponding GT box information, and the image and label information are combined into a dataset; the data is expanded by adding Gaussian noise and Poisson noise; the image is transformed by randomly changing the brightness, contrast and sharpness, and randomly flipping the color method, so that the data can achieve a balance between each defect class.
[0014] Furthermore, the implementation process of step (2) is as follows:
[0015] A point is randomly selected from the data set as the first cluster center. According to the width and height of the defect GT box, the Kmeans++ algorithm is used for clustering to obtain K cluster centers. Then, the clustering algorithm is continued to be used based on the obtained K to obtain K prior boxes.
[0016] Furthermore, the implementation process of the mosic-9 data enhancement method in step (4) is as follows:
[0017] First, select an image to start stitching, and randomly select 8 images from the dataset; then, adjust the images to the preset size, create a new canvas, place the first image at the center of the canvas, and stitch the remaining 8 images in a clockwise direction. According to the label information, calculate the GT frame after stitching; finally, randomly select a center point, misalign the stitched image and the canvas, update the GT frame, and randomly transform the image to get the final 9-stitched image.
[0018] Furthermore, the implementation process of introducing deformable convolution DCNV3 into the backbone network of feature extraction in step (4) is as follows:
[0019] Replace the C3 modules in the fourth, sixth, and eighth layers of the backbone network with DCNV3 modules.
[0020] Furthermore, the implementation process of adding the convolutional attention module Biformer to the feature fusion neck network Neck in step (4) is as follows:
[0021] Add an attention Biformer module under the tenth layer of the detection network head. This module implicitly encodes relative position information using 3×3 deep convolution at the beginning. Then, apply a two-layer routing attention mechanism BRA module and a 2-layer MLP module with an expansion ratio of e to perform cross-position relationship modeling and position-by-position embedding.
[0022] The Efficient decoupling head is introduced into the feature fusion of the head layer, and the Decoupled_Detect module of the detection network head is replaced by the Decoupled_Detect module.
[0023] Furthermore, the information of the GT box is extracted from the annotation information json file, and each target GT box is marked as (class, x-center, y-center, w, h), where class represents the category of the fabric defect contained in the target GT box, x-center and y-center represent the x-coordinate and y-coordinate of the center point of the target GT box, respectively, and w and h represent the length and width of the target GT box.
[0024] Furthermore, the DCNV3 module is defined as follows:
[0025]
[0026] Among them, G represents the number of groups, K represents the number of sampling points, and w g represents the shared projection weight within each group, m gk represents the normalized modulation factor of the kth sampling point in the gth group, Δp gk Indicates the offset of the corresponding sampling points of the g-th group.
[0027] Based on the same inventive concept, the present invention proposes a device, including a memory and a processor, wherein:
[0028] A memory for storing computer programs that can be run on the processor;
[0029] A processor is used to execute the steps of the above-mentioned fabric defect detection method based on improved YOLOv5 when running the computer program.
[0030] Based on the same inventive concept, the present invention proposes a storage medium having a computer program stored thereon, and when the computer program is executed by at least one processor, the steps of the above-mentioned fabric defect detection method based on the improved YOLOv5 are implemented.
[0031] Beneficial effects: Compared with the prior art, the present invention has the following beneficial effects: the improved YOLOV5 model proposed in the present invention integrates deformable convolution, attention mechanism and efficient detection head; in feature extraction, the characteristics of deformable convolution expand the receptive field for defective targets, thereby obtaining better feature extraction capabilities; in feature fusion, the use of dual-route attention mechanism can enable the network to strengthen feature fusion; replacing a more efficient Decoupled detection head in the detection network, due to the use of an anchor-based detection method, not only the efficient detection efficiency is improved, but also the accuracy is gained; the present invention realizes the detection of fabric defects, abandons the past reliance on manual detection, can not only improve the detection efficiency and fabric quality, but also reduce corporate losses and production costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 It is a flowchart of the defect detection method based on the improved YOLOv5;
[0033] Figure 2 A data enhancement diagram showing the noise adding method of the present invention;
[0034] Figure 3 A data enhancement diagram of the pixel transformation method of the present invention;
[0035] Figure 4 It is the K-means++ flow chart of the present invention;
[0036] Figure 5 It is the Mosaic9 data enhancement diagram of the present invention;
[0037] Figure 6 is a structural diagram of the attention module of the present invention;
[0038] Figure 7 It is a structural diagram of the Efficient decoupling head module of the present invention;
[0039] Figure 8 It is a comparison diagram of the training result diagram of the present invention, wherein (a) is a diagram showing the running result of the fabric defect data set using the basic YOLOv5 network, and (b) is a diagram showing the training result of the fabric defect data set using the YOLOv5+DCN+BiFormer+Decoupled fabric defect detection network;
[0040] Fig. 9This is a diagram showing the effect of fabric defect detection based on improved YOLOv5 of the present invention. DETAILED DESCRIPTION
[0041] The present invention will be further described in detail below in conjunction with the accompanying drawings.
[0042] The present invention proposes a defect detection method based on improved YOLOv5, such as Figure 1 As shown, the following steps are included:
[0043] Step 1: Collect fabric defect datasets, using the Alibaba Tianchi dataset, and convert the image data defect_Images and annotation information json files in the dataset into corresponding GT box information according to the model requirements, and combine the image and label information into a dataset. For the problem of unbalanced sample numbers in the dataset, select a data enhancement method suitable for fabric defect detection to expand it. The data enhancement method includes adding Gaussian noise and Poisson noise, such as Figure 2 As shown, the image is transformed pixel by pixel, including random changes in brightness, contrast and sharpness, and random color flipping methods, such as Figure 3 As shown, the data is balanced between each defect class.
[0044] For the information of the GT box extracted from the annotation information json file, each target GT box is marked as (class, x-center, y-center, w, h), class represents the category of fabric defects contained in the target GT box, x-center and y-center represent the x-coordinate and y-coordinate of the center point of the target GT box respectively, w and h represent the length and width of the target GT box; the appropriate data enhancement method is selected by inputting sample data into the basic YOLOv5s network for training and verification, and the appropriate data enhancement method is selected according to the experimental training accuracy.
[0045] In order to prove the effectiveness of data enhancement equalization, the experimental results are shown in Table 1. Through experiments, it can be found that data enhancement equalization is much better than unenhanced equalization, with the accuracy increased by 4.7 points and mAP@0.5 increased by 9.6 percentage points.
[0046] Table 1 Comparison of experimental results with and without data enhancement and equalization
[0047]
[0048] Step 2: Use the Kmeans++ clustering algorithm to cluster all defect GT boxes in the fabric dataset and obtain K prior boxes.
[0049] First, randomly select a point from the data set as the first cluster center. According to the width and height of the defective GT box, use the Kmeans++ algorithm to cluster and obtain K cluster centers. Then, based on the obtained K, continue to use the clustering algorithm to obtain K prior boxes. The specific process is as follows: Figure 4 As shown in the figure, the image size of the fabric defect dataset used is 2446×1000 pixels. The defect GT box is clustered by the Kmeans++ algorithm to obtain 7 prior box sizes (52,2393), (977,591), (62,576), (13,125), (26,291), (28,1280), (573,293).
[0050] Step 3: Divide the extended fabric image dataset according to a preset ratio to obtain a training set and a validation set.
[0051] Step 4: Build an improved YOLOv5 model, replace the online enhanced mosic-4 in the YOLOv5 model with the mosic-9 data enhancement method; introduce deformable convolution (DCNV3) in the backbone network of feature extraction; add a convolutional attention module (Biformer) to the neck network of feature fusion; introduce Decoupled_Detect in the detection network head.
[0052] The mosic-4 data enhancement method is an online data enhancement method built into YOLOv5. It reads four pictures, flips, scales, changes the color gamut, etc., and arranges them in four directions before combining the pictures and the GT frames. The same mosic-9 data enhancement method first selects an image to start stitching, and randomly extracts 8 images from the data set; then, the pictures are uniformly adjusted to the preset size, a new canvas is created, the first image is placed at the center of the canvas, and the remaining 8 images are stitched together in a clockwise direction. According to the label information, the spliced GT frame is calculated; finally, a center point is randomly selected, the stitched image and the canvas are misaligned, the GT frame is updated, and the image is randomly transformed to obtain the final 9-stitched image. After mosic-9 data enhancement, the following is shown: Figure 5 shown.
[0053] Introducing deformable convolution (DCNV3) in the backbone layer of the feature extraction backbone network is to replace the second C3 module, the third C3 module, and the fourth C3 module in the backbone network with the C3_DCNV3 module.
[0054] DCNV3 is based on DCNV2 and combines the related ideas of multi-head attention (MHSA). It has made three key improvements. Shared convolution weights, project the sampling points with the same weight, and then weight the projected feature vector with a position-aware learnable coefficient; introduce a multi-group mechanism to group DCNV3, perform different offset sampling, sampling vector projection, and factor modulation in each group, which enhances the expression ability of the DCNV3 operator; normalize the modulation scalar, and change the modulation scalar from sigmoid processing to softmax normalization, making the entire training process more stable.
[0055] The definition of DCNV3 is as follows:
[0056]
[0057] Among them, G represents the number of groups, K represents the number of sampling points, and w g represents the shared projection weight within each group, m gk represents the normalized modulation factor of the kth sampling point in the gth group, Δp gk Represents the offset of the corresponding sampling point of the g-th group. In this way, the DCNV3 operator not only makes up for the shortcomings of traditional convolution in long-range dependence and adaptive spatial aggregation, but also makes the variable convolution operator more suitable for large visual models. It can be said that a good balance is achieved between accuracy and computational complexity.
[0058] Adding a convolutional attention module (Biformer) to the Neck layer of the feature fusion neck network is to add an attention Biformer module to the tenth layer of the detection network head. This module uses a 3×3 deep convolution to implicitly encode relative position information at the beginning; then the Bi-Level Routing Attention (BRA) module and the 2-layer Multi-Layer Perceptron (MLP) module with an expansion rate of e are applied in sequence, which are used for cross-position relationship modeling and each position embedding, respectively, such as Figure 6 shown.
[0059] BRA is the key module of Biformer, and its core idea is to filter out the most irrelevant key-value pairs at the coarse-grained region level. This is achieved by first constructing a region-level association graph and then pruning it so that each node only retains top-k connections. Therefore, each region only needs to pay attention to the top-k routing regions. After determining the participating regions, the next step is to apply token-to-token attention, which is non-trivial because the key-value pairs are now assumed to be spatially dispersed.
[0060] The Efficient decoupling head is introduced into the feature fusion of the head layer, and the Decoupled_Detect module of the detection network head is replaced by the Decoupled_Detect module.
[0061] The detection head of YOLOv5 is a coupled head that shares parameters between the classification and positioning branches, while the Efficientdecoupled head uses a hybrid channel strategy to build a more efficient decoupled head. The detection head in the Efficientdecoupled head decouples the two branches and introduces an additional 3×3 convolutional layer in each branch to improve performance. The width of the head is scaled by the width multipliers of the Backbone and Neck, such as Figure 7 shown.
[0062] Step 5: Train the improved YOLOv5 network. The fabric defect training set obtained in step 3 is trained by the improved YOLOv5 model and verified by the validation set to obtain the optimal weight model of the improved YOLOv5 for fabric defect detection.
[0063] In order to demonstrate the performance of the proposed method, the processed fabric defect data set is trained using the YOLOv5+DCN+BiFormer+Decoupled fabric defect detection network. The experimental results are shown in Table 2. The data results include precision (Precision, P), recall (Recall, R), average precision when IoU is 0.5 (mAP@0.5), and average precision with a step size of 0.05 between IoU 0.5 and 0.95 (mAP@.5:.95).
[0064] Table 2 Experimental results statistics
[0065]
[0066] It can be seen from Table 2 that for fabric defect detection, the accuracy of the method of the present invention is increased by 0.3 percentage points when mosic-4 is replaced by mosic-9, and mAP@0.5 is reduced by 1.2 percentage points. The C3 module is replaced by C3_DCNV3, and the accuracy is increased by 2 percentage points, and mAP@0.5 is increased by 0.3 percentage points. The attention module Biformer is introduced, and the accuracy is increased by 2.3 percentage points, and mAP@0.5 is increased by 0.8 percentage points. Finally, the detection head Decoupled is replaced, and the accuracy is increased by 6.1 percentage points, and mAP@0.5 is increased by 3 percentage points. The above results verify the superiority of the method in fabric defect detection. Figure 8 (a) is the result of running the basic YOLOv5 network on the fabric defect dataset. Figure 8(b) is the training result of the fabric defect detection network using YOLOv5+DCN+BiFormer+Decoupled on the fabric defect dataset; it can be seen that there is a significant increase in accuracy.
[0067] Step 6: Input the preprocessed fabric defect image to be detected into the optimal weight model for defect detection. The fabric defect detection effect of improved YOLOv5 is shown in the figure below: Fig. 9 shown.
[0068] The present invention also proposes a device, comprising a memory and a processor, wherein: the memory is used to store a computer program that can be run on the processor; the processor is used to execute the steps of the above-mentioned fabric defect detection method based on the improved YOLOv5 when running the computer program.
[0069] The present invention also proposes a storage medium, on which a computer program is stored, and when the computer program is executed by at least one processor, the steps of the above-mentioned fabric defect detection method based on improved YOLOv5 are implemented.
[0070] So far, the technical solutions of the present invention have been described in conjunction with the specific experimental processes shown in the accompanying drawings, but the protection scope of the present invention is not limited to these specific implementations. Without departing from the principles of the present invention, those skilled in the art may make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.
Claims
1. A fabric defect detection method based on improved YOLOv5, characterized in that: The following steps are involved: (1) Collect fabric defect data sets and preprocess the data; (2) Use the Kmeans++ clustering algorithm to cluster all defect GT frames in the fabric defect dataset and obtain K prior frames; (3) dividing the extended fabric image dataset according to a preset ratio to obtain a training set and a validation set; (4) Build an improved YOLOv5 model: replace the online enhancement mosic-4 in the YOLOv5 model with the mosic-9 data enhancement method; introduce the deformable convolution DCNV3 in the backbone network of feature extraction; add the convolution attention module Biformer in the neck network of feature fusion; introduce Decoupled_Detect in the detection network head; (5) training the fabric defect training set obtained in step (3) through the improved YOLOv5 model and verifying it through the validation set to obtain the optimal weight model of the improved YOLOv5 for fabric defect detection; (6) inputting the preprocessed fabric defect image to be detected into the optimal weight model for defect detection; The implementation process of the mosic-9 data enhancement method in step (4) is as follows: First, select an image to start stitching, and randomly select 8 images from the data set; then, adjust the images to the preset size, create a new canvas, place the first image at the center of the canvas, stitch the remaining 8 images in a clockwise direction, and calculate the GT frame after stitching according to the label information; finally, randomly select a center point, displace the stitched image and the canvas, update the GT frame, and randomly transform the image to get the final 9-stitched image; The implementation process of introducing deformable convolution DCNV3 into the backbone network of feature extraction in step (4) is as follows: Replace the C3 modules in the fourth, sixth, and eighth layers of the backbone network with DCNV3 modules; The implementation process of adding the convolutional attention module Biformer to the feature fusion neck network Neck in step (4) is as follows: Add an attention Biformer module under the tenth layer of the detection network head. This module implicitly encodes relative position information using 3×3 deep convolution at the beginning. Then, apply a two-layer routing attention mechanism BRA module and a 2-layer MLP module with an expansion ratio of e to perform cross-position relationship modeling and position-by-position embedding. The Efficient decoupling head is introduced into the feature fusion of the head layer, and the Decoupled_Detect module of the detection network head is replaced by the Decoupled_Detect module.
2. The fabric defect detection method based on improved YOLOv5 according to claim 1, characterized in that: The implementation process of step (1) is as follows: According to the image data defect_Images and annotation information json file in the dataset, it is converted into the corresponding GT box information, and the image and label information are combined into a dataset; the data is expanded by adding Gaussian noise and Poisson noise; the image is transformed by randomly changing the brightness, contrast and sharpness, and randomly flipping the color method, so that the data can achieve a balance between each defect class.
3. The fabric defect detection method based on improved YOLOv5 according to claim 1, characterized in that: The implementation process of step (2) is as follows: A point is randomly selected from the data set as the first cluster center. According to the width and height of the defect GT box, the Kmeans++ algorithm is used for clustering to obtain K cluster centers. Then, the clustering algorithm is continued to be used based on the obtained K to obtain K prior boxes.
4. The fabric defect detection method based on improved YOLOv5 according to claim 2, characterized in that: For the GT box information extracted from the annotation information json file, each target GT box is marked as (class, x-center, y-center, w, h), where class represents the category of fabric defects contained in the target GT box, x-center and y-center represent the x-coordinate and y-coordinate of the center point of the target GT box respectively, and w and h represent the length and width of the target GT box.
5. The fabric defect detection method based on improved YOLOv5 according to claim 1, characterized in that: The DCNV3 module is defined as follows: Among them, G represents the number of groups, K represents the number of sampling points, and w g represents the shared projection weight within each group, m gk represents the normalized modulation factor of the kth sampling point in the gth group, Δp gk Indicates the offset of the corresponding sampling points of the g-th group.
6. A device, characterized in that: comprising a memory and a processor, wherein: A memory for storing computer programs that can be run on the processor; A processor, configured to execute the steps of the fabric defect detection method based on the improved YOLOv5 as claimed in any one of claims 1 to 5 when running the computer program.
7. A storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by at least one processor, the steps of the fabric defect detection method based on the improved YOLOv5 as described in any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Assembly body change detection method, device and medium based on attention mechanism
CA3121440A1
Fabric defect detection method based on DR-RSBU-YOLOv5
CN115187544A