A cloth defect detection method based on a lightweight cascade network
By using data enhancement technology and YOLOv5s-GSBiFPN lightweight detection algorithm in cloth defect detection, combined with the dual attention CSPGhostSE structure and cascading network architecture, the problems of insufficient data set samples, large model parameters and high error detection rate are solved, and efficient cloth defect detection is achieved.
Patent Information
- Application Number
- CN202210887548.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-26
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-07-26
AI Technical Summary
In the detection of fabric defects, the existing technology has problems such as insufficient data set samples, large number of detection model parameters, difficulty in deploying on low-performance equipment, and high false detection rates.
A cloth defect detection method based on lightweight cascade network is proposed, including data augmentation technology to generate new samples through Poisson fusion, designing the YOLOv5s-GSBiFPN lightweight detection algorithm, adopting a dual-way attention CSPGhostSE structure and a weighted bidirectional pyramid BiFPN structure, combining detection and classification cascade network architecture to reduce the error detection rate.
The data set is effectively expanded, the detection capability of small targets and difficult samples is improved, the number and size of model parameters is reduced, the deployment is realized on the low-performance device side, and the error detection rate is significantly reduced.
Smart Images

Figure CN115205274B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of defect detection in computer vision, and particularly to a method for detecting cloth defects based on a lightweight cascade network. Background Art
[0002] A cloth defect detection system based on computer vision usually monitors the quality of cloth by taking pictures of cloth images. Traditional cloth defect detection often uses image processing methods to extract features such as textures and edges in the cloth images, and then uses appropriate algorithms to process these features to effectively identify cloth defects. Traditional methods often require manual design of feature extractors. This method is very effective in specific scenarios, but has weak generalization ability and cannot adapt to changes in scenarios. With the development of deep learning, defect detection methods based on convolutional neural networks (CNNs) have demonstrated their powerful feature extraction capabilities, and can automatically learn the feature representations of defects from a large number of cloth image samples, far exceeding traditional methods in terms of detection accuracy, real-time performance, and wide adaptability. Deep learning-based methods have more powerful performance, but such methods require powerful computing power and storage to build an inference environment for neural network models. In the actual implementation and deployment of algorithms, restricted by current hardware conditions, their advantages cannot be well demonstrated. At the same time, the detection rate for small defect targets and difficult samples is relatively low, and misdetection is relatively serious for defect samples with unclear defect features. Summary of the Invention
[0003] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a cloth defect detection method based on a lightweight cascade network. First, in view of the situation that when constructing a cloth defect dataset, the number of cloth defect samples of certain specific categories is small and difficult to obtain, a data augmentation method is proposed. The defect area is intercepted from the existing sample images containing defects, and then randomly pasted back to the negative sample images without defects through Poisson fusion to form new cloth defect samples, expanding the dataset. Second, for the current cloth defect detection model, aiming at the problems of weak detection ability for small targets and difficult samples, as well as the large number of network parameters and large model size, which are difficult to be deployed on devices with low performance, the YOLOv5s-GSBiFPN lightweight cloth defect detection algorithm is proposed. A new dual-path attention CSPGhostSE structure is designed as the core feature extraction module of the backbone network, and together with the SPPF module, a new feature extraction network is formed. In the feature fusion stage, a three-layer stacked weighted bidirectional pyramid BiFPN structure is used, and weighted discrimination is made according to the different contributions of features with different resolutions to the final output of the network. After fully fusing multi-scale features, they are input into 3 detection heads to detect defect targets of large, medium and small scales respectively. The detection ability for small targets and difficult samples is significantly improved, the number of parameters and the model size are greatly reduced, and it is easy to be deployed on devices with low performance. Finally, aiming at the problem of easy false detection of samples with unclear defect features, a cascade network architecture of detection + classification is proposed. A secondary classification network is added after the detection network to further perform binary classification judgment on whether the predicted boxes output by the detection model are defects, filtering the detection results and greatly reducing the false detection rate.
[0004] A cloth defect detection method based on a lightweight cascade network, the technical solution adopted includes the following steps:
[0005] Step 1: Use an industrial camera and an LED lighting device to construct a well-lit and stable imaging shooting environment, and collect cloth images; in the shooting environment, the front lighting method is adopted, and the camera and the light source are placed on the same side and kept parallel to the cloth to be inspected.
[0006] Step 2: Collect defective cloth images containing different defects, and perform data augmentation on the cloth images to expand the number of dataset samples and balance the number of defect samples of different categories; the data augmentation method is to intercept the defect area from the existing sample images containing defects, and then randomly paste it back to the negative sample images without defects through Poisson fusion to form new cloth defect samples, expanding the dataset.
[0007] Furthermore, a segmented annotation strategy is adopted to annotate the defective areas of the cloth images, making the annotation rectangular boxes fit the defective areas. After the annotation of the samples containing defects is completed, cloth images without defects are added as negative samples according to a 1:1 ratio. Then, all samples are divided into a training set and a test set at an 8:2 ratio to construct a cloth defect dataset.
[0008] Step 3: Construct a lightweight cloth defect detection model of YOLOv5s-GSBiFPN. A lightweight GhostConv phantom convolution structure is introduced into the backbone network, and the SE attention module is embedded into the CSP core feature extraction network to form a new dual-path attention CSPGhostSE structure. Deep features are extracted by stacking multiple layers of CSPGhostSE structures, and then the features are input into the SPPF module to be mapped to a fixed dimension. The GhostConv phantom convolution, the CSPGhostSE structure, and the finally added SPPF structure together form a new feature extraction network. In the feature fusion stage, a three-layer stacked weighted bidirectional pyramid BiFPN structure is used. Weighted discrimination is made according to the different contributions of features with different resolutions to the final output of the network, fully fusing multi-scale features and inputting them into 3 detection heads to detect defective targets of large, medium, and small scales respectively.
[0009] In the dual-path attention CSPGhostSE structure, the ordinary convolution is replaced by the GhostConv phantom convolution in the lightweight network GhostNet, and the CSP structure idea is borrowed to construct the CSPGhost structure. While reducing the computational amount and the number of parameters of the network, it increases the gradient calculation path, avoids repeated calculation of gradient information during the backpropagation of the network, and significantly improves the learning ability of the network.
[0010] Furthermore, aiming at the problem that the decrease in the number of parameters caused by replacing the ordinary convolution with a lightweight network will have a negative impact on the accuracy of the detection network, the SE attention module is embedded into the GhosBottleneck core structure, enabling the network to better focus its attention on important features, improving the ability to capture features of difficult samples, and thus enhancing the overall performance of the model to make up for the decrease in detection accuracy caused by the reduction in the number of parameters.
[0011] The dual-path attention CSPGhostSE structure contains two branches, and each branch is composed of a GhostConv phantom convolution, an SE attention module, and a CBS module. Branch 1 contains a CBS module, a stacked multi-layer GhostBottleneck structure with a stride of 1, and a GhostBottleneck structure with a stride of 2 connected in sequence; Branch 2 contains a CBS module and a GhostBottleneck structure with a stride of 2 connected in sequence. The output features of the two branches are merged through a Concat operation, and after passing through a batch normalization layer and an activation function, they are input into the lower-level network.
[0012] The CBS module contains a Conv2d convolutional layer with a kernel size of 1*1, a batch normalization BN layer, and an activation function SiLU layer. The input features are transformed through a Conv2d convolutional layer with a kernel size of 1*1, and the number of convolutional kernels is set to 1 / 2 of the input feature map to adjust the number of output feature channels.
[0013] The GhostBottleneck structure with a stride of 1 contains two layers. The upper layer is two GhostConv phantom convolutions and an SE attention module connected in sequence; the lower layer is a short connection structure that directly maps the original input features; the features of the two layers are fused by addition. The GhostBottleneck structure with a stride of 2 adds a depthwise convolution DWConv layer with a stride of 2 in the middle of the two GhostConv phantom convolutions in the upper layer, and the SE attention module is located after the DWConv layer, which has a downsampling effect of 1 / 2.
[0014] Among them, the GhostConv phantom convolution structure contains a convolutional layer and a depthwise convolutional layer. The input features first pass through a convolutional layer with a kernel size of 3*3, and then are divided into two paths. One path uses depthwise convolution to perform a linear transformation on the feature map generated by the basic convolution to generate the other half of the Ghost convolution; the other path is a short connection structure that outputs its own features; then the feature maps obtained from the two parts are concatenated together by channel through a Concat operation.
[0015] Furthermore, in the Neck stage of the cloth defect detection network, the BiFPN weighted bidirectional pyramid structure is used to replace the original PAN structure. In multi-scale feature fusion, weighted discrimination is made according to the different contributions of features with different resolutions to the final output of the network, which promotes the efficient fusion of multi-scale features.
[0016] A multi-scale fast fusion method based on weights is proposed in the BiFPN, and the definition formula is as follows:
[0017]
[0018] Where \(o\) is the output of the fused feature, \(w\) i is the learnable weight, \(\epsilon\) is a very small value to avoid division by zero, and \(I\) i is the \(i\)-th input feature.
[0019] First, the weights are normalized to ensure that each weight value is between 0 and 1.
[0020] For the fusion of the multi-scale features, depthwise separable convolution is used for fusion. After convolution, a BN layer and a non-linear activation layer are connected in sequence. The intermediate layer feature \(P\) is defined by the formula:
[0021]
[0022] Where is the intermediate feature of the current layer, is the input feature of the current layer, is the input feature of the next layer, Conv is the convolution operation, and Resize is the sampling operation.
[0023] The feature of the next layer goes through the Resize operation and is then multiplied by the weight \(w\) 2 and fused with the feature of the current layer. The output formula of the current layer is defined as:
[0024]
[0025] Where is the output feature of the current layer, is the output feature of the previous layer.
[0026] The output feature of the previous layer goes through the Resize operation and is multiplied by the corresponding weight \(w'\) 3 and fused with the feature of the current layer.
[0027] Step 4: Optimize the loss function. The focal loss function Focal Loss is used to calculate the classification loss, which is defined as:
[0028]
[0029] Where \(y\) is the true sample label, \(y'\) is the output after the activation function; \(\gamma\) is the influencing factor, which is used to reduce the weight of easy-to-classify samples and increase the weights of difficult samples and misclassified samples; \(\alpha\) is the balancing factor, which is used to balance the weights of the two classes when the number of positive and negative samples is unbalanced; log is the logarithmic function; \(\gamma = 2\), \(\alpha = 0.25\) are the optimal values.
[0030] Step 5: Build a binary classification data set and construct a secondary classification network; the binary classification data sets are all from the cloth defect data set, which are obtained by cropping. The defect categories are not subdivided, and only two categories are included: defective samples and non-defective samples, which simplifies the classification problem.
[0031] The defective samples are expanded by 2 pixels in width and height according to the original annotation box information. When the width and height exceed the original image, no expansion is performed and the original image is cropped. The non-defective samples are obtained by randomly cropping the cloth image without defects based on the average size of the existing defective area.
[0032] The secondary classification network adopts the ResNet18 classification model structure, the input image size is 56*56, and the number of channels of the final output feature map is 512.
[0033] Step 6: Train the improved lightweight cloth defect detection model and the binary classification model; set the training hyperparameters: batch size, anchor box size, training iteration rounds, initial learning rate, number of target categories, and probability of random mosaic data enhancement, etc.
[0034] When the training iteration reaches the point where the model loss curve is close to 0 and tends to be flat, the training is stopped to obtain the optimal model.
[0035] Step 7: Construct a cascade network architecture for detection and classification; the cascade network architecture for detection and classification consists of a primary detection network and a secondary classification network; the primary detection network is responsible for locating and classifying defects in the cloth image and outputting a prediction box, and the secondary classification network further performs a binary classification judgment on whether the prediction box area output by the detection network contains defects, filters false detections, and only outputs the prediction results determined by the classification network to contain defects;
[0036] Step 8: Input the cloth image to be tested into the cascade network model to perform defect detection, and output the cloth defect detection results and defect target location information.
[0037] Preferably, the cloth defect images in step 2 include: polyester, cotton, satin and grey cloth. There are three types of defects in total: repeated / broken warp, broken weft and blurred weft; each image contains one or more types of defects.
[0038] Preferably, the data enhancement method in step 2 is to manually enhance a certain defect category with a small number of defects. The defect area of a specific category is cut from the original image, and then randomly pasted back to the negative sample image without defects through Poisson fusion to form a new cloth defect image and expand the data set.
[0039] Preferably, the GhostConv phantom convolution in step 3 includes two parts: the primary convolution part (PrimaryConvolution) and the cheap operation of linear transformation (Cheap Operation).
[0040] The Primary Convolution consists of ordinary components, uses a small number of convolutions, reduces the number of convolution kernels to half of the original, reduces the number of channels of the feature map by half, and correspondingly reduces the number of parameters by half.
[0041] The Cheap Operation consists of a convolution structure with a fixed filter size, is implemented using depthwise convolution, and performs a linear transformation on the feature map generated by the primary convolution to generate the other half of the Ghost feature map. Finally, the feature maps obtained from the two parts are concatenated together according to the channels.
[0042] Preferably, in step 3, the position where the SE module is added is after the two GhostConv phantom convolutions in the GhostBottleneck structure with a stride of 1, and after the DWConv in the GhostBottleneck structure with a stride of 2, giving greater weights to important channels.
[0043] Preferably, the BiFPN Layer in step 3 is stacked three times to achieve the fusion of higher-level information. Depthwise separable convolutions are used for fusion in BiFPN, and a BN layer and a non-linear activation layer are attached after the convolution.
[0044] Preferably, in step 6, the hyperparameter settings are as follows: the number of training iterations epoch is 300, the number of samples loaded in one training batchsize is 64, and the initial learning rate lr0 is 0.01. The anchor settings are [36, 20, 10, 114, 119, 13], [26, 154, 18, 422, 159, 47], and [401, 30, 35, 471, 431, 72] respectively.
[0045] Preferably, the model performance evaluation metrics include mAP, Top1 accuracy (Top-1), computational complexity (GFLOPs), number of parameters (Parameters), model size (Size), model inference speed (Speed), and model detection speed (FPS).
[0046] The beneficial effects of the present invention:
[0047] In view of the situation that when constructing a cloth defect dataset, the number of cloth defects in certain specific categories is small and difficult to obtain, the present invention proposes a data augmentation method. The defect regions of specific categories are intercepted from the original image, and then randomly pasted back into the negative sample images without defects through Poisson fusion to form new cloth defect samples, which are added to the dataset.
[0048] Secondly, aiming at the current cloth defect detection model, which has weak detection ability for small targets and difficult samples, and the problem that the number of network parameters is large and the model is large, making it difficult to deploy on devices with low performance, the YOLOv5s-GSBiFPN lightweight cloth defect detection algorithm is proposed. A lightweight GhostConv phantom convolution structure is introduced into the backbone network, and the SE attention module is embedded into the CSP core feature extraction network, and a new dual-path attention CSPGhostSE structure is designed. Then, deep features are extracted by stacking multiple layers of CSPGhostSE structures, and the features are input into the SPPF module to be mapped to a fixed dimension. In the feature fusion stage, a three-layer stacked weighted bidirectional pyramid BiFPN structure is used to fully fuse multi-scale features and input them into 3 detection heads to detect defect targets of large, medium and small scales respectively. It significantly improves the detection ability of small targets and difficult samples, greatly reduces the number of parameters and the model size, and is easy to deploy on devices with lower performance.
[0049] Finally, aiming at the problem that samples with unclear defect features are prone to false detection, a cascaded network architecture of detection + classification is proposed. A secondary classification network is added after the detection network to further perform binary classification judgment on whether the prediction boxes output by the detection model are defects, and filter the detection results, greatly reducing the false detection rate. Description of the Drawings
[0050] Figure 1 It is a schematic diagram of the overall process of the lightweight cascaded network cloth defect detection method proposed in the embodiment of the present invention.
[0051] Figure 2 It is a network structure diagram of the lightweight cloth defect detection YOLOv5s-GSBiFPN in the embodiment of the present invention.
[0052] Figure 3 It is a core structure diagram of the dual-path attention CSPGhostSE designed in the embodiment of the present invention.
[0053] Figure 4 It is a graph of the change of training loss of the YOLOv5s-GSBiFPN and the original YOLOv5s algorithm proposed in the embodiment of the present invention.
[0054] Figure 5This is the curve graph of the change of the mAP index during the training process of YOLOv5s-GSBiFPN and the original YOLOv5s algorithm proposed in the embodiments of the present invention.
[0055] Figure 6 This is the detection result graph of the present invention in the defective cloth image. Detailed implementation manners
[0056] The accompanying drawings are only for illustrative purposes and should not be construed as a limitation of this patent;
[0057] For those skilled in the art, it is understandable that some well-known contents in the accompanying drawings may be omitted.
[0058] To make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings:
[0059] The present invention applies the latest research results of object detection to the cloth defect detection task, and designs a lightweight detection algorithm suitable for multi-platform and multi-scenario deployment according to the diversified deployment requirements of enterprises.
[0060] As an implementable manner, as Figure 1 shown, a cloth defect detection method based on a lightweight cascade network includes the following steps:
[0061] Use an industrial camera and an LED lighting device to construct a well-illuminated and stable imaging shooting environment to collect cloth images; in the shooting environment, a front lighting method is adopted, and the camera and the light source are placed on the same side and kept parallel to the cloth to be inspected.
[0062] Collect defective cloth images containing different defects, and perform data augmentation on the cloth images to expand the number of samples in the data set and balance the number of defect samples of different categories. For the data augmentation, the defective area of a specific category is intercepted from the original image, and then randomly pasted back to the negative sample image without defects through Poisson fusion to form a new cloth defect sample and expand the data set.
[0063] Furthermore, label the cloth images containing defects. The sizes and dimensions of the cloth defect areas vary greatly and mostly appear in the form of extreme aspect ratios, showing a slender state in the image. In addition, there are also defects in the form of non-vertical and non-horizontal inclined states. Adopt a segmented labeling strategy to make the labeled rectangular frame fit the defect area to the greatest extent, reduce the interference of background information, and highlight the defect features.
[0064] After completing the labeling of the defective images, add cloth images without defects as negative samples in a 1:1 ratio; then divide all samples into a training set and a test set in an 8:2 ratio to construct a cloth defect data set;
[0065] Construct a lightweight fabric defect detection model YOLOv5s-GSBiFPN; introduce a lightweight GhostConv phantom convolution structure into the backbone network, and embed the SE attention module into the CSP core feature extraction network to form a new dual-path attention CSPGhostSE structure; stack multiple CSPGhostSE structures to extract deep features, and then input the features into the SPPF module to map to a fixed dimension. The GhostConv phantom convolution, CSPGhostSE structure, and the finally added SPPF structure together form a new feature extraction network. In the feature fusion stage, use a three-layer stacked weighted bidirectional pyramid BiFPN structure to fully fuse multi-scale features and input them into 3 detection heads to detect defect targets of large, medium, and small scales respectively.
[0066] Optimize the loss function, and use the focal loss function Focal Loss to calculate the classification loss, which is defined as:
[0067]
[0068] In the formula, y is the true sample label, and y′ is the output after the activation function; γ is the influencing factor, which is used to reduce the weight of easy-to-classify samples and increase the weights of difficult samples and misclassified samples; α is the balancing factor, which is used to balance the weights of the two categories when the number of positive and negative samples is unbalanced; log is the logarithmic function; γ = 2 and α = 0.25 are the optimal values.
[0069] Construct a binary classification dataset, construct a secondary classification network based on ResNet18, the input image size is 56*56, and the number of channels of the finally output feature map is 512.
[0070] The training data of the secondary classifier all come from the fabric defect dataset, which is obtained by cropping. The defect categories are not subdivided, and only defective samples and non-defective samples are included. The defect categories are no longer considered in detail, simplifying the classification problem.
[0071] The defective samples are obtained by cropping the original image according to the original annotation box information, expanding 2 pixels in width and height, and not expanding when exceeding the width and height of the original image; the non-defective samples are randomly cropped from the fabric image without defects according to the average size of the defect area.
[0072] Train the improved lightweight fabric defect detection model and the binary classification model; stop training when the training iteration reaches that the model loss curve is close to 0 and tends to be flat to obtain the optimal model;
[0073] Furthermore, a secondary classifier is added after the detection network to further classify and filter the prediction boxes output by the detection network, and only the prediction boxes determined by the classifier to be of the defect category are output. In cooperation with the detection network, the model misdetection is minimized to the greatest extent.
[0074] Input the image of the cloth to be measured into the cascaded network model for defect detection, and output the detection results of cloth defects and the location information of defect targets.
[0075] As an implementable manner, as Figure 2 shown, the embodiment of the present invention provides a lightweight cloth defect detection model YOLOv5s-GSBiFPN improved based on the YOLOv5s algorithm. This model aims at the problems of the current cloth defect detection model, such as the weak detection ability for small targets and difficult samples, and the large number of network parameters and large model size, which makes it difficult to deploy on platforms with low performance. The YOLOv5s algorithm is improved from two aspects: lightweight network structure and improved detection ability for small targets and difficult samples.
[0076] In the feature extraction stage of the backbone network, the lightweight network GhostConv phantom convolution is used to replace the original ordinary convolution, and the CSP structure idea in CSPNet is borrowed to construct the CSPGhost structure. While reducing the network calculation amount and parameter amount, the gradient calculation path is increased, avoiding repeated calculation of gradient information during backpropagation of the network, and improving the learning ability of the network.
[0077] Furthermore, aiming at the problem that the reduction of parameter amount will have a negative impact on the detection network accuracy when the lightweight network replaces the ordinary convolution, the SE attention module is embedded into the core structure of CSPGhost, enabling the network to better focus on important features, enhancing the ability to capture difficult sample features, thereby improving the overall performance of the model and making up for the decline in detection accuracy caused by the reduction of parameter amount.
[0078] As an implementable manner, the dual-path attention CSPGhostSE structure, as Figure 3 shown, includes two branches, and each branch is composed of GhostConv phantom convolution, SE attention module and CBS module; Branch 1 includes a CBS module connected in sequence, a stacked multi-layer GhostBottleneck structure with a stride of 1 and a GhostBottleneck structure with a stride of 2; Branch 2 includes a CBS module connected in sequence and a GhostBottleneck structure with a stride of 2; The output features of the two branches are merged through the Concat operation, and after passing through the batch normalization layer and activation function, they are input into the lower-layer network.
[0079] For the CSPGhostSE structure, the input features are first divided by the CBS module. The CBS module includes a Conv2d layer with a kernel size of 1*1, a batch normalization BN layer, and an activation function SiLU layer, where the number of convolution kernels in the Conv2d layer is half of the input feature channels. Instead of dividing the input features into two parts by channels as in the original CSPNet, the input features are transformed using a 1*1Conv2d convolution in two CBS modules. By controlling the number of convolution kernels, the output channel number is adjusted, achieving the same effect as halving the channels and further improving the feature reuse.
[0080] The position where the SE attention module is added is in the GhostBottleneck structure with a stride of 1. The GhostBottleneck structure with a stride of 1 also includes two layers. The upper layer is two consecutive GhostConv phantom convolutions and the SE attention module connected in sequence, and the lower layer is a short connection structure that outputs its own features. The two fuse the features of the two layers by addition. The GhostBottleneck structure with a stride of 2 adds a depthwise convolution DWConv layer with a stride of 2 in the middle of the two GhostConv phantom convolutions in the upper layer. The SE attention module is located after the DWConv layer and has a downsampling effect of 1 / 2.
[0081] For the GhostConv phantom convolution structure, first, ordinary convolution is used to generate the base features, also called the eigenfeature map. Then, linear transformation operations are performed on the base features obtained by convolution channel by channel to generate ghost features. The linear transformation in the GhostConv phantom convolution structure does not actually use linear operations such as translation, rotation, affine transformation, and wavelet transformation, but is implemented by convolution. The convolution structure itself can cover many linear operations, such as smoothing and blurring. Special linear operations require pre-setting parameter thresholds, etc., while the convolution structure can automatically adjust the weight parameters during network training, with better effects. By using depthwise convolution DWConv, linear transformation is performed on the base feature map channel by channel to generate the other half of the ghost feature map. Finally, the base features and ghost features are concatenated. The involved formulas are:
[0082] Y′=X*f′
[0083] In the formula, X is the input feature, f′∈R c×k×k×m is the filter, and the bias term is omitted here for simplicity. Y′ is the base feature generated by ordinary convolution; y′ i is the i-th original feature map in Y′; Φ i,j is the j-th linear operation used to generate the j-th phantom feature map y ij .
[0084] Furthermore, the BiFPN structure is composed of three stacked BiFPN Layers. BiFPN can be stacked as a network structure block to achieve the fusion of higher-level information.
[0085] The BiFPN deletes the nodes with only one input on the basis of the PAN structure. The nodes with only one input do not produce feature fusion and contribute little to the output of fusing different feature networks. Deleting the nodes with single input can also reduce the computational amount and simplify the network. Secondly, an additional skip connection is added in the BiFPN structure to fuse as much information as possible. The involved formulas are:
[0086]
[0087]
[0088]
[0089] Among them, o is the output of the fused feature, w i is the learnable weight, ∈ is a very small value to avoid the denominator being 0, and I i is the i-th input feature; is the intermediate feature of the current layer, is the input feature of the current layer, is the input feature of the next layer, Resize is the sampling operation to control the feature dimension, and Conv is the convolution operation; is the output feature of the current layer, is the output feature of the previous layer.
[0090] In the structure diagram, CSPGhostSE_3, the number after the underline is the number of module stacks. The SPP module specifically includes: the CBS module, four max-pooling layers Maxpool with different sizes and the Concat operation.
[0091] The four Maxpools with different sizes have filter sizes of 13*13, 9*9, 5*5, and 1*1 respectively, and the strides are all 2.
[0092] The output has three detection heads corresponding to three different-scale feature maps of large, medium, and small. Specific embodiments:
[0094] The sources of the experimental data are all collected by industrial cameras in the actual production workshop of textile enterprises. During the data construction process, it is relatively difficult to obtain defective cloth samples. Therefore, a part of the defective samples are taken from the defective products screened out by the cloth inspectors in their previous work and obtained by re-photographing with the camera; the number of defective images is small and is obtained by artificial manufacturing and then photographing according to the causes of defects in the actual production environment.
[0095] The cloth defect images include fabrics of polyester, cotton, silk and grey cloth materials. There are a total of three types of defects: heavy / broken warp, broken weft, and blurred weft. Each image contains one or more types of defects.
[0096] The cloth images with defects collected by the camera have an original resolution of 4096×2048. For the convenience of annotation and training, the original images were cropped to a fixed size of 512×512, resulting in a total of 21,193 images. And according to the 8:2 division ratio, they were divided into a training set of 16,953 images (including 3,672 negative samples without annotation information) and a test set of 4,240 images.
[0097] The software and hardware configurations are shown in Table 1:
[0098] Table 1 Software and Hardware Configurations
[0099]
[0100] The experimental hyperparameter settings are shown in Table 2:
[0101] Table 2 Experimental Hyperparameters
[0102]
[0103] The anchor settings are [36, 20, 10, 114, 119, 13], [26, 154, 18, 422, 159, 47], and [401, 30, 35, 471, 431, 72] respectively.
[0104] As an implementable way, as Figure 4 shown, for the improved core YOLOv5s-GSBiFPN detection network, the regression box prediction loss, classification loss, and confidence loss of different losses are all steadily decreasing and finally tend to converge. The variation range is relatively large in the first 50 epochs during the test stage, but generally shows a downward trend and tends to be stable after 150 epochs.
[0105] The experimental evaluation indicators mainly use mAP, Top1 accuracy (Top-1), computational volume (GFLOPs), number of parameters (Parameters), model size (Size), and model inference speed (Speed) as evaluation indicators.
[0106] Furthermore, the effects of the improved model were compared on the test set, and the results are shown in Table 3:
[0107] Table 3 Comparison of the Effects of YOLOv5 and the Improved Model
[0108]
[0109] Table 3 presents the experimental results according to the evaluation metrics. For all metrics except mAP and Top-1 where higher values indicate better performance, lower values are preferred for the rest of the metrics, and the bolded values represent the optimal ones. It can be seen from Table 3 that the combination of the lightweight network GhostNet, the CSP structure integrated with the SE layer, and the BiFPN bidirectional weighted pyramid structure achieves feature extraction ability comparable to that of convolution. While improving the detection accuracy, it significantly reduces the number of parameters and the computational complexity. For the mAP metric, YOLOv5s-GSBiFPN has increased by 1.8% compared to the original model, and for the Top-1 accuracy metric, it has increased by 3.5%.
[0110] Generally speaking, the improved algorithm has obvious advantages in terms of detection accuracy and detection speed.
[0111] As an implementable approach, as Figure 5 shown, from the change curve of the mAP value of the core detection network, it can be seen that the mAP value of the improved model YOLOv5s-GSBiFPN is in a steadily increasing state as the number of iterations increases. After 150 epochs of iteration, the mAP value of the improved model gradually becomes higher than that of the original model, and after 250 epochs of iteration, it is significantly higher than the original model, verifying the effectiveness of the improvement strategy.
[0112] Furthermore, the detection model and the classification model were cascaded for testing to analyze the false detection and missed detection situations of the models. The number of experimental data samples was 4240 (including positive and negative samples), among which the number of defect-free negative samples was 484, and the number of pictures with defects was 3756. The experiment counted the false detection number, missed detection number, recognition rate, and the detection speed FPS of different models. Taking pictures as the unit, if a defect is detected, it is regarded as a positive detection, and it is not further divided whether a picture contains multiple different defects. The results are shown in Table 4:
[0113] Table 4 Comparison results of model detections
[0114]
[0115] In the cascade network structure described above, the first-level detection network detects the cloth pictures to detect and locate the defects in the pictures; the second-level classifier further classifies and judges the detection boxes output by the detection network, and cooperates with the detection network to filter out false detections, and only outputs the prediction results judged as defects by the classifier. After adopting the cascade network strategy of detection + classification, on 4240 test pictures, the false detection number of the improved model was further reduced from 261 to 192. The defect recognition rate was significantly improved, and finally a recognition rate of 92.1% and a detection speed of 71 frames per second were obtained on 4240 test pictures.
[0116] Finally, it should be noted that although the above embodiments have been described in the text and drawings of the specification of this application, the patent protection scope of this application cannot be limited thereby. Any technical solutions obtained by equivalent structure or equivalent process substitution or modification based on the essential concept of this application and using the content recorded in the text and drawings of the specification of this application, as well as those directly or indirectly implementing the technical solutions of the above embodiments in other related technical fields, are all included in the patent protection scope of this application.
Claims
1. A cloth defect detection method based on a lightweight cascaded network, characterized in that, the method comprises the following steps: 1) Use an industrial camera and an LED lighting device to construct a well-lit and stable imaging shooting environment, and collect cloth images; in the shooting environment, a front lighting method is adopted, the camera and the light source are placed on the same side and kept parallel to the cloth to be inspected; 2) Collect defective cloth images containing different defects, and perform data augmentation on the cloth images to expand the number of samples in the data set and balance the number of defective samples in different categories; adopt a segmented annotation strategy to annotate the defective areas of the cloth images so that the annotation rectangular boxes fit the defective areas; after the samples containing defects are annotated, add cloth images without defects as negative samples according to a 1:1 ratio; then divide all samples into a training set and a test set according to an 8:2 ratio to construct a cloth defect data set; 3) Construct a YOLOv5s-GSBiFPN lightweight cloth defect detection model; construct a dual-path attention CSPGhostSE structure as the core feature extraction module of the backbone network; And extract deep features by stacking multiple layers of CSPGhostSE structures, and then map the features into a fixed dimension in the SPPF module; Use a three-layer stacked weighted bidirectional pyramid BiFPN structure in the feature fusion stage to fully fuse multi-scale features and input them into 3 detection heads to detect defect targets of large, medium and small scales respectively; The detailed steps include: 3.1) Construct a dual-path attention CSPGhostSE structure; the dual-path attention CSPGhostSE structure includes two branches, and each branch is composed of a GhostConv phantom convolution, an SE attention module and a CBS module; branch 1 includes a CBS module, a stacked multi-layer GhostBottleneck structure with a stride of 1, and a GhostBottleneck structure with a stride of 2 connected in sequence; branch two includes a CBS module and a GhostBottleneck structure with a stride of 2 connected in sequence; the output features of the two branches are merged through a Concat operation, and after passing through a batch normalization layer and an activation function, they are input into the lower layer network; The CBS module includes a Conv2d convolutional layer with a kernel size of 1*1, a batch normalization BN layer, and an activation function SiLU layer; The input features are transformed through a Conv2d convolutional layer with a kernel size of 1*1, and the number of convolutional kernels is set to 1 / 2 of the input feature map to adjust the number of output feature channels; The GhostBottleneck structure with a stride of 1 includes two layers; the upper layer is two GhostConv phantom convolutions and an SE attention module connected in sequence; The lower layer is a short connection structure that directly maps the original input features; Fuse the two layers of features by adding; the GhostBottleneck structure with a stride of 2 adds a depthwise convolution DWConv layer with a stride of 2 between the two upper GhostConv phantom convolutions, and the SE attention module is located after the DWConv layer; The GhostConv phantom convolution first uses ordinary convolution to generate basic features, then uses deep convolution on the basic features obtained by convolution, performs linear transformation on the basic feature map channel by channel, generates the other half of the ghost feature map, and finally splices the basic feature and the ghost feature map; 3.2) In the feature fusion stage, a three-layer stacked weighted bidirectional pyramid BiFPN structure is used to replace the original PAN structure. In the multi-scale feature fusion, weighted distinction is made according to the different contributions of different resolution features to the final output of the network, and multi-scale features are fused. The fusion formula is defined as: Among them, o is the output of the fused feature, w i is the learnable weight, ∈ is an extremely small value to avoid a zero denominator, I i is the i-th input feature; is the intermediate feature of the current layer, is the input feature of the current layer, is the input feature of the next layer, Resize is the sampling operation to control the feature dimension, and Conv is the convolution operation; is the output feature of the current layer, is the output feature of the previous layer; 4) Optimize the loss function and use the focal loss function to calculate the classification loss; 5) Establish a binary classification data set and construct a secondary classification network; the binary classification data set is obtained from the cloth defect data set by cropping, without subdividing the defect categories, and only includes two categories: defective samples and non-defective samples; the defective samples are obtained by cropping the original image by expanding 2 pixels in width and height according to the original annotation box information, and not expanding when exceeding the width and height of the original image; the non-defective samples are obtained by randomly cropping the cloth image without defects according to the average size of the existing defective area; the secondary classification network adopts the ResNet18 classification model structure; 6) Train the improved lightweight cloth defect detection model and the binary classification model; when the training iteration is close to 0 and the model loss curve tends to be flat, stop the training and obtain the optimal model; 7) Constructing a cascade network architecture of detection and classification; the cascade network architecture of detection and classification is composed of a primary detection network and a secondary classification network; the primary detection network is responsible for locating and classifying defects in the cloth image and outputting a prediction box, and the secondary classification network is responsible for further performing a binary classification judgment on whether the prediction box area output by the detection network contains defects, and filtering the detection results; 8) Input the cloth image to be tested into the cascade network model to perform defect detection, and output the detection results of the cloth defects and the defect target location information.
2. According to the cloth defect detection method based on lightweight cascade network as described in claim 1, It is characterized in that The data enhancement in step 2) includes: The defective area of a specific category is cut from the original image, and then randomly pasted back to the negative sample image without defects through Poisson fusion to form a new cloth defect sample and expand the data set; Using the mosaic data enhancement method, four pictures are randomly read, randomly scaled, and then spliced together, and the merged label information is processed; then the spliced pictures are randomly horizontally flipped and enhanced and affine transformed without color space transformation.