Cloth defect detection method and system based on multi-scale fusion attention mechanism
By improving the YOLOv7 model, combining the multi-scale fusion attention mechanism and Inner-SIoU Loss loss function, the problems of low efficiency and low accuracy of cloth defect detection in the existing technology are solved, and efficient and accurate cloth defect detection is achieved.
Patent Information
- Application Number
- CN202510196873.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-10
AI Technical Summary
In the detection of cloth defects, the prior art has problems such as low detection efficiency, low detection accuracy, complex texture background and large scale span in the detection of defects.
The cloth defect detection method based on a multi-scale fusion attention mechanism is adopted. By improving the YOLOv7 model, the PConv module, CA attention module and the reconstruction SPPCSPC module are introduced, and combined with the Inner-SIoU Loss loss function, the feature extraction ability and detection accuracy of the model are improved.
It realizes efficient and accurate cloth defect detection, and can quickly detect multiple defect types under limited computing resources, improving detection accuracy and speed.
Smart Images

Figure CN120125948A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of light industry production detection, and specifically relates to a small target cloth defect detection method and system based on improved YOLOv7. Background Art
[0002] The textile industry is a traditional pillar industry and an important people's livelihood industry in China's national economy, and it is also an advantageous industry in China's international competition, playing an important role in prospering the market, expanding exports, absorbing employment, increasing farmers' income, and promoting urbanization development. China's textile industry needs to continuously optimize its industrial structure and adopt advanced technologies to upgrade and transform traditional industries.
[0003] For the textile industry, the inspection and control of cloth quality are important components in the textile industry, directly affecting the production efficiency of enterprises. The cloth production process includes raw material processing, spinning, weaving, printing and dyeing, etc. Situations such as debris accidentally mixing into raw materials, operation errors, and mechanical equipment failures may all cause defects on the cloth surface. The more processing levels there are, the greater the probability of defects occurring. Common cloth defects include: holes, sizing spots, foreign fibers, thick warp, thick weft, knots, etc. Defects on the cloth surface not only affect its appearance but also significantly reduce its quality. Market surveys show that defects on the cloth surface will cause its selling price to drop by about 50%. Therefore, the defect detection of cloth is an important link in the quality control of cloth during the production process.
[0004] Traditional cloth defect detection is mainly completed through manual observation, which requires the human eye to observe the cloth and check whether there are defects on its surface. This type of method has many disadvantages: it requires a large amount of human resources; the defect detection workers are highly subjective; long-term high-intensity eye use causes irreversible damage to the eyes of inspectors; the detection efficiency is low, etc. According to relevant survey results, the highest manual detection speed is only 0.36 m / s, and even for skilled inspectors, the successful detection rate is only 60% - 75%, and most of the detected defects are relatively large in size. With the continuous improvement of the production efficiency of textile enterprises, relying solely on manual detection of cloth defects is difficult to adapt to the rapid development of the textile industry. Therefore, many computer vision technologies have been widely applied to cloth defect detection.
[0005] Traditional image processing algorithms aim at defect recognition, and separate the defect area from the cloth defect image through corresponding feature extraction algorithms to identify whether there are defects in the cloth image. After extracting the defect image features, subsequent cloth defect detection algorithms use various classification methods in machine learning to classify the defect types. The advantage of this type of method is high feature recognition accuracy and fast detection speed, but there are problems that some feature extraction algorithms can only cover some defect types and have poor detection effects for defects with complex texture backgrounds and large scale spans.
[0006] With the rapid development of deep learning in the field of computer vision, classification algorithms and object detection algorithms based on deep learning have achieved a leapfrog breakthrough in accuracy on numerous public datasets. Compared with traditional methods, deep learning models can automatically extract image features from a large amount of image data without manual participation. This feature has stronger generalization ability and robustness compared to a single manually designed method and can complete defect detection tasks with large scales and multiple categories. For the cloth defect detection task, there are significant differences in detection accuracy and speed when using different object detection algorithms. In actual cloth quality inspection applications, it is necessary to achieve as high a detection accuracy as possible while meeting the requirements of industrial production speed. Summary of the Invention
[0007] Object of the Invention: In order to overcome the deficiencies in the prior art, the present invention provides a cloth defect detection method based on a multi-scale fusion attention mechanism, which has powerful feature extraction ability, autonomous learning ability, and relatively high detection accuracy and speed, and can realize cloth defect detection by using limited computing power resources.
[0008] Technical Solution: To achieve the above object, the technical solution adopted by the present invention is as follows:
[0009] A cloth defect detection method based on a multi-scale fusion attention mechanism, comprising the following steps:
[0010] Step S1: Obtain cloth defect images, analyze the defect characteristics of the cloth defect dataset, and screen typical cloth defect types to construct a dataset.
[0011] Step S2: Expand the constructed dataset through a data augmentation method to solve the problem of sample imbalance.
[0012] Step S3: Improve the YOLOv7 model to obtain an improved YOLOv7 model.
[0013] Step S4: Use the expanded dataset to train the improved YOLOv7 model to obtain a trained improved YOLOv7 model.
[0014] Step S5: Detect cloth defects through the trained improved YOLOv7 model.
[0015] Preferably, the improved YOLOv7 model includes a first CBS layer 1, a first CBS layer 2, a first CBS layer 3, a first CBS layer 4, an ELANP module 1, a first MP1 module 1, an ELANP module 2, a CA attention module 1, a first MP1 module 2, an ELANP module 3, a CA attention module 2, a first MP1 module 3, an ELANP module 4, an SPPCSPC module, an UP upsampling module 1, a C fusion module 1, an ELAN module 1, an UP upsampling module 2, a C fusion module 2, an ELAN module 2, a second MP2 module 1, a C fusion module 3, an ELAN module 3, a second MP2 module 2, a C fusion module 4, an ELAN module 4, a Respconv module 1, and a CBM module 1, which are connected in sequence. The input of the Respconv module 1 is connected to the output of the SPPCSPC module. A second CBS layer 1 is provided between the ELANP module 3 and the C fusion module 1, and a second CBS layer 2 is provided between the ELANP module 2 and the C fusion module 2. The input of the C fusion module 3 is connected to the output of the ELAN module 1, and the input of the C fusion module 4 is connected to the output of the SPPCSPC module. The output of the ELAN module 3, the Respconv module 2, and the CBM module 2 are connected in sequence, and the input of the Respconv module 2 is connected to the output of the second CBS layer 1. The output of the ELAN module 2, the Respconv module 3, and the CBM module 3 are connected in sequence, and the input of the Respconv module 3 is connected to the output of the second CBS layer 2.
[0016] Preferably, the method for improving the YOLOv7 model in step S3 includes the following steps:
[0017] Step S31: Taking the YOLOv7 model as the basic architecture, both branches of the original ELAN first pass through a 1×1 ordinary convolution module to change the number of channels. One of the branches needs to further pass through a 3×3 ordinary convolution module to change the number of channels for feature extraction. Replace the 3×3 ordinary convolution Conv with a PConv convolution, and finally stack the four features together to obtain the final feature extraction result.
[0018] Step S32: Introduce a CA attention module, and introduce the CA attention module before the ELAN structure, MP structure, and CBS fusion in the Backbone. The CA attention module takes spatial information as part of the attention and is more sensitive to the defect location information in the cloth defect detection task.
[0019] Step S33: Reconstruct the SPPCSPC module, cut off 2 CBS layers before the pooling layer, and embed a CA attention mechanism before the pooling layer:
[0020] Crop off 2 CBS layers before the pooling layer to reduce the filtering of small target edge information by the convolutional layer, and only apply one layer of CBS to smooth the features, reducing the computational load and parameter scale. Introduce the CA attention mechanism, embed the attention mechanism before the pooling layer, enhance the attention to the regions with high feature weights, reduce the loss of effective features, highlight the key position information of the defects, and at the same time suppress the expression of confusing features.
[0021] S34: Use the BiFPN structure as the feature fusion network of the YOLOv7 network model to optimize the original FPN and PANet structures:
[0022] Step S35: Use the Inner-SIoU Loss function instead of the original loss function as the loss function for target box regression.
[0023] Preferably: The formula of the Inner-SIoU Losss loss function is as follows:
[0024]
[0025] where L Inner-SIoU represents the Inner-SIoU Loss function, Δ is the distance loss, and Ω is the shape loss. represents the abscissa of the centroid of the ground truth box, ω gt is the width of the ground truth box, ratio represents the scale factor for generating the auxiliary boundary, x c represents the abscissa of the centroid of the predicted box, ω represents the width of the predicted box. ordinate of the centroid of the ground truth box, h gt is the height of the ground truth box, y c ordinate of the centroid of the predicted box, h represents the height of the predicted box, y c ordinate of the centroid of the predicted box.
[0026] Preferably: The implementation process of the CA attention module in step S32: It is divided into two parts: coordinate information embedding and coordinate attention generation. By performing global average pooling on the input feature map in the horizontal and vertical directions respectively, the feature maps in the horizontal and vertical directions are obtained respectively. Use the pooling kernel to encode the features in the horizontal and vertical directions of the input feature tensor. The outputs of the c-th channel with height h and the c-th channel with width w are obtained by the following formulas respectively:
[0027]
[0028] Wherein, W and H represent the width and height of the pooling kernel. The above two transformations aggregate features along two spatial directions respectively, generating a pair of direction-aware feature maps. Then, a shared 1×1 convolution is used to transform the two feature maps, obtaining an intermediate feature map with a size of C / r×1×(H + W), where r is the downsampling ratio. Finally, the intermediate feature mapping information in the horizontal and vertical directions is represented by f, and the calculation formula of f is as follows:
[0029] f = δ(F ′ ([z h ,z w ))
[0030] Wherein, δ is a non-linear activation function. f is split into two independent tensors along the spatial dimension, and two 1×1 convolution transformations F h and F w transform f h and f w into tensors g h and g w with the same number of channels as the input X, and the calculation formulas are as follows:
[0031] g h = σ(F h (f h ))
[0032] g w = σ(F w (f w ))
[0033] Wherein, σ is the sigmoid activation function, and the output g h and g w are enhanced as attention weights. Finally, the output y c (i, j) of the coordinate attention block can be expressed as:
[0034]
[0035] Wherein, y c (i, j) represents the output feature map, x c (i, j) represents the input feature map, and represent the attention weights in the horizontal and vertical spatial directions.
[0036] Preferably: in step S2, for the data set constructed in step S1. Using offline enhancement method, the specific data enhancement method is as follows: random rotation and flipping. At the same time, combined with Mosaic enhancement and adaptive enhancement, a higher enhancement probability is set for the minority class. The Mosaic enhancement method is specifically: randomly crop four images, splice the four images together to obtain a new image, adjust the corresponding bounding box annotation, pass the new image to the network for learning, and enrich the background of the detected object by increasing the batch size in disguise to improve the training effect. Adaptive enhancement defines a series of enhancement operations and dynamically adjusts the intensity of these operations according to the current performance of the model. The expanded data set is divided into training set, validation set and test set in a ratio of 7:1:2.
[0037] Preferably: in step S4, the network model is trained: the improved YOLOv7 model is trained using the training set to obtain a trained improved YOLOv7 model, parameters are set for the configuration file of the improved YOLOv7 model, the yaml file with the set parameters and the improved YOLOv7 model are placed in a computer with a configured environment, and the marked images in the training set and the validation set are used for training. During the training process, the effect of the training at each stage is obtained, and the process monitoring parameters are set to observe the mAP value of the training, and the trained network model weights are saved after the training is completed.
[0038] Preferably: in step S1, obtain the cloth defect image and defect label data, including the minimum bounding box of the complete defect target, mark the specific location and defect category of the defect, and integrate all the information in a json file. By using matplotlib and seaborn drawing tools, parse the labeled data of each minimum bounding box, know the category and corresponding position coordinates of each defect, and screen out typical cloth defect types.
[0039] Preferred: Typical cloth defect types include knots, holes, three threads, coarse warp, coarse weft, and starch spots.
[0040] Another object of the present invention is to provide a cloth defect detection system based on a multi-scale fusion attention mechanism, comprising an input unit, an improved YOLOv7 model unit, and an output unit, wherein:
[0041] The input unit is used to input the cloth image to be detected.
[0042] The improved YOLOv7 model unit detects and identifies the cloth image to be detected by using the trained improved YOLOv7 model to obtain cloth defects.
[0043] The output unit is used to output cloth defects.
[0044] Compared with the prior art, the present invention has the following beneficial effects:
[0045] The present invention is improved based on the YOLOv7 network. First, by analyzing the characteristics of fabric defects, representative typical defect types are selected for subsequent model migration to other application scenarios of fabric defects. The dataset images are subjected to data augmentation processing to alleviate the imbalance between samples. From the perspective of lightweight models, PConv is introduced into the ELAN module of the YOLOv7 backbone network to reduce the number of parameters while ensuring the model accuracy and improving the detection speed. The CA attention mechanism is added to the backbone part to obtain cross-channel information, as well as perceptual and position-sensitive information, improving the detection accuracy and network expression ability. The SPPCSPC structure is reconstructed to reduce the loss of effective features and highlight the key position information of the defects. The original loss function is replaced with SIoU Loss to reduce the imbalance between samples and improve the positioning accuracy. In the NVIDIA GeForce RTX4060 environment, the mAP of this method for the pure-color fabric defect dataset with a denim material background is 92.9%. Experiments show that this method has high detection accuracy and fast inference speed, which helps textile factories accurately and real-time detect targets before the dyeing and bleaching process, reduce losses, and improve the quality of fabrics. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 It is a flowchart of the fabric defect detection method based on the multi-scale fusion attention mechanism in the present invention.
[0047] Figure 2 It is a distribution diagram of the aspect ratios of the original 34 defects in the fabric dataset of the present invention.
[0048] Figure 3 It is a distribution diagram of the aspect ratios of the six selected typical fabric defects in the present invention.
[0049] Figure 4 It is a diagram of the improved YOLOv7 network model in the present invention.
[0050] Figure 5 It is a diagram of the training verification loss and evaluation index results of the improved YOLOv7 model in the present invention.
[0051] Figure 6 It is a diagram (partial) of the fabric defect detection results after using the improved YOLOv7 network in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0052] The present invention is further explained below in conjunction with the accompanying drawings and specific embodiments. It should be understood that these examples are only used to illustrate the present invention and are not used to limit the scope of the present invention. After reading the present invention, various equivalent forms of modifications to the present invention by those skilled in the art all fall within the scope defined by the claims attached to this application.
[0053] The present invention designs a cloth defect detection method based on a multi-scale fusion attention mechanism, such as Figure 1 As shown in the figure, first, the defect characteristics of the cloth defect dataset are analyzed, and typical cloth defect types are screened to construct a dataset. Secondly, the screened dataset is enhanced and divided into training set, validation set and test set. Then the YOLOv7 model is improved to improve the model's target detection capability. Finally, the improved YOLOv7 model is trained using the dataset, and the trained model is evaluated, adjusted and optimized, which specifically includes the following steps:
[0054] Step S1: Obtain a cloth defect image, analyze the defect characteristics of the cloth defect dataset, and select typical cloth defect types to construct a dataset. Typical cloth defect types include knots, holes, three threads, coarse warp, coarse weft, and starch spots.
[0055] In another embodiment, a cloth defect image and defect label data are obtained, including the minimum bounding box of the complete defect target, and the specific location and defect category of the defect are marked, and all information is integrated in a json file. By using matplotlib and seaborn drawing tools, the annotation data of each minimum bounding box is parsed, the category of each defect and the corresponding position coordinates are known, and typical cloth defect types are screened out.
[0056] The data set used in the experiment of this embodiment covers many types of defects and has high resolution. Each picture has one or more defects of different types. There may also be multiple defects of the same type, with a total of 34 defect categories. With the help of matplotlib and seaborn drawing tools, the aspect ratio information of each defect type is parsed, and the aspect ratio distribution map of various defects of cloth is drawn. Defects with smaller differences are classified into one category, and screened on the basis of the original data set. Finally, 6 common defects (knots, holes, three wires, coarse warp, coarse weft, and starch spots) of denim material solid color cloth are selected for detection and research.
[0057] Step S2: Use data augmentation methods to expand the constructed data set and solve the sample imbalance problem.
[0058] In another embodiment, for the data set constructed in step S1. Using an offline enhancement method, the specific data enhancement method is as follows: random rotation and flipping. At the same time, combined with Mosaic enhancement and adaptive enhancement, a higher enhancement probability is set for the minority class. The Mosaic enhancement method is specifically: randomly cropping four images, splicing the four images together to obtain a new image, adjusting the corresponding bounding box annotations, passing the new image to the network for learning, and enriching the background of the detected object by increasing the batch size in disguise to improve the training effect. Adaptive enhancement defines a series of enhancement operations and dynamically adjusts the intensity of these operations according to the current performance of the model. The expanded data set is divided into training set, validation set and test set in a ratio of 7:1:2.
[0059] In this embodiment, an offline enhancement method is adopted. The specific data enhancement methods are as follows: random rotation. Random rotation generally refers to the random rotation of the input image [0, 360); flipping, including horizontal flipping and vertical flipping, horizontal flipping is to flip the image along the Y axis, and vertical flipping is to flip the image along the X axis. Combined with Mosaic enhancement and adaptive enhancement, four images are randomly cropped, and the four images are spliced together to obtain a new image. The corresponding bounding box annotations are adjusted, and the new image is passed to the network for learning. By increasing the batch size in disguise, the background of the detected object is enriched to improve the training effect. Adaptive enhancement defines a series of enhancement operations (such as rotation, flipping, color adjustment, etc.), and dynamically adjusts the intensity of these operations according to the current performance of the model. Finally, the data set is divided into training set, validation set and test set in a ratio of 7:1:2.
[0060] Step S3: Improve the YOLOv7 model to obtain an improved YOLOv7 model.
[0061] The improved YOLOv7 model includes a first CBS layer 1, a first CBS layer 2, a first CBS layer 3, a first CBS layer 4, an ELANP module 1, a first MP1 module 1, an ELANP module 2, a CA attention module 1, a first MP1 module 2, an ELANP module 3, a CA attention module 2, a first MP1 module 3, an ELANP module 4, an SPPCSPC module, an UP upward adoption module 1, a C fusion module 1, an ELAN module 1, an UP upward adoption module 2, a C fusion module 2, an ELAN module 2, a second MP2 module 1, a C fusion module 3, an ELAN module 3, a second MP2 module 2, a C fusion module 4, an ELAN module 4, a Respconv module 1, and a CBM module 1, which are connected in sequence. The input of the Respconv module 1 is connected to the output of the SPPCSPC module. A second CBS layer 1 is provided between the ELANP module 3 and the C fusion module 1, and a second CBS layer 2 is provided between the ELANP module 2 and the C fusion module 2. The input of the C fusion module 3 is connected to the output of the ELAN module 1, and the input of the C fusion module 4 is connected to the output of the SPPCSPC module. The output of the ELAN module 3, the Respconv module 2, and the CBM module 2 are connected in sequence, and the input of the Respconv module 2 is connected to the output of the second CBS layer 1. The output of the ELAN module 2, the Respconv module 3, and the CBM module 3 are connected in sequence, and the input of the Respconv module 3 is connected to the output of the second CBS layer 2.
[0062] A method for improving the YOLOv7 model, with the YOLOv7 model as the basic architecture, replaces some standard convolutions in the backbone network with PConv modules to reduce network computational complexity and improve network detection speed; adds CA to enhance the ability to extract position features of tiny defects on the cloth; reconstructs the SPPCSPC module to improve the ability to detect small targets; optimizes BiFPN to achieve simple and fast multi-scale feature fusion; introduces a new loss function SIOU with angular loss to promote the fitting of the ground truth box and the prediction box and improve the accuracy of defect prediction. Specifically, it includes the following steps:
[0063] Step S31: With the YOLOv7 model as the basic architecture, both branches of the original ELAN first pass through a 1×1 ordinary convolution module to change the number of channels, and one of the branches needs to further pass through a 3×3 ordinary convolution module to change the number of channels for feature extraction. Replace the 3×3 ordinary convolution Conv with a PConv convolution, and finally stack the four features together to obtain the final feature extraction result. Using the PConv module can enhance the ability to generate feature maps while keeping the network lightweight, thus bringing important benefits to the training and inference of the cloth defect detection model.
[0064] Step S32: Introduce the CA attention module before the fusion of the ELAN structure, MP structure, and CBS in the Backbone. The CA attention module takes spatial information as part of the attention and is more sensitive to the defect location information in the cloth defect detection task.
[0065] The implementation process of the CA attention module is divided into two parts: coordinate information embedding and coordinate attention generation. By performing global average pooling on the input feature map in the horizontal and vertical directions respectively, feature maps in the horizontal and vertical directions are obtained respectively. Pooling kernels are used to encode the features of the input feature tensor in the horizontal and vertical directions. The outputs of the c-th channel with height h and the c-th channel with width w are obtained by the following formulas respectively:
[0066]
[0067] In the formula, W and H represent the width and height of the pooling kernel. The above two transformations aggregate features along two spatial directions respectively, generating a pair of direction-aware feature maps. Then, a shared 1×1 convolution is used to transform the two feature maps, obtaining an intermediate feature map of size C / r×1×(H + W), where r is the downsampling ratio. Finally, the intermediate feature mapping information in the horizontal and vertical directions is represented by f, and the calculation formula of f is as follows:
[0068] f = δ(F ′ ([z h ,z w ))
[0069] In the formula, δ is a non-linear activation function. Split f along the spatial dimension into two independent tensors, and two 1×1 convolution transformations F h and F w transform f h and f w into tensors g h and g w with the same number of channels as the input X. The calculation formulas are as follows respectively:
[0070] g h = σ(F h (f h ))
[0071] g w = σ(F w (f w ))
[0072] In the formula, σ is the sigmoid activation function, and the output g h and g w are enhanced as attention weights. Finally, the output y c(i,j) can be expressed as:
[0073]
[0074] where y c (i,j) represents the output feature map, and x c (i,j) represents the input feature map, and represent the attention weights in the horizontal and vertical spatial directions.
[0075] Step S33: Reconstruct the SPPCSPC module, cut off 2 CBS layers before the pooling layer, and embed the CA attention mechanism before the pooling layer;
[0076] Cut off 2 CBS layers before the pooling layer to reduce the filtering of small target edge information by the convolutional layer, only apply one layer of CBS to smooth the features, reduce the computational amount and parameter scale. Introduce the CA attention mechanism, embed the attention mechanism before the pooling layer, enhance the attention to the regions with high feature weights, reduce the loss of effective features, highlight the key position information of the defects, and at the same time suppress the expression of confusing features.
[0077] S34: Use the BiFPN structure as the feature fusion network of the YOLOv7 network model to optimize the original FPN and PANet structures:
[0078] The BiFPN structure is a weighted bidirectional feature pyramid network. Based on the FPN structure, each node of the BiFPN network will fuse different feature layers by weighted fusion of the input feature vectors. Based on the PANet structure, the BiFPN network repeats the top-down and bottom-up bidirectional fusion. Finally, three BiFPN basic structures are stacked to output the fused low-dimensional and high-dimensional features. The BiFPN uses the bidirectional fusion idea to reconstruct the top-down and bottom-up bidirectional channels in addition to the forward propagation, fuse the feature information of different scales from the backbone network, unify the feature resolution scale through upsampling and downsampling, and add double lateral connections between the features of the same scale to alleviate the loss of feature information caused by too many network levels.
[0079] Step S35: Use the Inner-SIoU Loss function to replace the original loss function as the loss function for target box regression:
[0080] The original YOLOv7 uses CIoU Loss as the coordinate loss function, and its formula is as follows:
[0081]
[0082] where b and b gt represent B and B respectively gtThe center point, ρ 2 is the Euclidean distance, c is the diagonal distance of the minimum bounding rectangle of B and B gt μ is a positive balance parameter, and v represents the consistency of the aspect ratio between the gt and the predicted bounding box. For the parameters μ and v, we have:
[0083]
[0084] In the formula, w and h respectively represent the width and height of the predicted bounding box B, w gt and h gt respectively represent the width and height of the predicted bounding box B gt
[0085] CIoU Loss does not consider the mismatch between the directions of the predicted bounding box and the ground truth box, which affects the convergence speed and efficiency. In this paper, Inner-SIoU Loss is used to replace CIoU Loss. Among them, SIoU Loss consists of four parts: angular loss Λ, distance loss Δ, shape loss, and intersection over union loss. The formula of SIoU Loss is as follows:
[0086] Angular loss
[0087]
[0088] Distance loss
[0089]
[0090] Shape loss
[0091]
[0092] In the formula, θ controls the degree of attention to the shape loss. To avoid excessive attention to the shape loss and reduce the movement of the predicted bounding box, the range of the θ parameter is between 2 and 6. The definition of SIoU Loss is:
[0093]
[0094] IoU inner is defined as:
[0095]
[0096] inter and union are defined as:
[0097]
[0098] union = (ω gt × h gt )×(ratio) 2 +(ω×h)×(ratio) 2 -inter
[0099] b r , b l , b b, b t is defined as follows:
[0100]
[0101] The definition of Inner-SIoU Loss is:
[0102]
[0103] In summary:
[0104]
[0105] Among them, L Inner-SIoU represents the Inner-SIoU Loss function, L SIoU represents the SIoU Loss function, IoU represents the intersection over union loss between the predicted bounding box and the ground truth bounding box, IoU inner represents the inner intersection over union between the predicted bounding box and the ground truth bounding box, Λ is the angular loss, representing the angular difference between the predicted bounding box and the ground truth bounding box, c h represents the height difference between the center point of the ground truth bounding box and the center point of the predicted bounding box, γ represents the distance from the center point of the ground truth bounding box to the center point of the predicted bounding box, α represents the angle between the predicted bounding box and the ground truth bounding box, Δ is the distance loss, representing the distance difference between the predicted bounding box and the ground truth bounding box, σ represents the angular loss function, p x represents the horizontal axis distance loss, c x represents the width of the minimum bounding rectangle of the predicted bounding box and the ground truth bounding box, p y represents the vertical axis distance loss, c y represents the height of the minimum bounding rectangle of the predicted bounding box and the ground truth bounding box, represents the horizontal axis coordinate of the center point of the bounding box, b cx represents the horizontal axis coordinate of the center point of the predicted bounding box, represents the vertical axis coordinate of the center point of the ground truth bounding box, b cy represents the vertical axis coordinate of the center point of the predicted bounding box, Ω is the shape loss, representing the shape difference between the predicted bounding box and the ground truth bounding box, ω w represents the width difference, v h represents the height difference, w represents the width of the predicted bounding box, w gt the width of the ground truth bounding box, h represents the height of the predicted bounding box, hgt Denote the height of the ground truth box, inter denote the area of the intersection region between the predicted box and the ground truth box under the setting of the auxiliary bounding box, union denote the area of the union region between the predicted box and the ground truth box under the setting of the auxiliary bounding box. Denote the abscissa of the centroid of the ground truth box. The ordinate of the centroid of the ground truth box, x c Denote the abscissa of the centroid of the predicted box, y c The ordinate of the centroid of the predicted box. Respectively denote the coordinates of the left, right, upper, and lower boundaries of the auxiliary bounding box of the ground truth box, b l , b r b t , b b Respectively denote the coordinates of the left, right, upper, and lower boundaries of the predicted auxiliary bounding box. In the formula, ratio denotes the scale factor for generating the auxiliary bounding box, and usually takes values in [0.5, 1.5].
[0106] Step S4: Use the augmented dataset to train the improved YOLOv7 model to obtain the trained improved YOLOv7 model.
[0107] In another embodiment, train the network model: Set parameters for the improved YOLOv7 network configuration file, put the yaml file with the set parameters and the improved YOLOv7 network structure into a computer with a configured environment, and use the marked pictures in the training set and the validation set for training. During the training process, obtain the training effect of each stage, and set process monitoring parameters to observe the mAP value of the training. After the training is completed, save the weights of the trained network model.
[0108] Step S5: Detect fabric defects through the trained improved YOLOv7 model.
[0109] The experimental settings basically adopt the official recommended parameter settings of YOLOv7. The improved and original algorithms use the same hyperparameters and pre-trained weights, adopt adaptive anchors, adopt mosaic data augmentation, the input image size is 640×640, the batch-size is 8 during training, the batchsize is 1 during testing, the epoch is 200, the initial learning rate is 0.015, the learning rate momentum is 0.937, and the conf-thres is set to 0.40 during inference.
[0110] The present invention also conducts performance verification on the improved YOLOv7 network model and the original YOLOv7 network model. The experimental environment is: Under the Win10 operating system, the CPU model is Core TMThe i5-12400F has a running memory of 16GB. The GPU model is NVIDIA GeForce RTX 4060, with 8GB of dedicated GPU memory and 16GB of GPU memory. Based on the Pytorch 1.12.0 deep learning framework, the programming language is Python 3.11, and the experiment is conducted in the Pycharm integrated development environment with the CUDA version being 11.7.99. The experimental results are as follows.
[0111] The mAP value of the improved YOLOv7 model is 90.5%, the mAP value of the original YOLOv7 is 79.3%, an increase of 11.2 percentage points; the accuracy value of the improved YOLOv7 model is 90.5%, the accuracy value of the original YOLOv7 model is 73.8%, an increase of 16.7 percentage points; the recall rate of the improved YOLOv7 model is 83.7%, the recall rate of the original YOLOv7 model is 71.2%, an increase of 12.5 percentage points. Generally speaking, the improved YOLOv7 model is superior to the original YOLOv7 model in multiple performance evaluation indicators.
[0112] Output the detection results. The detection result diagram (partial) is as Figure 6 shown. It can be seen that the cloth defect detection model with the multi-scale fusion attention mechanism of the present invention can accurately identify cloth defects.
[0113] In another embodiment, a cloth defect detection system based on a multi-scale fusion attention mechanism is provided, including an input unit, an improved YOLOv7 model unit, and an output unit, where:
[0114] The input unit is used to input the cloth image to be detected.
[0115] The improved YOLOv7 model unit detects and identifies the cloth image to be detected through the trained improved YOLOv7 model to obtain cloth defects.
[0116] The output unit is used to output cloth defects.
[0117] By obtaining the printed cloth defect images, an initial data set is established; the data set is subjected to clustering analysis to obtain the clustering centers, and the clustering center values are input into the YOLOv7 network; on the basis of the YOLOv7 network structure, some standard convolutions of the backbone network are replaced with PConv modules to reduce the network calculation amount and improve the network detection speed; a CA attention mechanism is added to enhance the ability to extract the position features of small cloth defects; the SPPCSPC module is reconstructed to improve the small target detection ability; BiFPN is optimized to achieve simple and fast multi-scale feature fusion; a new loss function SIOU with angular loss is introduced to promote the fitting of the true box and the predicted box and improve the accuracy of defect prediction.
[0118] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A cloth defect detection method based on multi-scale fusion attention mechanism, characterized in that: The following steps are involved: Step S1: Acquire cloth defect images, analyze defect characteristics of cloth defect datasets, and select typical cloth defect types to construct a dataset; Step S2: Expand the constructed data set through data enhancement methods to solve the problem of sample imbalance; Step S3: improving the YOLOv7 model to obtain an improved YOLOv7 model; Step S4: train the improved YOLOv7 model using the expanded data set to obtain a trained improved YOLOv7 model; Step S5: Detect cloth defects using the trained improved YOLOv7 model.
2. The cloth defect detection method based on the multi-scale fusion attention mechanism according to claim 1 is characterized in that: The improved YOLOv7 model includes the first CBS layer 1, the first CBS layer 2, the first CBS layer 3, the first CBS layer 4, the ELANP module 1, the first MP1 module 1, the ELANP module 2, the CA attention module 1, the first MP1 module 2, the ELANP module 3, the CA attention module 2, the first MP1 module 3, the ELANP module 4, the SPPCSPC module, the UP upward adoption module 1, the C fusion module 1, the ELAN module 1, the UP upward adoption module 2, the C fusion module 2, the ELAN module 2, the second MP2 module 1, the C fusion module 3, the ELAN module 3, the second MP2 module 2, the C fusion module 4, the ELAN module 4, the Respconv module 1, and the CBM module 1. The input of nv module one is connected to the output of SPPCSPC module; a second CBS layer one is arranged between the ELANP module three and the C fusion module one, and a second CBS layer two is arranged between the ELANP module two and the C fusion module two; the input of the C fusion module three is connected to the output of the ELAN module one, and the input of the C fusion module four is connected to the output of the SPPCSPC module; the output of the ELAN module three, the Respconv module two, and the CBM module two are connected in sequence, and the input of the Respconv module two is connected to the output of the second CBS layer one; the output of the ELAN module two, the Respconv module three, and the CBM module three are connected in sequence, and the input of the Respconv module three is connected to the output of the second CBS layer two.
3. The cloth defect detection method based on multi-scale fusion attention mechanism according to claim 2 is characterized in that: The method for improving the YOLOv7 model in step S3 includes the following steps: Step S31: Using the YOLOv7 model as the basic architecture, the two branches of the original ELAN first change the number of channels through a 1×1 ordinary convolution module, and one of the branches needs to change the number of channels through a 3×3 ordinary convolution module for feature extraction; the 3×3 ordinary convolution Conv is replaced with a PConv convolution, and finally the four features are superimposed together to obtain the final feature extraction result; Step S32: Introduce the CA attention module, and introduce the CA attention module before the ELAN structure of Backbone is fused with the MP structure and CBS; the CA attention module takes spatial information as part of the attention and is more sensitive to the defect position information in the cloth defect detection task; Step S33: Reconstruct the SPPCSPC module, cut out the two CBS layers before the pooling layer, and embed the CA attention mechanism before the pooling layer: The two CBS layers before the pooling layer are cut off to reduce the filtering of the edge information of small objects by the convolution layer. Only one layer of CBS smoothing features is applied to reduce the amount of calculation and parameter scale. The CA attention mechanism is introduced and embedded before the pooling layer to enhance the attention to the areas with high feature weights, reduce the loss of effective features, highlight the key location information of the defects, and suppress the expression of confusing features. S34: Use the BiFPN structure as the feature fusion network of the YOLOv7 network model to optimize the original FPN and PANet structures: Step S35: Use the Inner-SIoU Loss loss function instead of the original loss function as the loss function for target box regression.
4. The cloth defect detection method based on multi-scale fusion attention mechanism according to claim 3 is characterized by: The formula of Inner-SIoU Loss loss function is as follows: Among them, L Inner-SIoU represents the Inner-SIoU Loss loss function, Δ is the distance loss, Ω is the shape loss, Represents the horizontal coordinate of the centroid of the real box, ω gt The width of the real box, ratio represents the scale factor for generating auxiliary boundaries, x c represents the horizontal coordinate of the centroid of the prediction box, ω represents the width of the prediction box, The vertical coordinate of the centroid of the real box, h gt Indicates the height of the real box, y c The vertical coordinate of the centroid of the prediction box, h represents the height of the prediction box, y c The vertical coordinate of the predicted box centroid.
5. The cloth defect detection method based on multi-scale fusion attention mechanism according to claim 4 is characterized in that: The implementation process of the CA attention module in step S32 is divided into two parts: coordinate information embedding and coordinate attention generation; by performing global average pooling on the input feature map in the horizontal and vertical directions respectively, the feature maps in the horizontal and vertical directions are obtained respectively, and the pooling kernel is used to encode the features of the input feature tensor in the horizontal and vertical directions. The outputs of the c-th channel with a height of h and the c-th channel with a width of w are obtained by the following formulas: In the formula, W and H represent the width and height of the pooling kernel. The above two transformations aggregate features along two spatial directions respectively to produce a pair of feature maps based on direction perception. Then, a shared 1×1 convolution is used to change the two feature maps to obtain an intermediate feature map of size C / r×1×(H+W), where r is the downsampling ratio. Finally, the intermediate feature mapping information in the horizontal and vertical directions is represented by f, and the calculation formula of f is as follows: f=δ(F ′ ([z h ,z w ])) Where δ is a nonlinear activation function; f is split into two independent tensors along the spatial dimension, and two 1×1 convolution transformations F h and F w f h and f w Transformed into a tensor g with the same number of channels as the input X h and g w , the calculation formulas are as follows: g h =σ(F h (f h )) g w =σ(F w (f w )) Where σ is the sigmoid activation function, and the output g h and g w is enhanced as the attention weight; finally, the output y of the coordinate attention block c (i,j) can be expressed as: In the formula, y c (i,j) represents the output feature map, x c (i,j) represents the input feature map, and Represents the attention weights in the horizontal and vertical spatial directions.
6. The cloth defect detection method based on multi-scale fusion attention mechanism according to claim 5 is characterized in that: In step S2, for the dataset constructed in step S1, offline enhancement is used. The specific data enhancement methods are as follows: random rotation and flipping. At the same time, Mosaic enhancement and adaptive enhancement are combined to set a higher enhancement probability for the minority class. The Mosaic enhancement method is specifically as follows: randomly crop four images, splice the four images together to get a new image, adjust the corresponding bounding box annotations, pass the new image to the network for learning, and enrich the background of the detected object by increasing the batch size in disguise to improve the training effect. Adaptive enhancement defines a series of enhancement operations, and dynamically adjusts the intensity of these operations according to the current performance of the model. The expanded dataset is divided into training set, validation set and test set in a ratio of 7:1:
2.
7. The cloth defect detection method based on multi-scale fusion attention mechanism according to claim 6 is characterized in that: In step S4, the network model is trained: the improved YOLOv7 model is trained using the training set to obtain a trained improved YOLOv7 model, parameters are set for the configuration file of the improved YOLOv7 model, the yaml file with the set parameters and the improved YOLOv7 model are placed in a computer with a configured environment, and the marked images in the training set and the validation set are used for training. During the training process, the effect of training at each stage is obtained, and the process monitoring parameters are set to observe the mAP value of the training. After the training, the trained network model weights are saved.
8. The cloth defect detection method based on multi-scale fusion attention mechanism according to claim 7 is characterized in that: In step S1, the cloth defect image and defect label data are obtained, including the minimum bounding box of the complete defect target, and the specific location and defect category of the defect are marked. All information is integrated in a json file; by using matplotlib and seaborn drawing tools, the annotation data of each minimum bounding box is parsed, the category of each defect and the corresponding position coordinates are known, and typical cloth defect types are screened out.
9. The cloth defect detection method based on multi-scale fusion attention mechanism according to claim 8, characterized in that: Typical types of cloth defects include knots, holes, three threads, coarse warp, coarse weft, and starch spots.
10. A system based on the cloth defect detection method based on multi-scale fusion attention mechanism according to claim 1, characterized in that: It includes input unit, improved YOLOv7 model unit and output unit, where: The input unit is used to input the cloth image to be detected; The improved YOLOv7 model unit detects and identifies the cloth image to be detected through the trained improved YOLOv7 model to obtain cloth defects; The output unit is used to output cloth defects.
Citation Information
Cited By
High-precision and high-adaptability intelligent detection and diagnosis method and system for electronic product defects
CN120673172A