A textile surface defect detection method and related equipment
Through the improved Yolov7 network and Transformer prediction head, the problems of low mechanization and missed detection of traditional textile defect detection are solved, and efficient textile surface defect detection is achieved.
Patent Information
- Application Number
- CN202310476352.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-28
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2043-04-28
AI Technical Summary
Traditional textile defect detection relies on artificial vision, and has low mechanization, low efficiency, high false detection and missed detection rates, making it difficult to meet the needs of complex fabric patterns and small defect detection.
The improved Yolov7 network is adopted, including input module, backbone module, head module, detection module and Transformer prediction head. Through feature extraction, fusion and small defect detection, combined with the K-means clustering algorithm and global attention module, the detection accuracy of small defects of textile fabrics is improved.
It improves the accuracy and performance of textile surface defect detection, shortens detection time, improves detection efficiency, and realizes automated defect identification and positioning.
Smart Images

Figure CN116385426B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of textile detection, and in particular to a textile surface defect detection method and related equipment. Background Art
[0002] With the rapid development of the textile industry, people are increasingly strict in controlling the quality of fabrics. In the textile industry, various adverse factors such as human error, machine failure, yarn breakage, etc. can easily cause fabric defects and affect product quality, thereby causing huge economic losses to the company. Therefore, fabric defect detection is one of the most important inspection items in textile inspection.
[0003] Defects come in a variety of types and shapes, and traditional inspection processes rely primarily on the human eye. Currently, the vast majority of textile companies still rely on manual visual inspection, where workers visually observe and judge fabric quality based on their own experience. This process suffers from low mechanization, slow manual inspection speed, and subjective factors that influence workers. This leads to false detections and missed detections, often failing to ensure precision and accuracy, and resulting in low efficiency. Defects with complex textures, varied patterns, and subtle color differences are particularly difficult for the human eye to detect, making this process far from meeting the needs of industrial production applications.
[0004] The presence of defects has a decisive impact on the quality and price of textile end products. If defective products are used in aviation, military, and medical applications, they will cause immeasurable and irreparable losses. Therefore, fabric defect detection is particularly important. However, due to the complex texture structures of various fabrics and the high similarity between noise and subtle defects, defect detection is greatly complicated.
[0005] Traditional fabric defect detection methods are generally unsupervised methods based on segmentation. Segmentation-based image defect detection technology relies on image quality and contrast, requiring a significant difference between the defect and the fabric texture background. In other words, traditional detection methods only achieve good detection results when the characteristics of the background and the defect are significantly different. However, due to the wide variety of fabric types and the continuous improvement of production technology, the types of defects are increasing and the defects are getting smaller. The human eye cannot easily distinguish between the defect and the background. In this case, traditional segmentation methods are prone to detecting patterned image parts as defects, resulting in a high false detection rate, missing small target defects, and failing to detect all fabric image defects. Summary of the Invention
[0006] The present invention provides a method for detecting surface defects of textiles and related equipment, the purpose of which is to improve the detection accuracy and performance of textile defects.
[0007] In order to achieve the above object, the present invention provides a method for detecting surface defects of textiles, comprising:
[0008] Step 1: Collect a test set of textile surface defect images;
[0009] Step 2: Input the surface defect image test set into the improved Yolov7 network for training to obtain the optimal model for target detection;
[0010] Step 3: Input the surface image of the textile to be inspected into the optimal target detection model to perform surface defect detection and obtain the inspection result;
[0011] The improved Yolov7 network consists of an input module, a backbone module for feature extraction, a head module for feature fusion, a detection module for detecting small defects, and a prediction Transformer head for defect category prediction.
[0012] The output end of the input module is connected to the input end of the backbone module, the first output end of the backbone module is connected to the first input end of the detection module, the second output end of the backbone module is connected to the input end of the head module, the output end of the head module is connected to the second input end of the detection module, the output end of the detection module is connected to the input end of the head module, and the output end of the detection module and the output end of the head module are both connected to the Transformer prediction head;
[0013] The input module serves as the input end of the improved Yolov7 network, and the output end of the Transformer prediction head serves as the output end of the improved Yolov7 network.
[0014] Specifically, the backbone modules include:
[0015] The first CBS submodule, the second CBS submodule, the third CBS submodule, the fourth CBS submodule, the first high-efficiency layer aggregation network submodule, the first MPC submodule consisting of a maximum pooling layer and a convolutional layer, the second high-efficiency layer aggregation network submodule, the second MPC submodule consisting of a maximum pooling layer and a convolutional layer, the third high-efficiency layer aggregation network submodule, the third MPC submodule consisting of a maximum pooling layer and a convolutional layer, and the fourth high-efficiency layer aggregation network submodule connected in sequence;
[0016] The input end of the first CBS sub-module is connected to the output end of the input module, the output end of the first high-efficiency layer aggregation network sub-module is connected to the first input end of the detection module, the output end of the second high-efficiency layer aggregation network sub-module is connected to the first input end of the head module, the output end of the third high-efficiency layer aggregation network sub-module is connected to the second input end of the head module, and the output end of the fourth high-efficiency layer aggregation network sub-module is connected to the third input end of the head module.
[0017] More specifically, the header module includes:
[0018] The fifth CBS submodule, the sixth CBS submodule, the seventh CBS submodule, the eighth CBS submodule, the first splicing submodule, the second splicing submodule, the third splicing submodule, the fourth splicing submodule, the first upsampling submodule, the second upsampling submodule, the fifth high-efficiency layer aggregation network submodule, the sixth high-efficiency layer aggregation network submodule, the seventh high-efficiency layer aggregation network submodule, the eighth high-efficiency layer aggregation network submodule, the fourth MPC submodule consisting of a maximum pooling layer and a convolutional layer, the fifth MPC submodule consisting of a maximum pooling layer and a convolutional layer, the first convolution submodule, the second convolution submodule, the third convolution submodule and the spatial pyramid pooling submodule;
[0019] The input end of the fifth CBS submodule is connected to the output end of the second high-efficiency layer aggregation network submodule, the output end of the fifth CBS submodule is connected to the input end of the first splicing submodule, the input end of the seventh CBS submodule is connected to the output end of the third high-efficiency layer aggregation network submodule, the output end of the seventh CBS submodule is connected to the input end of the third splicing submodule, the input end of the spatial pyramid pooling submodule is connected to the output end of the fourth high-efficiency layer aggregation network submodule, and the output end of the pyramid pooling submodule is respectively connected to the input end of the eighth CBS submodule and the input end of the fourth splicing submodule;
[0020] The output end of the eighth CBS submodule is connected to the input end of the second upsampling submodule, the output end of the second upsampling submodule is connected to the third splicing submodule, the fifth high-efficiency layer aggregation network submodule, the sixth CBS submodule, the first upsampling submodule, and the first splicing submodule in sequence, and the output end of the first splicing submodule is connected to the second input end of the detection module;
[0021] The first output end of the detection module is connected to the input end of the sixth high-efficiency layer aggregation network submodule, the output end of the sixth high-efficiency layer aggregation network submodule is respectively connected to the input end of the first convolution submodule and the input end of the fourth MPC submodule, the output end of the fourth MPC submodule and the output end of the sixth CBS submodule are both connected to the input end of the second splicing submodule, the output end of the second splicing submodule is connected to the input end of the seventh high-efficiency layer aggregation network submodule, the output end of the seventh high-efficiency layer aggregation network submodule is respectively connected to the input end of the second convolution submodule and the input end of the fifth MPC submodule, the output end of the fifth MPC submodule is connected to the input end of the fourth splicing submodule, the output end of the fourth splicing submodule is connected to the input end of the eighth high-efficiency layer aggregation network submodule, the output end of the eighth high-efficiency layer aggregation network submodule is connected to the input end of the third convolution submodule, and the output end of the third convolution submodule, the output end of the second convolution submodule, and the output end of the first convolution submodule are all connected to the input end of the Transformer prediction head.
[0022] More specifically, the detection module includes:
[0023] The ninth CBS submodule, the tenth CBS submodule, the fifth splicing submodule, the sixth splicing submodule, the third upsampling submodule, the ninth high-efficiency layer aggregation network submodule, the tenth high-efficiency layer aggregation network submodule, the sixth MPC submodule consisting of a maximum pooling layer and a convolutional layer, and the fourth convolution submodule;
[0024] The input end of the ninth CBS submodule is connected to the output end of the first high-efficiency layer aggregation network submodule, the output end of the ninth CBS submodule and the output end of the third upsampling submodule are both connected to the input end of the fifth splicing submodule, the output end of the fifth splicing submodule is connected to the input end of the ninth high-efficiency layer aggregation network submodule, the output end of the ninth high-efficiency layer aggregation network submodule is respectively connected to the input end of the fourth convolution submodule and the input end of the sixth MPC submodule, the output end of the fourth convolution submodule is connected to the input end of the Transformer prediction head, the output end of the sixth MPC submodule is connected to the input end of the sixth splicing submodule, the output end of the sixth splicing submodule is connected to the input end of the tenth high-efficiency layer aggregation network submodule, the input end of the tenth high-efficiency layer aggregation network submodule is connected to the output end of the first splicing submodule, the output end of the tenth high-efficiency layer aggregation network submodule is connected to the input end of the tenth CBS submodule, and the output end of the tenth CBS submodule is respectively connected to the input end of the third upsampling submodule and the input end of the sixth splicing submodule.
[0025] More specifically, the Transformer prediction head consists of:
[0026] The first Transformer prediction head, the second Transformer prediction head, the third Transformer prediction head, and the fourth Transformer prediction head;
[0027] The input of the first Transformer prediction head is connected to the output of the third convolutional submodule;
[0028] The input of the second Transformer prediction head is connected to the output of the second convolutional submodule;
[0029] The input of the third Transformer prediction head is connected to the output of the first convolutional submodule;
[0030] The input of the fourth Transformer prediction head is connected to the output of the fourth convolution submodule.
[0031] Furthermore, the improved Yolov7 network also includes: a first global attention module, a second global attention module, a third global attention module, a fourth global attention module and a fifth global attention module;
[0032] The input end of the first global attention module is connected to the output end of the fourth efficient layer aggregation network submodule, and the output end of the first global attention module is connected to the input end of the spatial pyramid pooling submodule;
[0033] The input end of the second global attention module is connected to the output end of the second efficient layer aggregation network submodule, and the output end of the second global attention module is connected to the input end of the fifth CBS submodule;
[0034] The input end of the third global attention module is connected to the output end of the third efficient layer aggregation network sub-module, and the output end of the third global attention module is connected to the input end of the seventh CBS sub-module;
[0035] The input end of the fourth global attention module is connected to the input end of the spatial pyramid pooling sub-module, and the output end of the fourth global attention module is connected to the input end of the eighth CBS sub-module;
[0036] The input end of the fifth global attention module is connected to the output end of the first efficient layer aggregation network sub-module, and the output end of the fifth global attention module is connected to the input end of the ninth CBS sub-module.
[0037] Furthermore, the K-means clustering algorithm is used in the optimal model for object detection to obtain the anchor boxes of the surface image samples;
[0038] A small target detection frame and a slender defect detection layer are added to the anchor frame to detect the surface image.
[0039] The present invention also provides a textile surface defect detection device, comprising:
[0040] An acquisition module, used for acquiring a test set of textile surface defect images;
[0041] The training module is used to input the surface defect image test set into the improved Yolov7 network for training to obtain the optimal model for target detection;
[0042] The detection module is used to input the surface image of the textile to be detected into the target detection optimal model to perform surface defect detection and obtain the detection result;
[0043] The improved Yolov7 network consists of an input module, a backbone module for feature extraction, a head module for feature fusion, a detection module for detecting small defects, and a prediction Transformer head for defect category prediction.
[0044] The output end of the input module is connected to the input end of the backbone module, the first output end of the backbone module is connected to the first input end of the detection module, the second output end of the backbone module is connected to the input end of the head module, the output end of the head module is connected to the second input end of the detection module, the output end of the detection module is connected to the input end of the head module, and the output end of the detection module and the output end of the head module are both connected to the Transformer prediction head;
[0045] The input module serves as the input end of the improved Yolov7 network, and the output end of the Transformer prediction head serves as the output end of the improved Yolov7 network.
[0046] The present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the method for detecting surface defects of textiles is implemented.
[0047] The present invention also provides a terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, a method for detecting surface defects of textiles is implemented.
[0048] The above solution of the present invention has the following beneficial effects:
[0049] The present invention collects a test set of surface defect images of textiles; inputs the surface defect image test set into an improved Yolov7 network for training to obtain an optimal target detection model; the optimal target detection model includes: an input module, a backbone module for feature extraction, a head module for feature fusion, a detection module for detecting small defects and a prediction Transformer prediction head for defect category prediction; the surface image of the textile to be detected is input into the optimal target detection model for surface defect detection to obtain a detection result; compared with the existing technology, the Transformer prediction head is used instead of the original prediction head to further explore the prediction potential of the model, and a detection module is added to make the model more sensitive to small defects of textiles, thereby improving the accuracy and performance of target detection, shortening the defect detection time of the textile surface, and improving the detection efficiency, which can meet the needs of automatic detection and defect identification and positioning of production line input images.
[0050] Other beneficial effects of the present invention will be described in detail in the subsequent specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 A schematic diagram of a flow chart of an embodiment of the present invention;
[0052] Figure 2 This is a structural diagram of the optimal model for target detection in an embodiment of the present invention;
[0053] Figure 3 is a schematic diagram of mixing and combining image features in an embodiment of the present invention;
[0054] Figure 4 This is a structural diagram of the Transformer prediction head in an embodiment of the present invention. DETAILED DESCRIPTION
[0055] To make the technical problems, technical solutions, and advantages to be solved by the present invention more clear, the following is a detailed description with reference to the accompanying drawings and specific embodiments. It is obvious that the embodiments described are only some of the embodiments of the present invention, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0056] In the description of the present invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings and are intended solely to facilitate and simplify the description of the present invention. They are not intended to indicate or imply that the devices or components referred to must have, be constructed, or operate in a specific orientation, and therefore should not be construed as limitations on the present invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0057] In the description of the present invention, it should be noted that, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood broadly. For example, they may refer to a locking connection, a detachable connection, or an integral connection; they may refer to a mechanical connection or an electrical connection; they may refer to a direct connection or an indirect connection through an intermediate medium; and they may refer to internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on the specific circumstances.
[0058] In addition, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0059] The present invention aims to solve the existing problems and provides a method and related equipment for detecting surface defects of textiles. Figure 1 As shown, an embodiment of the present invention provides a method for detecting surface defects of textiles, comprising:
[0060] Step 1: Collect a test set of textile surface defect images;
[0061] Step 2: Input the surface defect image test set into the improved Yolov7 network for training to obtain the optimal model for target detection;
[0062] Step 3: Input the surface image of the textile to be inspected into the optimal target detection model to perform surface defect detection and obtain the inspection result;
[0063] like Figure 2 As shown in the figure, the improved Yolov7 network includes: an input module, a backbone module for feature extraction, a head module for feature fusion, a detection module for detecting small defects, and a prediction Transformer prediction head for defect category prediction;
[0064] The output end of the input module is connected to the input end of the backbone module, the first output end of the backbone module is connected to the first input end of the detection module, the second output end of the backbone module is connected to the input end of the head module, the output end of the head module is connected to the second input end of the detection module, the output end of the detection module is connected to the input end of the head module, and the output end of the detection module and the output end of the head module are both connected to the Transformer prediction head;
[0065] The input module serves as the input end of the improved Yolov7 network, and the output end of the Transformer prediction head serves as the output end of the improved Yolov7 network.
[0066] Specifically, step 1 includes:
[0067] 300 images of each type of surface defect of textiles were collected, totaling 3000 images. In order to improve the performance of the model, data augmentation was used to increase the number of defect images of each type to 3000. The dataset was enhanced and labeled for each type of defect image to obtain a surface defect image test set. The dataset was enhanced by rotation, forging defect images and scaling, random brightness and darkness, Mixup, etc. The data labeling can be processed online using the target detection and labeling tool Labelimg.
[0068] Specifically, the traditional Yolov7 network includes a head module and a backbone module. The backbone module is used to extract features. The head module is a network layer that mixes and combines image features and passes the image features to the prediction layer to predict the image features, generate bounding boxes and predict categories, such as Figure 3As shown, the hybrid and combined image features adopt the PaFPN structure. The 32-fold downsampled feature map C5, the final output of the backbone module, is then passed through the Spatial Pyramid Pooling Cross Stage Partial (SPPCSP). The number of channels is changed from 1024 to 512. It is first fused with C4 and C3 from top to bottom to obtain P3, P4 and P5; then it is fused with P4 and P5 from bottom to top. That is, the features extracted by the backbone network are subjected to multi-scale fusion processing, which can fuse the semantic information of high and low layers, make full use of shallow information, and better detect small targets. Finally, the image features are passed to the prediction layer to predict the image features, generate bounding boxes and predict categories.
[0069] Specifically, the surface defect image test set is scaled to a preset size and then input into the improved Yolov7 network for training, which includes:
[0070] 1. Calculate the required content of the loss function value Loss of the improved Yolov7 network. The loss function value consists of three parts: classification loss, localization loss, which is used to characterize the error between the predicted box and the true box, and confidence loss, which is used to characterize the target of the box:
[0071] LOSS = target confidence loss * 0.1 + category confidence loss * 0.125 + coordinate regression loss * 0.05
[0072] The target confidence loss and category confidence loss are calculated using the binary cross entropy loss function (BCE), which is defined as:
[0073]
[0074] in, Indicates the probability that the model predicts a certain sample, and y(i) represents the sample label (the label value is 0, 1);
[0075] The coordinate regression loss is calculated using CIOU Loss, which is defined as follows:
[0076]
[0077]
[0078] Among them, IoU is an indicator used to measure the degree of overlap between the predicted box and the real box, ρ is the center distance between the predicted box and the real box, the diagonal length of the minimum enclosing rectangle of the predicted box and the real box, and the aspect ratio similarity of the predicted box and the target box is an influencing factor of the aspect ratio similarity.
[0079] 2. The matching process of positive samples in the surface defect image test set;
[0080] (1) For each real frame, roughly match the prior frame and feature points by coordinates and width and height;
[0081] (2) Use adaptive multi-positive sample matching to accurately select how many prior boxes each true box corresponds to. The steps are as follows:
[0082] a. Calculate the degree of overlap between each real frame and the current feature point prediction frame;
[0083] b. Calculate the intersection over union (IOU) of the 20 predicted boxes with the highest overlap with the real box and add up the k value of each real box. The k value represents that each real box has k feature points corresponding to it.
[0084] c. Calculate the category prediction accuracy of each true frame and the current feature point prediction frame;
[0085] d. Calculate the Cost matrix:
[0086]
[0087] Where λ is the balance coefficient, is the category loss, is the regression loss;
[0088] e. Take the k points with the lowest cost as the positive samples of the real box;
[0089] 3. Calculate the loss function value Loss.
[0090] In order to speed up the training efficiency of the model, the number of positive samples is increased. During training, each real box can be predicted by multiple prior boxes. For each real box, the IOU and type are calculated based on the predicted box adjusted by the prior box to obtain the cost, and then the prior box that best suits the real box is found.
[0091] According to the loss function, the hyperparameters of the improved Yolov7 network are continuously adjusted until the requirements are met and the optimal model for target detection is obtained.
[0092] Specifically, the backbone modules include:
[0093] The first CBS submodule, the second CBS submodule, the third CBS submodule, the fourth CBS submodule, the first efficient layer aggregation network submodule, the first MPC submodule composed of a maximum pooling layer and a convolutional layer, the second efficient layer aggregation network submodule, the second MPC submodule composed of a maximum pooling layer and a convolutional layer, the third efficient layer aggregation network submodule, the third MPC submodule composed of a maximum pooling layer and a convolutional layer, and the fourth efficient layer aggregation network submodule are connected in sequence; wherein, a CBS submodule consists of a convolutional layer Conv+a batch normalization layer BN+activation function SiLU, and the efficient layer aggregation network is ELAN (Efficient Long-range Attention Network).
[0094] The input end of the first CBS sub-module is connected to the output end of the input module, the output end of the first high-efficiency layer aggregation network sub-module is connected to the first input end of the detection module, the output end of the second high-efficiency layer aggregation network sub-module is connected to the first input end of the head module, the output end of the third high-efficiency layer aggregation network sub-module is connected to the second input end of the head module, and the output end of the fourth high-efficiency layer aggregation network sub-module is connected to the third input end of the head module.
[0095] Specifically, the header module includes:
[0096] The fifth CBS submodule, the sixth CBS submodule, the seventh CBS submodule, the eighth CBS submodule, the first splicing submodule cat, the second splicing submodule cat, the third splicing submodule cat, the fourth splicing submodule cat, the first upsampling submodule, the second upsampling submodule, the fifth high-efficiency layer aggregation network submodule, the sixth high-efficiency layer aggregation network submodule, the seventh high-efficiency layer aggregation network submodule, the eighth high-efficiency layer aggregation network submodule, the fourth MPC submodule composed of the maximum pooling layer and the convolution layer, the fifth MPC submodule composed of the maximum pooling layer and the convolution layer, the first convolution submodule, the second convolution submodule, the third convolution submodule and the spatial pyramid pooling submodule.
[0097] The input end of the fifth CBS submodule is connected to the output end of the second high-efficiency layer aggregation network submodule, the output end of the fifth CBS submodule is connected to the input end of the first splicing submodule cat, the input end of the seventh CBS submodule is connected to the output end of the third high-efficiency layer aggregation network submodule, the output end of the seventh CBS submodule is connected to the input end of the third splicing submodule cat, the input end of the spatial pyramid pooling submodule is connected to the output end of the fourth high-efficiency layer aggregation network submodule, and the output end of the pyramid pooling submodule is respectively connected to the input end of the eighth CBS submodule and the input end of the fourth splicing submodule cat;
[0098] The output end of the eighth CBS submodule is connected to the input end of the second upsampling submodule, the output end of the second upsampling submodule is connected to the third splicing submodule cat, the fifth high-efficiency layer aggregation network submodule, the sixth CBS submodule, the first upsampling submodule, and the first splicing submodule cat in sequence, and the output end of the first splicing submodule cat is connected to the second input end of the detection module;
[0099] The first output end of the detection module is connected to the input end of the sixth high-efficiency layer aggregation network submodule, the output end of the sixth high-efficiency layer aggregation network submodule is respectively connected to the input end of the first convolution submodule and the input end of the fourth MPC submodule, the output end of the fourth MPC submodule and the output end of the sixth CBS submodule are both connected to the input end of the second splicing submodule cat, the output end of the second splicing submodule cat is connected to the input end of the seventh high-efficiency layer aggregation network submodule, the output end of the seventh high-efficiency layer aggregation network submodule is respectively connected to the input end of the second convolution submodule and the input end of the fifth MPC submodule, the output end of the fifth MPC submodule is connected to the input end of the fourth splicing submodule cat, the output end of the fourth splicing submodule cat is connected to the input end of the eighth high-efficiency layer aggregation network submodule, the output end of the eighth high-efficiency layer aggregation network submodule is connected to the input end of the third convolution submodule, the output end of the third convolution submodule, the output end of the second convolution submodule, and the output end of the first convolution submodule are all connected to the input end of the Transformer prediction head.
[0100] Specifically, the detection module includes:
[0101] The ninth CBS submodule, the tenth CBS submodule, the fifth splicing submodule cat, the sixth splicing submodule cat, the third upsampling submodule, the ninth high-efficiency layer aggregation network submodule, the tenth high-efficiency layer aggregation network submodule, the sixth MPC submodule consisting of a maximum pooling layer and a convolutional layer, and the fourth convolution submodule;
[0102] The input end of the ninth CBS submodule is connected to the output end of the first high-efficiency layer aggregation network submodule, the output end of the ninth CBS submodule and the output end of the third upsampling submodule are both connected to the input end of the fifth splicing submodule cat, the output end of the fifth splicing submodule cat is connected to the input end of the ninth high-efficiency layer aggregation network submodule, the output end of the ninth high-efficiency layer aggregation network submodule is respectively connected to the input end of the fourth convolution submodule and the input end of the sixth MPC submodule, the output end of the fourth convolution submodule is connected to the input end of the Transformer prediction head, the output end of the sixth MPC submodule is connected to the input end of the sixth splicing submodule cat, and the sixth splicing submodule ca The output end of t is connected to the input end of the tenth high-efficiency layer aggregation network sub-module, the input end of the tenth high-efficiency layer aggregation network sub-module is connected to the output end of the first splicing sub-module cat, the output end of the tenth high-efficiency layer aggregation network sub-module is connected to the input end of the tenth CBS sub-module, and the output end of the tenth CBS sub-module is respectively connected to the input end of the third upsampling sub-module and the input end of the sixth splicing sub-module cat; since low-level high-resolution feature maps are more sensitive to tiny objects, a detection module is added in the embodiment of the present invention to detect tiny defects; the output end of the detection module is also connected to a Transformer detection head, and maintains the same processing method as the head module.
[0103] Specifically, the fourth MPC submodule, the fifth MPC submodule, and the sixth MPC submodule have one more forward feedback compared to the first MPC submodule, the second MPC submodule, and the third MPC submodule;
[0104] The number of outputs selected by the second branches of the first high-efficiency layer aggregation network submodule, the second high-efficiency layer aggregation network submodule, the third high-efficiency layer aggregation network submodule, and the fourth high-efficiency layer aggregation network submodule are all two;
[0105] The number of outputs selected from the second branches of the fifth high-efficiency layer aggregation network submodule, the sixth high-efficiency layer aggregation network submodule, the seventh high-efficiency layer aggregation network submodule, the eighth high-efficiency layer aggregation network submodule, the ninth high-efficiency layer aggregation network submodule, and the tenth high-efficiency layer aggregation network submodule are all 4.
[0106] Specifically, the Transformer prediction head includes:
[0107] The first Transformer prediction head, the second Transformer prediction head, the third Transformer prediction head, and the fourth Transformer prediction head;
[0108] The input of the first Transformer prediction head is connected to the output of the third convolutional submodule;
[0109] The input of the second Transformer prediction head is connected to the output of the second convolutional submodule;
[0110] The input of the third Transformer prediction head is connected to the output of the first convolutional submodule;
[0111] The input of the fourth Transformer prediction head is connected to the output of the fourth convolution submodule.
[0112] like Figure 4 As shown, in this embodiment of the present invention, a Transformer prediction head is used to replace the original prediction head. The Transformer architecture consists of two main blocks: a multi-head attention block and a feedforward neural network. Within the Transformer prediction head, the sequence-generated EmbeddedPatches first pass through a normalization layer laynorm, then undergo a multi-head self-attention operation, and then pass through a normalization layer layNorm before being output through a multi-layer perceptron (MLP). The LayerNorm layer and the Dropout layer help the network converge better and prevent overfitting. The multi-head attention block can help the current node not only focus on the current pixel, but also obtain the semantics of the context. Based on Yolov7, this embodiment of the present invention only applies the Transformer encoder to the head module to form the Transformer prediction head. It is applied at the end of the model because the feature map resolution at the end is low. Applying the Transformer prediction head on the low-resolution feature map can reduce computational and memory costs.
[0113] Specifically, the improved Yolov7 network also includes: a first global attention module, a second global attention module, a third global attention module, a fourth global attention module and a fifth global attention module;
[0114] The input end of the first global attention module is connected to the output end of the fourth efficient layer aggregation network submodule, and the output end of the first global attention module is connected to the input end of the spatial pyramid pooling submodule;
[0115] The input end of the second global attention module is connected to the output end of the second efficient layer aggregation network submodule, and the output end of the second global attention module is connected to the input end of the fifth CBS submodule;
[0116] The input end of the third global attention module is connected to the output end of the third efficient layer aggregation network sub-module, and the output end of the third global attention module is connected to the input end of the seventh CBS sub-module;
[0117] The input end of the fourth global attention module is connected to the input end of the spatial pyramid pooling sub-module, and the output end of the fourth global attention module is connected to the input end of the eighth CBS sub-module;
[0118] The input end of the fifth global attention module is connected to the output end of the first efficient layer aggregation network sub-module, and the output end of the fifth global attention module is connected to the input end of the ninth CBS sub-module.
[0119] In the improved Yolov7 network of the embodiment of the present invention, a GAMAttention global attention module is added at the connection between the head module and the backbone module and after the spatial pyramid pooling submodule. The GAMAttention global attention module uses a channel attention mechanism and a spatial attention mechanism. The processing of CAM is as follows: for the input feature map, the dimension conversion is first performed, and the dimension-converted feature map is input to the MLP (Multi-layer Perceptron) multi-layer perceptron, and then converted to the original dimension for Sigmoid processing and output;
[0120] Convolution processing is mainly used for SAM processing. The number of channels is first reduced and then increased. First, the number of channels is reduced by convolution with a convolution kernel of 7 to reduce the amount of calculation. Then, a convolution operation with a convolution kernel of 7 is performed to increase the number of channels and keep the number of channels consistent.
[0121] Specifically, in the optimal model of target detection in an embodiment of the present invention, the K-means clustering algorithm is used to obtain the anchor boxes of the surface image samples and add them to the initial anchors setting; a small target detection box and a slender defect detection layer are added to the anchor boxes, and a total of four layers [74, 75, 76, 77] are used to detect the surface image.
[0122] Specifically, the model evaluation criteria of the embodiment of the present invention are the mean average precision (mAP), precision, and recall. The superiority of the method proposed in the embodiment of the present invention is verified by calculating the mean average precision (mAP), precision, and recall. The specific calculation process is as follows:
[0123] The calculation steps for the average accuracy mAP of all categories are:
[0124] 1) Filter prediction boxes with low confidence
[0125] First, iterate over each ground-truth box object in the image, then read the detection box of this category detected by the algorithm detector, and then filter out the boxes whose confidence scores are lower than the confidence threshold;
[0126] 2) Calculate IoU
[0127] Sort the remaining detection boxes by confidence score from high to low. First, determine whether the IoU between the detection box with the highest confidence score and the ground truth box is greater than the IoU threshold. If the IoU is greater than the set IoU threshold, it is judged as TP and the ground truth box is marked as detected. Subsequent redundant detection boxes of the same ground truth box are considered FP. If the IoU is less than the IoU threshold, it is directly judged as FP.
[0128] 3) Calculate AP (Average Precision)
[0129] Based on the TP and FP obtained in step 2, combined with formulas 5 and 6, we can calculate Precision and Recall. We plot Precision as the ordinate and Recall as the abscissa to obtain the precision / recall curve, abbreviated as PR curve. The area (integral) under the PR curve is the AP.
[0130] The average AP value of all classes is taken to get mAP (meanAverage Precision).
[0131] Accuracy calculation formula:
[0132]
[0133] Recall calculation formula:
[0134]
[0135] Among them, TP: the number of samples correctly classified as positive; they are actually positive samples and are also classified as positive samples by the model; FP: the number of samples incorrectly classified as positive; they are actually negative samples, but are classified as positive samples by the model; TN: the number of samples correctly classified as negative; they are actually negative samples and are also classified as negative samples by the model; FN: the number of samples incorrectly classified as negative; they are actually positive samples, but are classified as negative samples by the model.
[0136] The embodiment of the present invention can also deploy the optimal target detection model to a local computer system. The local work computer loads all online cameras, converts the image format to JPEG through a program, and calls the detection model. When a defective image is detected, the defect detection result is marked in the original input image and displayed on the display screen. The worker is automatically paused to prompt the worker to decide whether to continue. At the same time, the detected image is automatically saved locally and automatically uploaded to the cloud (a server with a specified IP address) at a customizable time to expand the cloud sample library for automatic model learning. When there are no defects, the image is saved randomly and the next image is automatically detected. The model is self-updated and the upgrade accuracy threshold can be customized. The current model accuracy is marked. The cloud model initiates an accuracy comparison at a customizable time. When the accuracy of the cloud model accuracy comparison mark exceeds the set threshold, the existing local model is automatically replaced. At the same time, images collected from the production line are uploaded to the cloud server, and the training samples are expanded in real time. The model can continuously self-learn, continuously improve the accuracy of defect detection, and continuously cover new defect types. It supports any image input device and does not require camera calibration, saving costs and debugging costs. The model is self-trained and automatically upgraded in the cloud without human intervention.
[0137] The present invention collects a test set of surface defect images of textiles; inputs the surface defect image test set into an improved Yolov7 network for training to obtain an optimal target detection model; the improved Yolov7 network includes: an input module, a backbone module for feature extraction, a head module for feature fusion, a detection module for detecting small defects and a prediction Transformer prediction head for defect category prediction; the surface image of the textile to be detected is input into the optimal target detection model for surface defect detection to obtain a detection result; compared with the existing technology, the Transformer prediction head is used to replace the original prediction head to further explore the prediction potential of the model, and a detection module is added to make the model more sensitive to small defects of textiles, and multiple fabric detection points can be detected at the same time, and multi-size target detection output detection results are output, thereby improving the accuracy and performance of target detection, shortening the defect detection time of the textile surface, and improving the detection efficiency, and being able to meet the needs of automatic detection and defect identification and positioning of production line input images.
[0138] An embodiment of the present invention further provides a textile surface defect detection device, comprising:
[0139] An acquisition module, used for acquiring a test set of surface images of textiles;
[0140] The training module is used to input the surface image test set into the improved Yolov7 network for training to obtain the optimal model for object detection;
[0141] The detection module is used to input the surface image of the textile to be detected into the target detection optimal model to perform surface defect detection and obtain the detection result;
[0142] The improved Yolov7 network consists of an input module, a backbone module for feature extraction, a head module for feature fusion, a detection module for detecting small defects, and a prediction Transformer head for defect category prediction.
[0143] The output end of the input module is connected to the input end of the backbone module, the first output end of the backbone module is connected to the first input end of the detection module, the second output end of the backbone module is connected to the input end of the head module, the output end of the head module is connected to the second input end of the detection module, the output end of the detection module is connected to the input end of the head module, and the output end of the detection module and the output end of the head module are both connected to the Transformer prediction head;
[0144] The input module serves as the input end of the improved Yolov7 network, and the output end of the Transformer prediction head serves as the output end of the improved Yolov7 network.
[0145] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiments of the embodiments of the present invention. Their specific functions and technical effects can be found in the method embodiment section and will not be repeated here.
[0146] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of the embodiments of the present invention. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0147] An embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, a method for detecting surface defects of textile fabrics is implemented.
[0148] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the embodiments of the present invention implement all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can implement the steps of the above-mentioned various method embodiments. Among them, the computer program includes computer program code, which can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium can at least include: any entity or device capable of carrying computer program code to a construction device / terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electric carrier signal, a telecommunication signal and a software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electric carrier signals and telecommunication signals.
[0149] An embodiment of the present invention further provides a terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, a method for detecting surface defects of textiles is implemented.
[0150] It should be noted that the terminal device may be a mobile phone, tablet computer, laptop computer, ultra-mobile personal computer (UMPC), netbook, personal digital assistant (PDA), and other terminal devices. For example, the terminal device may be a station (ST, STAION) in a WLAN, a cellular phone, a cordless phone, a Session Initiation Protocol (SIP) phone, a wireless local loop (WLL) station, a personal digital assistant (PDA), a handheld device with wireless communication capabilities, a computing device or other processing device connected to a wireless modem, a computer, a laptop computer, a handheld communication device, a handheld computing device, a satellite wireless device, etc. The embodiments of the present invention do not impose any restrictions on the specific type of the terminal device.
[0151] The processor may be a central processing unit (CPU), other general-purpose processors, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.
[0152] In some embodiments, the memory may be an internal storage unit of the terminal device, such as a hard disk or memory of the terminal device. In other embodiments, the memory may also be an external storage device of the terminal device, such as a plug-in hard disk equipped on the terminal device, a smart memory card (SMC, Smart Media Card), a secure digital (SD, Secure Digital) card, a flash card, etc. Furthermore, the memory may include both an internal storage unit of the terminal device and an external storage device. The memory is used to store an operating system, an application program, a boot loader (BootLoader), data, and other programs, such as the program code of the computer program. The memory may also be used to temporarily store data that has been output or is to be output.
[0153] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiments of the embodiments of the present invention. Their specific functions and technical effects can be found in the method embodiment section and will not be repeated here.
[0154] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.
Claims
1. A method for detecting surface defects of textiles, characterized in that: include: Step 1: Collect a test set of textile surface defect images; Step 2: Input the surface defect image test set into the improved Yolov7 network for training to obtain the optimal model for target detection; Step 3: Inputting the surface image of the textile to be inspected into the target detection optimal model to perform surface defect detection and obtain a detection result; The improved Yolov7 network includes: an input module, a backbone module for feature extraction, a head module for feature fusion, a detection module for detecting small defects, and a Transformer prediction head for defect category prediction; The output end of the input module is connected to the input end of the backbone module, the first output end of the backbone module is connected to the first input end of the detection module, the second output end of the backbone module is connected to the input end of the head module, the output end of the head module is connected to the second input end of the detection module, the output end of the detection module is connected to the input end of the head module, and the output end of the detection module and the output end of the head module are both connected to the Transformer prediction head; The input module serves as the input end of the improved Yolov7 network, and the output end of the Transformer prediction head serves as the output end of the improved Yolov7 network.
2. The method for detecting textile surface defects according to claim 1, wherein: The backbone module includes: The first CBS submodule, the second CBS submodule, the third CBS submodule, the fourth CBS submodule, the first high-efficiency layer aggregation network submodule, the first MPC submodule consisting of a maximum pooling layer and a convolutional layer, the second high-efficiency layer aggregation network submodule, the second MPC submodule consisting of a maximum pooling layer and a convolutional layer, the third high-efficiency layer aggregation network submodule, the third MPC submodule consisting of a maximum pooling layer and a convolutional layer, and the fourth high-efficiency layer aggregation network submodule connected in sequence; The input end of the first CBS sub-module is connected to the output end of the input module, the output end of the first high-efficiency layer aggregation network sub-module is connected to the first input end of the detection module, the output end of the second high-efficiency layer aggregation network sub-module is connected to the first input end of the head module, the output end of the third high-efficiency layer aggregation network sub-module is connected to the second input end of the head module, and the output end of the fourth high-efficiency layer aggregation network sub-module is connected to the third input end of the head module.
3. The method for detecting textile surface defects according to claim 2, wherein: The head module includes: The fifth CBS submodule, the sixth CBS submodule, the seventh CBS submodule, the eighth CBS submodule, the first splicing submodule, the second splicing submodule, the third splicing submodule, the fourth splicing submodule, the first upsampling submodule, the second upsampling submodule, the fifth high-efficiency layer aggregation network submodule, the sixth high-efficiency layer aggregation network submodule, the seventh high-efficiency layer aggregation network submodule, the eighth high-efficiency layer aggregation network submodule, the fourth MPC submodule consisting of a maximum pooling layer and a convolutional layer, the fifth MPC submodule consisting of a maximum pooling layer and a convolutional layer, the first convolution submodule, the second convolution submodule, the third convolution submodule and the spatial pyramid pooling submodule; The input end of the fifth CBS submodule is connected to the output end of the second high-efficiency layer aggregation network submodule, the output end of the fifth CBS submodule is connected to the input end of the first splicing submodule, the input end of the seventh CBS submodule is connected to the output end of the third high-efficiency layer aggregation network submodule, the output end of the seventh CBS submodule is connected to the input end of the third splicing submodule, the input end of the spatial pyramid pooling submodule is connected to the output end of the fourth high-efficiency layer aggregation network submodule, and the output end of the pyramid pooling submodule is respectively connected to the input end of the eighth CBS submodule and the input end of the fourth splicing submodule; The output end of the eighth CBS submodule is connected to the input end of the second upsampling submodule, the output end of the second upsampling submodule is connected to the third splicing submodule, the fifth high-efficiency layer aggregation network submodule, the sixth CBS submodule, the first upsampling submodule, and the first splicing submodule in sequence, and the output end of the first splicing submodule is connected to the second input end of the detection module; The first output end of the detection module is connected to the input end of the sixth high-efficiency layer aggregation network submodule, the output end of the sixth high-efficiency layer aggregation network submodule is respectively connected to the input end of the first convolution submodule and the input end of the fourth MPC submodule, the output end of the fourth MPC submodule and the output end of the sixth CBS submodule are both connected to the input end of the second splicing submodule, the output end of the second splicing submodule is connected to the input end of the seventh high-efficiency layer aggregation network submodule, the output end of the seventh high-efficiency layer aggregation network submodule is respectively connected to the input end of the second convolution submodule and the input end of the fifth MPC submodule, the output end of the fifth MPC submodule is connected to the input end of the fourth splicing submodule, the output end of the fourth splicing submodule is connected to the input end of the eighth high-efficiency layer aggregation network submodule, the output end of the eighth high-efficiency layer aggregation network submodule is connected to the input end of the third convolution submodule, and the output end of the third convolution submodule, the output end of the second convolution submodule, and the output end of the first convolution submodule are all connected to the input end of the Transformer prediction head.
4. The method for detecting textile surface defects according to claim 3, wherein: The detection module includes: The ninth CBS submodule, the tenth CBS submodule, the fifth splicing submodule, the sixth splicing submodule, the third upsampling submodule, the ninth high-efficiency layer aggregation network submodule, the tenth high-efficiency layer aggregation network submodule, the sixth MPC submodule consisting of a maximum pooling layer and a convolutional layer, and the fourth convolution submodule; The input end of the ninth CBS submodule is connected to the output end of the first high-efficiency layer aggregation network submodule, the output end of the ninth CBS submodule and the output end of the third upsampling submodule are both connected to the input end of the fifth splicing submodule, the output end of the fifth splicing submodule is connected to the input end of the ninth high-efficiency layer aggregation network submodule, the output end of the ninth high-efficiency layer aggregation network submodule is respectively connected to the input end of the fourth convolution submodule and the input end of the sixth MPC submodule, the output end of the fourth convolution submodule is connected to the input end of the Transformer prediction head, the output end of the sixth MPC submodule is connected to the input end of the sixth splicing submodule, the output end of the sixth splicing submodule is connected to the input end of the tenth high-efficiency layer aggregation network submodule, the input end of the tenth high-efficiency layer aggregation network submodule is connected to the output end of the first splicing submodule, the output end of the tenth high-efficiency layer aggregation network submodule is connected to the input end of the tenth CBS submodule, and the output end of the tenth CBS submodule is respectively connected to the input end of the third upsampling submodule and the input end of the sixth splicing submodule.
5. The method for detecting textile surface defects according to claim 4, characterized in that: The Transformer prediction head includes: The first Transformer prediction head, the second Transformer prediction head, the third Transformer prediction head, and the fourth Transformer prediction head; The input end of the first Transformer prediction head is connected to the output end of the third convolution submodule; The input end of the second Transformer prediction head is connected to the output end of the second convolution submodule; The input end of the third Transformer prediction head is connected to the output end of the first convolution submodule; An input end of the fourth Transformer prediction head is connected to an output end of the fourth convolution submodule.
6. The method for detecting textile surface defects according to claim 4, wherein: The improved Yolov7 network further includes: a first global attention module, a second global attention module, a third global attention module, a fourth global attention module and a fifth global attention module; The input end of the first global attention module is connected to the output end of the fourth efficient layer aggregation network submodule, and the output end of the first global attention module is connected to the input end of the spatial pyramid pooling submodule; The input end of the second global attention module is connected to the output end of the second efficient layer aggregation network sub-module, and the output end of the second global attention module is connected to the input end of the fifth CBS sub-module; The input end of the third global attention module is connected to the output end of the third efficient layer aggregation network submodule, and the output end of the third global attention module is connected to the input end of the seventh CBS submodule; The input end of the fourth global attention module is connected to the input end of the spatial pyramid pooling sub-module, and the output end of the fourth global attention module is connected to the input end of the eighth CBS sub-module; The input end of the fifth global attention module is connected to the output end of the first efficient layer aggregation network sub-module, and the output end of the fifth global attention module is connected to the input end of the ninth CBS sub-module.
7. The method for detecting textile surface defects according to claim 1, wherein: Using a K-means clustering algorithm in the improved Yolov7 network to obtain an anchor frame of the surface image sample; A small target detection frame and a slender defect detection layer are added to the anchor frame to detect the surface image.
8. A textile surface defect detection device, characterized in that: include: An acquisition module, used for acquiring a test set of textile surface defect images; A training module is used to input the surface defect image test set into the improved Yolov7 network for training to obtain an optimal model for target detection; a detection module, configured to input the surface image of the textile to be detected into the target detection optimal model to perform surface defect detection and obtain a detection result; The improved Yolov7 network includes: an input module, a backbone module for feature extraction, a head module for feature fusion, a detection module for detecting small defects, and a prediction Transformer prediction head for defect category prediction; The output end of the input module is connected to the input end of the backbone module, the first output end of the backbone module is connected to the first input end of the detection module, the second output end of the backbone module is connected to the input end of the head module, the output end of the head module is connected to the second input end of the detection module, the output end of the detection module is connected to the input end of the head module, and the output end of the detection module and the output end of the head module are both connected to the Transformer prediction head; The input module serves as the input end of the improved Yolov7 network, and the output end of the Transformer prediction head serves as the output end of the improved Yolov7 network.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for detecting surface defects of textiles according to any one of claims 1 to 7 is implemented.
10. A terminal device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the textile surface defect detection method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Lightning arrester surface defect detection method based on YOLOv7
CN115953408A
Insulator defect detection method and device, medium and program product
CN115984226A