Yarn quality online detection method based on lightweight UNet
Through the MiniUNet network based on lightweight UNet, the problem of yarn quality detection in the prior art is solved, and high-precision and fast yarn quality monitoring is achieved, which is suitable for mobile device deployment.
Patent Information
- Application Number
- CN202411846979.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-16
- Publication Date
- 2025-05-16
AI Technical Summary
Existing yarn quality detection technology is difficult to detect yarn quality in real time on the spinning equipment production line, resulting in high cost and low efficiency of manual inspection, and poor detection effect of traditional machine vision solutions, and high missed detection rates and error detection rates.
The yarn quality online detection method based on lightweight UNet is adopted, and the semantic segmentation detection of yarn pictures is carried out through the MiniUNet network, and a lightweight yarn quality monitoring network is built using technical means such as cross-bar convolution module, Soft-CBAM attention mechanism and depth separation convolution.
It realizes high accuracy and fast online detection of yarn quality, with only 35.11K parameters, only 0.27MB memory usage, average transfer ratio (MIoU) can reach 84.50%, and F1 score can reach 91.24%, making it suitable for deployment on mobile inspection equipment.
Smart Images

Figure CN120013852A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of yarn detection and deep learning, and specifically relates to an online yarn quality detection method based on lightweight UNet. Background Art
[0002] Online yarn quality detection technology is mainly used to detect abnormalities in the yarn production process of the chemical fiber industry. These production abnormalities are mainly caused by equipment failure, equipment aging, manual operation errors, etc., and the nozzle area in the spinning equipment mainly needs to detect yarn quality defects such as floating yarns and abnormal silk paths. Existing yarn quality detection technologies are mainly based on capacitive, photoelectric, traditional image processing and other methods, but they are all for quality detection of yarns after production, and lack of yarn quality detection in the production line.
[0003] In addition, most yarn production is completed automatically in a 24-hour non-stop manner, with a small amount of manual operation during the period. For abnormal detection in the production line, it is mainly based on manual inspection, which has problems such as high labor cost and low detection efficiency. It is limited by subjective factors such as fatigue and lack of concentration of personnel, coupled with environmental factors such as unstable light on site, which is very easy to miss the detection phenomenon, resulting in defective products and reduced production efficiency. Therefore, the existing technology tends to adopt a combination of computer technology and mobile robot technology to reduce or even replace manual inspection. It mainly adopts a production detection method based on traditional machine vision algorithms, deploys the detection algorithm on mobile devices, obtains images of the yarn production process through industrial cameras on mobile devices, and then sends the images to the corresponding algorithm for judgment. The algorithm based on traditional machine vision mainly judges whether it is abnormal by judging the deviation distance of the yarn. This type of method obtains the yarn position through edge processing, and then judges whether there is an abnormality by the yarn position spacing. However, the background and environment of yarn production equipment are complex, and the diameter of the yarn is extremely small during the production process. The traditional edge algorithm or template matching algorithm is prone to losing and misjudging the yarn position. At the same time, traditional machine vision solutions are easily affected by the background and shooting clarity, resulting in poor detection results, high missed detection rate and false detection rate, which are not suitable for large-scale promotion.
[0004] In addition, deep learning algorithms have made achievements in many fields with their strong adaptability and excellent detection effects, especially semantic segmentation algorithms, which have powerful fine segmentation capabilities and are very suitable for image segmentation with unclear features. Therefore, it is of practical significance to apply semantic segmentation algorithms to defect segmentation in yarn production. However, the UNet semantic segmentation algorithm based on convolutional neural networks is often used for fine segmentation of images such as medical images, but it has problems such as large algorithm size, high computing power requirements, and slow detection speed, which makes it unsuitable for direct deployment on mobile inspection equipment. Summary of the invention
[0005] The technical problem to be solved by the present invention is to provide an online yarn quality detection method based on lightweight UNet, which is used to detect the yarn quality at the oil nozzle during the yarn spinning process in a lightweight and high-precision manner.
[0006] In order to solve the above technical problems, the present invention provides an online yarn quality detection method based on lightweight UNet, the process comprising: online collection of yarn pictures of the yarn production process, reducing the size of the collected pictures in a computer and inputting them into a MiniUNet network trained offline, and outputting a semantic segmentation detection result map with abnormal positions and types;
[0007] The MiniUNet network is based on the UNet network and adopts an asymmetric structure. Each downsampling module of the Encoder includes a cross-strip convolution module and a maximum pooling operation; Bottleneck includes a Soft-CBAM attention mechanism and nearest neighbor interpolation; each upsampling module of the Decoder includes a depthwise separable convolution, an activation function GELU, and a Pixelshuffle operation; a jump connection is used between the downsampling module and the upsampling module.
[0008] As an improvement of the yarn quality online detection method of a lightweight UNet of the present invention:
[0009] The Soft-CBAM attention mechanism is improved based on the CBAM attention mechanism, and the calculation process is as follows:
[0010]
[0011]
[0012]
[0013] Where: is the original feature map; is the output attention value; OutCA(F) is the channel attention output value; OutSA(F) is the spatial attention output value; σ is the Sigmoid function; W0 and W1 are the weights of MLP; and is the average pooling feature along the channel; and is the maximum pooling feature along the channel; is the soft spatial pooling feature; To pool features along the channel width; It is the pooling feature along the channel height direction.
[0014] As a further improvement of the yarn quality online detection method of a lightweight UNet of the present invention:
[0015] The calculation process of the cross-strip convolution module is:
[0016]
[0017] Among them, Conv k1×k2 It represents the convolution with kernel k1×k2, x is the input, y is the output, GELU is the nonlinear activation function based on Gaussian error function, and BN is batch normalization.
[0018] As a further improvement of the yarn quality online detection method of a lightweight UNet of the present invention:
[0019] The offline training process of the MiniUNet network is:
[0020] Collect pictures of the enterprise's yarn production process, scale the collected pictures on the computer, and then manually annotate them and divide them into training set and test set in a ratio of 0.8:0.2;
[0021] The images in the training set are input into the MiniUNet network. During the training process, the loss function value is calculated and the model parameters are optimized by back-propagation iteration. After each round of training, the test set is input to check the training effect and the corresponding F1 score is calculated. The training ends after the preset number of epochs is reached, and the best weight is retained based on the best test result.
[0022] As a further improvement of the yarn quality online detection method of a lightweight UNet of the present invention:
[0023] The loss function is:
[0024]
[0025] Among them, Loss is the total loss, is the cross entropy loss function, is the dice loss, is the prediction result, and y is the label map.
[0026] The beneficial effects of the present invention are mainly reflected in:
[0027] 1. The present invention constructs a cross-strip convolution module and applies it to the Encoder part of the UNet network, so that the Encoder is lightweight while ensuring the feature extraction capability;
[0028] 2. The present invention improves the CBAM mechanism, strengthens the performance of the attention mechanism, and applies it to the Bottleneck part of the UNet network, which simplifies the processing of high-level features and greatly reduces the number of model parameters without affecting the detection accuracy.
[0029] 3. The present invention uses deep separable convolution to reconstruct the Decoder of the UNet network, removing the redundant 3×3 ordinary convolution, making the Decoder more lightweight while maintaining the same decoding effect;
[0030] 4. The present invention designs an asymmetric structure based on UNet, and constructs a lightweight yarn quality online monitoring network MiniUNet by introducing cross-strip convolution, Soft-CBAM attention mechanism and depth-separable convolution. The network has only 35.11K parameters and 0.27MB memory usage. The average intersection over union (MIoU) can reach 84.50%, and the F1 score can reach 91.24%, which makes the model have better performance in the image semantic segmentation task of yarn quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The specific implementation modes of the present invention are further described in detail below with reference to the accompanying drawings.
[0032] Figure 1 This is a schematic diagram of the basic UNet structure;
[0033] Figure 2 It is a schematic diagram of the structure of MiniUNet of the present invention;
[0034] Figure 3 It is a schematic diagram of the structure of the cross-strip convolution module of the present invention;
[0035] Figure 4 Schematic diagram of the CBAM attention mechanism structure;
[0036] Figure 5 It is a schematic diagram of the structure of Soft-CBAM of the present invention;
[0037] Figure 6 This is a schematic diagram of Silk Road anomalies and floating silk;
[0038] Figure 7 It is a flow chart of the offline training process of the MiniUNet network of the present invention;
[0039] Figure 8 This is a schematic diagram of the detection effect of the MiniUNet of the present invention and several other mainstream algorithms on yarn defects. DETAILED DESCRIPTION
[0040] The present invention is further described below in conjunction with specific embodiments, but the protection scope of the present invention is not limited thereto:
[0041] Example 1: A method for online yarn quality detection based on lightweight UNet, the specific process is as follows:
[0042] 1. MiniUNet network
[0043] The present invention is based on the UNet network and adopts an asymmetric structure design: a cross-strip convolution module is proposed to reconstruct the encoder, which greatly reduces the network parameters and can speed up the yarn quality detection speed; at the same time, the CBAM attention mechanism is improved to process low-resolution high-level features, which simplifies the processing process of high-level features and improves the feature processing capability; finally, a deep separable convolution is introduced to reconstruct the decoder, which further reduces the network parameters.
[0044] 1.1. Build a basic UNet network
[0045] The UNet network is mainly composed of four parts: Encoder, Bottleneck, Decoder, and Skip. The UNet network structure is as follows: Figure 1 As shown in the figure. Both the Encoder and Decoder parts are composed of multiple convolutional modules (ConvBlock), each of which consists of a 3×3 ordinary convolution and a ReLU (LinearRectification Function) activation function. Every two convolutional modules of the Encoder are connected by maximum pooling, which is mainly responsible for image feature extraction; every two convolutional modules of the Decoder are connected by transposed convolution, which is mainly used to generate output results based on feature images; the output of each downsampling module of the Encoder is connected to the input of the corresponding upsampling module of the Decoder through Skip to increase the information of low-level features; Bottleneck is composed of 2 convolutional modules and is responsible for processing high-level features.
[0046] 1.2 Lightweight Encoder
[0047] In order to achieve the lightweight of the Encoder of the UNet network, the present invention proposes a cross-strip convolution module for the reconstruction of the Encoder of the basic UNet network: replace the two ordinary convolution modules in each downsampling module in the encoder with a cross-strip convolution module of the present invention. The cross-strip convolution module is a new design of the present invention. Compared with the 3×3 ordinary convolution in the original encoder, the cross-strip convolution module composed of 1×7 and 7×1 strip convolutions has a larger receptive field and stronger feature extraction capability; at the same time, the cross-strip convolution module adopts a depth-separable design, with smaller overall parameters and calculations, and is more lightweight; the specific design of the cross-strip convolution module is as follows:
[0048] First, two sets of symmetrical strip convolutions with shared parameters are used to obtain image features along the axial direction to ensure that the encoder pays equal attention to horizontal and vertical features; then the extracted results are superimposed and passed through the nonlinear activation function GELU (Gaussian Error Linear Unit) and then superimposed with the local residual; finally, the number of channels is expanded through point-by-point convolution blocks. The overall structure of the cross strip convolution module is as follows: Figure 3 As shown, the formula is as follows:
[0049]
[0050] Among them, Conv k1×k2 It represents the convolution with kernel k1×k2, x is the input, y is the output, GELU is the nonlinear activation function based on Gaussian error function, and BN is batch normalization.
[0051] The cross-strip convolution module uses GELU as the activation function instead of ReLU because ReLU is not differentiable at x=0, while GELU is differentiable at 0 and is smoother; at the same time, because GELU does not force all values less than or equal to 0 to be set to 0, in the yarn image, the defects are less obvious than the background. When ReLU is used to force negative values to zero, it is easy to lose the defect features. The GELU used in the present invention makes the optimization process more stable and can better retain information. The formulas of the two are as follows:
[0052]
[0053] GELU=&xΦ(x)=xP(X≤x),&where X~N(0,1) (3)
[0054] Where x is the input and Φ(x) is the cumulative distribution function of the standard normal distribution.
[0055] 1.3 Bottleneck structure based on Soft-CBAM attention mechanism
[0056] The traditional CBAM attention mechanism consists of channel attention and spatial attention, and the way to obtain channel attention and spatial attention is through strong maximum pooling (MaxPool) and average pooling (AvgPool), such as Figure 4 As shown, the formula is as follows:
[0057]
[0058]
[0059]
[0060] in, is the original feature map; σ is the Sigmoid function; Represents element-by-element multiplication of matrices; and is the weight of MLP; and is the average pooling feature along the channel;
[0061] and is the maximum pooling feature along the channel. is the output attention value; Out CA (F) is the channel attention output value; Out SA (F) is the spatial attention output value.
[0062] The traditional CBAM attention mechanism method may cause problems such as excessive feature loss. Therefore, this paper proposes a softer attention mechanism Soft-CBAM, such as Figure 5 As shown:
[0063] (1) For channel attention, a soft spatial pooling channel is added in parallel with the original maximum pooling channel and average pooling channel, including an FC layer (fully connected layer) and an activation function ReLU. The FC layer is used to obtain additional channel weights, compress the data of each channel to 1×1, and then pass through the activation function ReLU before superimposing it into the original weight;
[0064] (2) For spatial attention, two convolution kernels are added, 1×C and C×1 strip convolutions, which slide along the height and width of the image respectively. By multiplying the pixels at the same position on each channel with the learnable weights, additional spatial information is obtained and spliced into the results obtained by directly taking the maximum pooling and average pooling of the traditional spatial attention mechanism, where C is the number of channels.
[0065] The calculation process of the Soft-CBAM attention mechanism is as follows:
[0066]
[0067]
[0068]
[0069] Where: is the soft spatial pooling feature; To pool features along the channel width; It is the pooling feature along the channel height direction.
[0070] The two 3×3 ordinary convolution modules in the Bottleneck structure of the UNet network are replaced with the Soft-CBAM attention mechanism to simplify high-level feature processing. The design of the Soft-CBAM attention mechanism increases the way to obtain channel information, retains high-level feature information as much as possible, and greatly reduces the computational pressure in the low-resolution high-level feature processing process.
[0071] 1.4. Decoder with depthwise separable convolution
[0072] The depthwise separable convolution consists of one depthwise convolution and one pointwise convolution. The image information is processed by the depthwise convolution. The convolution operation is performed on each channel with the same number of convolution kernels as the number of input feature layers, and then the channel information is processed by pointwise convolution. The overall effect is similar to that of an ordinary convolution, but the number of parameters and the amount of calculation are greatly reduced. Therefore, it is often used in lightweight design. It is introduced in the present invention. Each upsampling module uses a depthwise separable convolution with layer normalization and GELU activation function to replace the two convolution modules in the decoder of the basic Unet network, and reconstructs it to form a lightweight Decoder.
[0073] Assuming the image size is C×H×W, the calculation formula for the parameters (Parameters, Params) and calculation amount (Floating Point Operations, FLOPs) of the depthwise separable convolution is as follows:
[0074]
[0075]
[0076] Among them, C in is the number of input channels, C out is the number of output channels, K 1 ×K 2 is the height and width of the convolution kernel, H out ×W out The height and width of the output image, g is the number of grouped convolution groups, and g is related to C in the depth convolution inSimilarly, in point-by-point convolution, the convolution kernel size is 1×1 and g is 1. It can be proved by the calculation formula 【1】 , depthwise separable convolution has lower computational complexity and parameter count than conventional convolution.
[0077] In addition, the decoder part of the lightweight UNet network (MiniUNet network) of the present invention adopts Pixelshuffle (Sub-Pixel Convolutional Neural Network) as an upsampling method to replace the original transposed convolution. Compared with transposed convolution, Pixelshuffle increases the resolution of the image by rearranging pixels, integrates multiple layers of features, can better retain the detailed information of the yarn, and avoids complex convolution calculations, reducing the amount of calculation.
[0078] In the BottleNeck part, the nearest neighbor interpolation is used to replace the original transposed convolution. The BottleNeck part uses the attention mechanism, which focuses more on extracting the important parts of the existing features. The nearest neighbor interpolation fills the target pixel with the pixel value of the nearest neighbor, which can obtain the processing results of low-resolution features more completely; in addition, the nearest neighbor interpolation is simpler to implement and has higher computational efficiency, which can effectively improve the processing speed of the model.
[0079] 1.5. In combination with steps 1.1-1.4, the present invention uses a lightweight UNet network as an online detection network for yarn quality (MiniUNet network), such as Figure 2 As shown in the figure, the MiniUNet network uses the U-shaped structure and jump connection of the UNet network, and the rest of the parts are lightweight and simplified. The Encoder includes 4 downsampling modules, and the Decoder includes 4 downsampling modules. Jump connections are used between the downsampling modules and the upsampling modules. Each downsampling module of the Encoder uses a cross-strip convolution module and a maximum pooling operation, which ensures the feature extraction capability while being lightweight; Bottleneck uses the Soft-CBAM attention mechanism and nearest neighbor interpolation to simplify the high-level feature processing process and further reduce the network parameters and calculation amount; each upsampling module of the Decoder uses a depth-separable convolution, activation function GELU and a Pixelshuffle operation, discarding the redundant convolution structure, while ensuring that the Decoder can effectively generate segmentation maps through high-level features. Finally, the constructed MiniUNet has only 35.11K parameters and 0.27MB of memory usage.
[0080] 2. Model offline training
[0081] 2.1、Dataset establishment
[0082] The dataset was collected from a chemical fiber production enterprise. The dataset images were taken by a CMOS industrial array camera to capture the nozzle area in the spinning equipment. In order to achieve real-time monitoring of yarn production, the camera was deployed on a mobile device located in the nozzle area, and the image acquisition pixel was set to 1920×1440. In the computer, the collected images were reduced to 640×480 pixel images and annotated with labelme to produce a dataset. The Silk Road anomaly was labeled as siluyichang, the floating silk was labeled as piaosi, and the rest was labeled as background. There are a total of 2863 images, including 2290 training sets and 573 test sets. The Silk Road anomaly and floating silk images are shown in the figure below. Figure 6 shown.
[0083] 2.2 Offline training and testing process
[0084] Input the labeled images in the training set into the MiniUNet network for training, such as Figure 7 As shown in the figure, the encoder first extracts features, then processes high-order features through the Soft-CBAM attention mechanism, and finally generates output results through the decoder. The model output is compared with the label map, and then the weights are continuously optimized according to the Loss function. The loss function used is:
[0085]
[0086] Among them, Loss is the total loss, is the cross entropy loss function (CE), is the Dice Loss, is the prediction result, and y is the label map.
[0087] The Adam optimizer with a learning rate of 0.0001 and a momentum of 0.9 was used for model training. The batch size (BatchSize) was 2, and the total number of training rounds was greater than or equal to 100. During the training process, the loss function value was calculated and the model parameters were optimized by back propagation. The model weights with the best training effect were saved.
[0088] The testing process is to put the labeled test set images into the trained model, generate and save the prediction results, compare the prediction results with the label images, and calculate the evaluation indicators to achieve the preset goals.
[0089] 3. Online use
[0090] Collect pictures of the nozzle area of the spinning equipment during the yarn production process, reduce the collected pictures to 640×480 pixels in the computer, and input them into the MiniUNet network that has been trained and tested offline in step 2, and output the detection result map with the abnormal location and type (including three types of silk path abnormality, floating silk and normal).
[0091] experiment
[0092] 1. The experiment uses the data set of Example 1
[0093] 2. Evaluation indicators
[0094] The experimental evaluation indicators include F1 score (F1 Score), mean intersection over union (MIoU), parameters (Params, Parameters), memory usage (MU, Memory Usage), computing amount (FLOPs, Floating Point Operations) and frame rate (FPS, Frames Per Second). The results of parameters, memory usage and computing amount are calculated by traversing the model using random inputs of size (1, 3, 640, 480); the result of frame rate is the average of three results with an error of no more than 5Hz after multiple rounds of testing while the device remains in the same operating state.
[0095] Among them, the calculation formulas for F1 score, average intersection-over-union ratio and FPS are as follows:
[0096]
[0097]
[0098]
[0099] Among them, n is the number of categories, m is the number of test set images, and IoU i is the intersection over union (IoU) of the i-th category, time j It is the total time from inputting the jth picture to being processed by the model and outputting it.
[0100] The formulas for precision, recall, and intersection-over-union are as follows:
[0101]
[0102]
[0103]
[0104] Among them, TP is the number of positive samples predicted correctly; FP is the number of positive samples incorrectly predicted as negative samples; TN is the number of negative samples predicted correctly; FN is the number of negative samples incorrectly predicted as positive samples.
[0105] 3. Comparative experiment
[0106] Under the same data set, the comparative test results of the present invention and four mainstream semantic segmentation algorithms are shown in Table 1. By analyzing Table 1, it can be seen that the MiniUNet network proposed in the present invention greatly reduces the number of parameters and memory usage of the model while ensuring the segmentation accuracy, and improves the frame rate. Among them, the number of parameters is only 35.11K, and the memory usage is only 0.27MB.
[0107] Table 1 Performance comparison of mainstream semantic segmentation algorithms
[0108]
[0109] As shown in Table 1, the extremely small model inevitably brings a certain performance loss. The F1 score of the MiniUNet network only drops by 1.52% compared to the optimal model UNet, and the average intersection-and-union ratio only drops by 2.42%. However, the MiniUNet network has the smallest parameter amount, memory usage and the highest frame rate. The parameter amount is only 35K, the memory usage is only 0.27MB, and the parameter amount and memory usage are approximately 1 / 1000 and 1 / 876 of the UNet model, respectively. In addition, compared with EGEUNet with the lowest computational amount, the F1 score of MiniUNet is almost the same, and MIoU is better, the frame rate is increased by about 80Hz, and the detection task can be completed more efficiently and in real time. Therefore, the MiniUNet of the present invention is better in comprehensive performance indicators and is more suitable for deployment on mobile devices for field detection of fabrics.
[0110] In order to intuitively demonstrate the detection effect of the algorithm of the present invention, it is proved that the improved algorithm proposed in this paper reduces a large number of parameters and memory usage, while having high efficiency close to several large-scale algorithms, and is more suitable for deployment on mobile devices with low computing power. Figure 8 An example comparison of the detection results of MiniUNet and several mainstream algorithms on the experimental data set is shown, which intuitively shows that the present invention can achieve the same or better detection effect as the existing algorithms with extremely small number of parameters and memory usage.
[0111] 4. Analysis of experimental results based on different attention mechanisms
[0112] In order to measure the impact of the Soft-CBAM attention mechanism in simplifying high-level feature processing, the F1 score, average intersection-over-union ratio, number of parameters, memory usage and amount of computation are selected as evaluation indicators, and different attention mechanisms are combined with the lightweight UNet after Bottleneck replacement. The test results are shown in Table 2, where "None" (i.e., no attention mechanism) represents the use of two ordinary convolutions.
[0113] Table 2 Results of different attention mechanisms replacing Bottleneck
[0114]
[0115] As can be seen from Table 2, using the attention mechanism for high-level feature processing can greatly reduce the number of parameters and memory usage. At the same time, the Soft-CBAM attention mechanism designed in this paper has better high-level feature processing capabilities among several attention mechanisms. Compared with the convolution with high parameter count, the F1 score is improved by 0.28% and the average intersection-over-union ratio is improved by 0.44%.
[0116] Finally, it should be noted that the above examples are only some specific embodiments of the present invention. Obviously, the present invention is not limited to the above embodiments, and there are many variations. All variations that can be directly derived or associated with the content disclosed by a person skilled in the art should be considered as the protection scope of the present invention.
[0117] 【1】Howard AG, Zhu M, Chen B, et al.MobileNets: Efficient ConvolutionalNeural Networks for Mobile Vision Applications[J].ArXiv,2017,abs / 1704.04861.
Claims
1. A yarn quality online detection method based on lightweight UNet, characterized by: The process includes online collection of yarn images during the yarn production process, reducing the size of the collected images in a computer and inputting them into the offline trained MiniUNet network, and outputting a semantic segmentation detection result map with abnormal locations and types; The MiniUNet network is based on the UNet network and adopts an asymmetric structure. Each downsampling module of the Encoder includes a cross-strip convolution module and a maximum pooling operation; Bottleneck includes a Soft-CBAM attention mechanism and nearest neighbor interpolation; each upsampling module of the Decoder includes a depthwise separable convolution, an activation function GELU, and a Pixelshuffle operation; a jump connection is used between the downsampling module and the upsampling module.
2. According to claim 1, a yarn quality online detection method based on lightweight UNet is characterized in that: The Soft-CBAM attention mechanism is improved based on the CBAM attention mechanism, and the calculation process is as follows: Where: is the original feature map; is the output attention value; Out CA (F) is the channel attention output value; Out SA (F) is the spatial attention output value; σ is the Sigmoid function; W0 and W1 are the weights of MLP; and is the average pooling feature along the channel; and is the maximum pooling feature along the channel; is the soft spatial pooling feature; To pool features along the channel width; It is the pooling feature along the channel height direction.
3. According to claim 2, a method for online detection of yarn quality based on lightweight UNet is characterized in that: The calculation process of the cross-strip convolution module is: Among them, Conv k1×k2 It represents the convolution with kernel k1×k2, x is the input, y is the output, GELU is the nonlinear activation function based on Gaussian error function, and BN is batch normalization.
4. According to claim 3, a method for online detection of yarn quality based on lightweight UNet is characterized in that: The offline training process of the MiniUNet network is: Collect pictures of the enterprise's yarn production process, scale the collected pictures on the computer, and then manually annotate them and divide them into training set and test set in a ratio of 0.8:0.2; The images in the training set are input into the MiniUNet network. During the training process, the loss function value is calculated and the model parameters are optimized by back-propagation iteration. After each round of training, the test set is input to check the training effect and the corresponding F1 score is calculated. The training ends after the preset number of epochs is reached, and the best weight is retained based on the best test result.
5. The method for online yarn quality detection based on lightweight UNet according to claim 4 is characterized in that: The loss function is: Among them, Loss is the total loss, is the cross entropy loss function, is the dice loss, is the prediction result, and y is the label map.