Fiber yarn variety identification method and device
By constructing a lightweight YOLO-FSL network model, the problem of deploying the tube yarn variety recognition model on edge equipment is solved, and efficient tube yarn variety recognition is achieved.
Patent Information
- Application Number
- CN202510424955.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-08-26
AI Technical Summary
The existing cylindrical yarn tube variety identification model has complex structure and high computing resource requirements, making it difficult to deploy on edge devices.
The lightweight feature extraction modules C3k2-FG and HS-FPN network are used as neck networks and SRD detection heads based on weight sharing mechanisms to build a YOLO-FSL network model and compress it through sparse training, model pruning and fine-tuning.
While maintaining high accuracy, it significantly reduces the amount of parameters, calculation amount and model size, realizes edge equipment deployment, and improves the efficiency of identifying yarn varieties.
Smart Images

Figure CN120544170A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of cheese yarn variety identification, and in particular to a fiber yarn variety identification method and device. Background Art
[0002] Cheese is the product of the winding process in textile mills. By rewinding the bobbins into conical or cylindrical tubes, efficient yarn storage and management are achieved. Cheese plays a connecting role in the spinning and weaving processes, directly affecting the quality of downstream processes such as weaving and knitting. Currently, textile companies mainly rely on manual identification and sorting of cheese types based on the color, pattern, yarn color, and label content of the yarn tubes. However, this manual identification method is costly and inefficient, and may lead to the mixing of cheeses from different batches, thereby affecting the subsequent weaving process and even resulting in high compensation costs.
[0003] Because different batches of yarn may have the same color, relying solely on yarn color for identification has limited application scenarios. Label content identification is difficult for the human eye to observe, and in the event of equipment failure, it is difficult for employees to conduct timely rechecks. Therefore, designing an accurate and efficient bobbin variety recognition model that can automatically complete variety identification and replace manual operations has become the key to the intelligent upgrade and competitiveness of textile companies.
[0004] With the continuous development of deep learning technology, deep learning-based object detection technology has been widely applied across various industries. In the textile industry, object detection technology is primarily used for textile quality inspection, especially in automated inspection, replacing traditional human-eye recognition methods. However, research on bobbin recognition remains relatively scarce. Currently, deep learning-based object detection algorithms are mainly divided into single-stage detection algorithms, such as the YOLO series and SSD; two-stage object detection algorithms, such as the R-CNN series; and Transformer-based object detection algorithms, such as the DETR series. Single-stage object detection algorithms often have high real-time performance and simple model structures, making them more suitable for bobbin type recognition. Zhang et al. proposed a deep learning-based multi-label recognition model, YoloColor-Net, which uses a lightweight DSConv module to replace standard convolutions in the backbone layer to reduce the number of parameters. They also added an improved attention mechanism (ICBAM) to the backbone and neck layers to improve bobbin recognition accuracy. While reducing both parameter and computational complexity, the model achieves a bobbin detection accuracy of 99.3%. Xu et al., based on the AlexNet model, used 3×3 convolutional kernels throughout all convolutional layers, chaining multiple kernels in series, to extract more abstract, high-level features of the object. They also integrated sliding averages and L2 regularization to improve generalization. This model boasts high recognition rate and detection efficiency. Huang et al. replaced the backbone network of the original Faster R-CNN model with ResNet50, achieving a mean average performance (MAP) of 99.9% on a self-built bobbin dataset. Dai et al., addressing the poor adaptability of traditional photoelectric and visual empty tube detection methods, proposed a lightweight empty tube detection algorithm based on YOLOv8, achieving a detection speed of 223 FPS, meeting the requirements of industrial deployment and real-time detection. Hu et al. proposed a lightweight textile defect detection model based on the CSPNet attention mechanism, demonstrating a balance between performance and deployment cost. Wang et al. improved YOLOv5 using a bidirectional feature fusion pyramid and the CA attention mechanism, achieving a 2.3% improvement in accuracy over the baseline network on a self-built dataset. Su et al. designed a balanced group softmax module and an importance-based sample weighting module based on the Faster-RCNN framework, which improved the detection accuracy and outperformed other models in the digital printed fabric defect detection dataset.
[0005] Currently, the cheese yarn bobbin variety recognition model has achieved a high level of accuracy, but its structure is often complex, making it difficult to achieve a balance between detection performance and lightweightness. It also requires high computing resources and occupies large amounts of storage and memory, making it difficult to deploy on edge devices. Summary of the Invention
[0006] The purpose of this application is to overcome the problem in the existing technology that it cannot be deployed on edge devices due to the complex structure of the recognition model, high computing resource requirements, and large storage and memory usage, and to provide a fiber yarn variety identification method and device.
[0007] In a first aspect, a method for identifying fiber yarn varieties is provided, comprising:
[0008] Obtain annotated images of cheese yarn and construct a dataset;
[0009] Constructing the YOLO-FSL network model includes the following steps: using the YOLOv11n model as the base model, using the lightweight feature extraction module C3k2-FG to replace the original C3K2 module, using the HS-FPN network as the neck network, and constructing an SRD detection head based on the weight sharing mechanism;
[0010] Using the data set to train the YOLO-FSL network model to obtain a yarn variety recognition model;
[0011] compressing the variety identification model to obtain a lightweight variety identification model;
[0012] The lightweight variety recognition model is used to identify yarn varieties.
[0013] In some possible implementations, the construction of the lightweight feature extraction module C3k2-FG includes: based on the Fasterblock residual structure composed of partial convolution PConv and standard convolution SConv, constructing a residual module FGblock by introducing a convolutional gated linear unit mechanism, and replacing the original Bottleneck structure with the residual module FGblock to obtain the lightweight feature extraction module C3k2-FG.
[0014] In some possible implementations, the HS-FPN network includes:
[0015] Feature selection module, which uses channel attention module to select input feature map Processing, performing global average pooling and global maximum pooling in parallel, and fusing the resulting features, using the Sigmoid activation function to generate channel attention weights , to achieve adaptive selection of channel dimension, where C, H and W represent the number of channels, the height and width of the feature map respectively;
[0016] Feature fusion module, which uses high-level features as weights to filter the necessary semantic information contained in low-scale features. Up-sampled by 3×3 transposed convolution with a stride of 2 to obtain ; Low-level features Dimension alignment is performed by bilinear interpolation to obtain The CA module converts high-level features into attention weights and filters low-level features. The filtered low-scale features are fused with high-level features to enhance the feature expression of the model. The formula for the fusion process of filtered low-scale features and high-level features is as follows:
[0017]
[0018]
[0019] Among them, f b is the high-level feature of the input; T conv represents a 3×3 transposed convolution with a stride of 2, which is used to upsample high-level features; BL represents bilinear interpolation, which is used to assist in feature size alignment; f att is the high-level feature after upsampling box alignment; f s is the low-level feature of the input; CA is the Channel Attention module, which is used to generate attention weights; f out is the feature of the final fusion output.
[0020] In some possible implementations, the SRD detection head consists of three groups of detection heads, each group uses P3, P4 and P5 level features for target recognition, and processes the received P3, P4 and P5 level input features respectively using 3x3 reparameterized convolution. The features are aggregated using the Conv_GN module with 3x3 reparameterized convolution and 1x1 convolution kernel, and the features extracted by shared convolution are input into the classification head and regression head, and the Scale layer is added to scale the features at different levels.
[0021] In some possible implementations, the variety recognition model is compressed to obtain a lightweight variety recognition model, including: sparse training, model pruning, and model fine-tuning.
[0022] In some possible implementations, the sparse training includes normalizing the input of each neuron in each layer of the neural network in each batch. The normalization formula is:
[0023]
[0024]
[0025] in, represents the input of the BN layer, and represents the output of the BN layer, Represents the average value of the activation function input on the BN layer, Represents the standard deviation of the activation function input on the BN layer, is the scaling factor, is the deviation factor; is the data after batch normalization, is a small positive number;
[0026] The L1 norm regularization constraint of the scaling factor γ is introduced into the objective function. The network weights and the scaling factor are updated synchronously through a joint optimization strategy. The gradient penalty generated by the constraint term forces the γ parameter to shrink toward the origin of the coordinate system to complete the channel sparse operation. The loss function formula for model sparse training is:
[0027]
[0028]
[0029] in, represents the loss term of network training, Indicates the sparse processing of L1 regularization on the scaling factor, x is the input of the training, y is the target of the training, are the training weights in the network, represents the penalty term for sparse training, is the sparsity regularization coefficient.
[0030] In some possible implementations, the model pruning includes: removing channels whose γ is close to 0 that are screened out after sparse training; and the model fine-tuning includes: adaptively adjusting the weight distribution of retained parameters so that the remaining network can effectively compensate for the feature expression capabilities of the removed channels.
[0031] In a second aspect, a fiber yarn variety identification device is provided, comprising:
[0032] The dataset construction module is used to obtain the annotated images containing cheese yarn and then construct the dataset;
[0033] The model construction module is used to build the YOLO-FSL network model. The specific steps include: taking the YOLOv11n model as the base model, using the lightweight feature extraction module C3k2-FG to replace the original C3K2 module, using the HS-FPN network as the neck network, and building an SRD detection head based on the weight sharing mechanism;
[0034] A training module, configured to train the YOLO-FSL network model using the data set to obtain a yarn variety recognition model;
[0035] A model compression module, used for compressing the variety identification model to obtain a lightweight variety identification model;
[0036] The identification module is used to identify the yarn variety using the lightweight variety identification model.
[0037] In a third aspect, a computer-readable storage medium is provided, wherein the computer-readable medium stores program code for execution by a device, the program code including steps for executing the method in any one of the implementations of the first aspect.
[0038] In a fourth aspect, an electronic device is provided, comprising a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements a method as in any one of the implementations of the first aspect described above.
[0039] The present application has the following beneficial effects: In the present application, a lightweight feature extraction module C3k2-FG, an HSFPN network as a neck network and an SRD detection head based on a weight sharing mechanism are used to construct a lightweight YOLO-FSL network model. After training, compression and fine-tuning, a lightweight variety recognition model that can identify cheese yarn varieties is obtained. While maintaining a high-precision model, the amount of parameters, calculation amount and model size are significantly reduced, thereby enabling edge device deployment and improving the recognition efficiency of cheese yarn varieties. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] The drawings that constitute a part of this application are used to provide a further understanding of this application. The illustrative embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation on this application.
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0042] Figure 1 This is a flow chart of the fiber yarn variety identification method of Example 1 of the present application;
[0043] Figure 2 This is a technical roadmap of the fiber yarn variety identification method of Example 1 of the present application;
[0044] Figure 3 2 is a schematic structural diagram of a cheese yarn image acquisition device in the fiber yarn variety identification method of Example 1 of the present application;
[0045] Figure 4 Schematic diagram of different varieties of cheese yarns in the fiber yarn variety identification method of Example 1 of the present application;
[0046] Figure 5 Schematic diagram of the YOLO-FSL network structure in the fiber yarn variety identification method of Example 1 of the present application;
[0047] Figure 6 Schematic diagram of the structure of the C3k2-FG module in the fiber yarn variety identification method of Example 1 of the present application;
[0048] Figure 7 1 is a schematic diagram of the PConv structure in the fiber yarn variety identification method of Example 1 of the present application;
[0049] Figure 8 1 is a schematic diagram of the structure of the CGLU module in the fiber yarn variety identification method of Example 1 of the present application;
[0050] Figure 9 Schematic diagram of the CA module knot in the fiber yarn variety identification method of Example 1 of the present application;
[0051] Figure 10(a) is the detection head architecture diagram of YOLOv11;
[0052] FIG10( b ) is a diagram of the SRD detection head structure in the fiber yarn variety identification method of Example 1 of the present application;
[0053] Figure 11 This is a compression flow chart of a yarn variety identification model in the fiber yarn variety identification method of Example 1 of the present application;
[0054] Figure 12 This is a schematic diagram of channel pruning in the fiber yarn variety identification method of Example 1 of the present application;
[0055] FIG13( a ) is a diagram showing the BN distribution change when λ is 0.01 in the fiber yarn variety identification method of Example 1 of the present application;
[0056] FIG13( b ) is a diagram showing the BN distribution change when λ is 0.1 in the fiber yarn variety identification method of Example 1 of the present application;
[0057] FIG13( c ) is a diagram showing the BN distribution change when λ is 0.2 in the fiber yarn variety identification method of Example 1 of the present application;
[0058] FIG13( d ) is a diagram showing the BN distribution change when λ is 0.3 in the fiber yarn variety identification method of Example 1 of the present application;
[0059] Figure 14 This is a comparison diagram of channels before and after pruning in the fiber yarn variety identification method of Example 1 of the present application;
[0060] Figure 15This is a structural block diagram of a fiber yarn variety identification device according to Example 2 of the present application;
[0061] Figure 16 This is a schematic diagram of the internal structure of the electronic device of Example 4 of the present application.
[0062] Reference numerals:
[0063] 100. Dataset construction module; 200. Model construction module; 300. Training module; 400. Model compression module; 500. Recognition module; 601. Conveyor; 602. Darkroom; 603. Bracket; 604. Color industrial camera; 605. Lens; 606. Light source; 607. Package yarn; 608. Computer. DETAILED DESCRIPTION
[0064] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0065] Example 1
[0066] like Figure 1 As shown, a fiber yarn variety identification method involved in Example 1 of the present application includes:
[0067] S100, obtaining annotated images containing cheese yarn and constructing a data set;
[0068] In the research of lightweight cheese yarn bobbin variety detection model, this embodiment mainly focuses on four aspects: data set acquisition, model lightweight design, model compression and model performance evaluation. The technical route of this embodiment is as follows: Figure 2 shown.
[0069] Figure 3 A homemade cheese yarn image acquisition lab platform is demonstrated, consisting of a color industrial camera 604, lens 605, light source 606, bracket 603, darkroom 602, cheese yarn 607, conveyor 601, and computer 608. Light source 606 is an LED, providing sufficient lighting for the color industrial camera. The offset between light source 606 and camera 604 does not obstruct the camera's image capture. The color industrial camera 604 used in the experiment is a Daheng MARS-1840-63GTC model, and images are saved in JPG format. The color industrial camera 604 is mounted directly above the cheese yarn conveyor belt to capture images of the top surface of the cheese yarn 607 and upload them to computer 608 for processing.
[0070] Construct a rich variety of cheese yarn dataset. The image data are all from a textile company in Shaoxing, Zhejiang Province, China. Considering the wide variety of cheese yarns, but the limited batches of cheese yarns in a textile factory at a single time, the data collection is divided into two phases. The first phase collected 6 categories in October 2024; the second phase collected 9 categories in December 2024. Figure 4 Figure 2 shows schematic diagrams of different types of cheese yarn. Due to the reuse of cheese tubes and the squeezing during the stacking process, some cheese tubes are damaged. Based on actual production conditions, these damaged images are also retained in the dataset.
[0071] In this example, the labelimg annotation tool was used to annotate all cheese yarn images. During the annotation process, the smallest outer rectangle of each cheese yarn tube was used as the labeled area, and the remaining unlabeled parts were considered as the background. To ensure the accuracy and consistency of the annotation, all images were independently completed by two annotators and cross-checked. If any annotation discrepancies were found, the annotators would discuss and ultimately reach a consensus. After the annotation was completed, the 3,000 images were divided into a training set, a validation set, and a test set in a 6:2:2 ratio. Table 1 shows the detailed composition of the cheese yarn image detection dataset.
[0072] Table 1: Dataset composition
[0073]
[0074] S200, constructing a YOLO-FSL network model, specifically including the following steps: using the YOLOv11n model as the basic model, adopting the lightweight feature extraction module C3k2-FG to replace the original C3K2 module, using the HS-FPN network as the neck network, and constructing an SRD detection head based on the weight sharing mechanism;
[0075] In visual inspection tasks based on deep learning, the size and complexity of the model directly affect the effect of practical applications. The YOLOv11n model was selected as the basic model because of its high precision, limited parameters, low computational requirements and compact model size, but it also has rich redundant functions to ensure the model's comprehensive interpretation of the input image. In order to ensure that the model maintains high precision while reducing structural redundancy, a new cheese yarn tube recognition model named YOLO-FSL is proposed in this embodiment. Specifically, a lightweight feature extraction module (C3k2-FG) is used to replace the original C3k2 module; the HSFPN neck network is introduced to enhance the yarn tube feature expression ability and reduce model parameters; the detection head (SRD) is redesigned to reduce the number of convolutional layers through a weight sharing mechanism. The proposed YOLO-FSL structure diagram is shown in the figure below. Figure 5 As shown in Figure 2. The size of the input image is 640×640×3.
[0076] Specifically, C3k2 is the core feature extraction module of the YOLOv11 network. Its traditional stacked convolution structure is prone to feature redundancy. This embodiment is based on the Fasterblock residual structure composed of partial convolution PConv and standard convolution SConv. By introducing the Convolutional Gated Linear Unit (CGLU) mechanism, a new residual module FGblock is designed to replace the original Bottleneck structure. The new feature extraction module is named C3k2-FG. The improvement scheme is as follows: Figure 6 shown.
[0077] PConv reduces computational complexity through a selective channel processing strategy: convolution operations are performed only on some channels of the input feature map, and the remaining channels maintain the same mapping, such as Figure 7 shown. Figure 7 middle 、 、 and The convolution kernel size, the number of input channels, the height, and the width of the feature map during PConv operations are represented by c, while the number of input channels of the feature map during SConv operations is represented by c. This operation significantly reduces memory access while ensuring consistency between input and output channels. The floating-point calculation formulas for PConv and Sconv are:
[0078] (1)
[0079] (2)
[0080] It can be seen from the formula that the floating-point calculation amount of PConv is the floating-point calculation amount of SConv. , the use of PConv greatly reduces the amount of calculation.
[0081] The residual module FGblock includes the Pconv module, the CGLU module, and the DropPath module. The CGLU module (Convolutional GLU) is a channel mixer that combines convolution operations with gated linear units (GLU). It bridges the gap between the GLU and SE mechanisms, extracts local neighborhood features through convolution operations, dynamically generates channel attention weights related to the current spatial position, suppresses irrelevant features, strengthens key information, enhances local modeling and feature extraction capabilities, and improves model robustness. The structure diagram is shown in Figure 8.
[0082] After adopting C3k2-FG design in the backbone network, compared with YOLOv11n, the model precision and recall rate are improved by 0.6% and 0.4% respectively, while the model floating-point operation amount is reduced.
[0083] Neck networks play an important role in improving model expressiveness and classification accuracy, but their high computational complexity and parameter count often lead to decreased inference speed and increased computing resource consumption. Given the resource constraints of edge computing devices, developing lightweight neck network architectures has become an important direction for model optimization. This study introduces HS-FPN as a neck network. HS-FPN is a new type of neck network proposed in 2024. Its structure, shown in Figure 8, consists of two core components: a feature selection module and a feature fusion module.
[0084] In the feature selection module, the channel attention module (CA) is used to input feature maps Processing is performed, where C, H, and W represent the number of channels, the height, and the width of the feature map, respectively. Global Average Pooling (GAP) and Global Max Pooling (GMP) are performed in parallel, and the resulting features are fused to reduce the dimension and number of feature maps, while eliminating redundant information and reducing the number of model parameters. Finally, the Sigmoid activation function is used to generate channel attention weights. , achieving adaptive selection of channel dimension.
[0085] In the feature fusion process, the SFF multi-level feature fusion mechanism uses high-level features as weights to filter the necessary semantic information contained in low-scale features. The low-level features are aligned by bilinear interpolation and up-sampled by 3×3 transposed convolution (T-Conv) with a stride of 2 to obtain Subsequently, the CA module converts high-level features into attention weights and filters low-level features. Finally, the filtered low-scale features are fused with high-level features to enhance the feature expression of the model. The feature fusion process formula is as follows:
[0086] (3)
[0087] (4)
[0088] Among them, f b is the high-level feature of the input; T conv represents a 3×3 transposed convolution with a stride of 2, which is used to upsample high-level features; BL represents bilinear interpolation, which is used to assist in feature size alignment; fatt is the high-level feature after upsampling box alignment; f s is the low-level feature of the input; CA is the Channel Attention module, which is used to generate attention weights; f out is the feature of the final fusion output.
[0089] As shown in Figure 10(a), YOLOv11's detection head uses an independent detection head branch design for multi-level feature extraction. It consists of three groups of detection heads, each of which utilizes P3, P4, and P5-level features for target recognition. However, this architecture may result in low model parameter utilization. In addition, YOLOv11 adopts a decoupled structure, processing classification and regression tasks separately. While this helps improve task focus, it also results in a large computational burden. To achieve a lightweight detection head design, this embodiment proposes a new SRD detection head based on a weight sharing mechanism, taking into account the similarity in scale and proportion of the detection targets. The structure of this detection head is shown in Figure 10(b).
[0090] First, 3x3 reparameterized convolutions are used to process the input features at the P3, P4, and P5 levels. This is because different features may require different convolution kernels to capture complex patterns, and a single shared convolution parameter may not be able to fully capture these differences, thus limiting the model's expressive power. Reparameterized convolutions can offset the negative impact of shared convolutions in lightweight designs. By introducing more learnable parameters, the network can more effectively extract features from the data, thereby mitigating the potential accuracy loss associated with lightweight models.
[0091] Next, a Conv_GN module with a 3x3 reparameterizable convolution and a 1x1 convolution kernel is used to aggregate features and reduce the interference of redundant information. GroupNorm has been shown to improve the detection head's performance in both localization and classification. Finally, the features extracted through the shared convolution are input to the classification and regression heads. To address the issue of inconsistent object scales processed by different detection heads, a Scale layer is added to scale features at different levels, further improving detection accuracy.
[0092] S300, using the data set to train the YOLO-FSL network model to obtain a yarn variety recognition model;
[0093] Specifically, the labeled dataset is divided into training set, validation set and test set according to the ratio of 6:2:2 and input into the YOLO-FSL network model to complete the training and testing of the model, thereby obtaining a yarn variety recognition model that can identify the variety of package yarn.
[0094] S400, compressing the variety identification model to obtain a lightweight variety identification model;
[0095] In this embodiment, a lightweight design strategy is adopted to significantly reduce the computational complexity and memory usage of the model. However, the problem of structural redundancy still exists. In view of this, in order to simultaneously reduce model complexity, reduce the amount of floating-point operations, and improve inference efficiency, this embodiment introduces a channel pruning algorithm based on L1 regularization to implement structured pruning on the optimized model. The basic principle of the algorithm is that the scaling factor γ in the batch normalization (BN) layer is strongly correlated with the importance of the feature channel. The larger the γ, the more important the channel is in the model. When the γ value approaches zero, the contribution of the corresponding channel to the model output is negligible, so it can be safely pruned to achieve model compression.
[0096] The channel pruning algorithm process is shown in Figure 11, which includes three core stages: sparsity training, model pruning, and fine-tuning.
[0097] The internal covariate shift (ICS) problem exists during deep neural network training, which forces higher layers to constantly adapt to changes in lower layer parameters. To address this problem, the inputs to each neuron in each layer of the neural network are normalized in each batch to follow a standard normal distribution, thus maintaining a consistent distribution of input data across each layer. The formula is:
[0098] (5)
[0099] (6)
[0100] In the above formula, represents the input of the BN layer, and represents the output of the BN layer, Represents the average value of the activation function input on the BN layer, Represents the standard deviation of the activation function input on the BN layer, is the scaling factor, is the deviation factor; is the data after batch normalization, A small positive number.
[0101] To select channels with smaller γ, an L1-norm regularization constraint on the scaling factor γ is introduced into the objective function. The network weights and scaling factor are updated synchronously through a joint optimization strategy. The gradient penalty generated by this constraint forces the γ parameter to shrink toward the origin, completing the channel sparsification operation. The loss function formula for model sparse training is:
[0102] (7)
[0103] (8)
[0104] Where, represents the loss term of network training, Indicates the sparse processing of L1 regularization on the scaling factor, x is the input of the training, y is the target of the training, are the training weights in the network, represents the penalty term for sparse training, is the sparsity regularization coefficient.
[0105] After sparse training, channels with γ close to 0 are screened out, which have low contribution to the model and can be removed. Figure 12 This paper demonstrates the principle of model channel pruning. Through pruning, the pruned network model forms a sparse and compact architecture, significantly reducing the number of model parameters and floating-point operations.
[0106] While increasing the pruning rate can significantly reduce model complexity, excessively high pruning rates can lead to accuracy degradation. Therefore, fine-tuning training is essential. Fine-tuning adaptively adjusts the weight distribution of retained parameters so that the remaining network can effectively compensate for the feature expression capabilities of the removed channels. This minimizes performance losses caused by channel pruning while maintaining model lightweight.
[0107] S500: Identify yarn varieties using the lightweight variety identification model.
[0108] After compressing the variety recognition model to obtain a lightweight variety recognition model, the lightweight variety recognition model can be deployed on the edge device to run. The picture containing the cheese yarn to be identified is input into the lightweight variety recognition model to realize automatic identification of the cheese yarn variety.
[0109] In order to verify the performance of the model, the following comparative experiments are conducted on the model, experimental configuration and parameter configuration:
[0110] To ensure the repeatability and fairness of the bobbin variety identification experiment, all comparative experiments in this example were completed on a unified hardware platform (see Table 2 for details). The experimental benchmark model uses the standard version of the YOLOv11n architecture, and performance comparisons are performed under the same training strategy. During model training, the global number of iterations is set to 300 epochs, each epoch contains 29 batches, and each batch inputs 64 images (batch size = 64). For the optimization strategy, the stochastic gradient descent (SGD) optimizer was selected based on its parameter adjustment flexibility and training efficiency advantages. The learning rate scheduling adopts a dynamic attenuation mechanism (ReduceLROnPlateau), and the initial value is set to = 0.01.
[0111] Table 2: Equipment hardware configuration and software environment
[0112]
[0113] In order to systematically evaluate the comprehensive performance of the YOLO-FSLP model, this experiment uses precision (P), recall (R), and mean average precision (mAP@0.5) as model performance evaluation indicators. Their mathematical definitions are shown in Equations (9)-(12):
[0114] (9)
[0115] (10)
[0116] (11)
[0117] (12)
[0118] In the above formula, TP represents the number of correctly detected positive samples; FP refers to the number of negative samples misclassified as positive; and FN represents the number of undetected positive samples. AP reflects the accuracy of single-category detection, while mAP achieves a comprehensive multi-category evaluation by taking the arithmetic average of the AP values across all categories.
[0119] In addition, we introduce indicators to evaluate model complexity and real-time performance, mainly including parameter count (Params), floating-point operations (FLOPs), model size (model size), and frames per second (FPS). The formulas for calculating Params and FLOPs in the convolutional layer are as follows:
[0120] (13)
[0121] (14)
[0122] Where: and is the convolution kernel size; represents the feature map input, Indicates the number of output channels; and is the spatial dimension of the feature map.
[0123] Comparative experiments of different neck networks:
[0124] In order to illustrate the advantages of the HSFPN neck network model, this experiment compares the effects of different neck networks on the recognition of bobbin varieties. Six different types of neck networks, PAFPN, BIFPN, CGAF, CGAF, MAFPN and HSFPN, are selected. The experimental results are shown in Table 3.
[0125] Experimental results show that the HSFPN neck network demonstrates significant advantages across multiple metrics. In terms of detection performance, HSFPN achieved a precision (P) of 99.8% and a recall (R) of 100%, tied for best with BIFPN. Its mAP50 of 99.5% was on par with models like PAFPN and BIFPN. Notably, while maintaining high-precision features, HSFPN reduced its parameter count (1.81M) by 6.2% compared to the next-best BIFPN (1.93M) and its computational load (5.5 GFLOPs) by 12.7% compared to PAFPN (6.3 GFLOPs), demonstrating superior lightweight performance. In terms of real-time performance, HSFPN achieved an inference speed of 90 FPS (batch size = 1), a 28.6%-38.5% improvement over other networks. The model size was compressed to 3.8MB, a 5% reduction compared to BIFPN. In summary, HSFPN effectively reduces the model complexity and significantly improves the inference efficiency by optimizing the bobbin feature selection and multi-scale feature fusion while ensuring the detection accuracy.
[0126] Table 3: Comparative experiments of different neck networks
[0127]
[0128] Comparative experiments of different models:
[0129] To further verify the superiority of the algorithm in this paper, this example compares YOLO-FSLP with other algorithms in the YOLO series, including YOLOv12n, YOLOv11n, YOLOv8n, YOLOv7-tiny, and other excellent algorithms.
[0130] The results in Table 4 show that in terms of detection accuracy, the lightweight variety recognition model of this embodiment achieved a precision (P) of 99.0% and a recall (R) of 99.1%, comparable to the DETR model (P = 99.6% and R = 99.8%). YOLO-FSLP demonstrates significant advantages in terms of lightweightness: the model parameters are only 0.34MB, an 86.7% reduction compared to the next-best YOLOv12n and approximately 0.5% of the DETR model; computational complexity is reduced to 1.8GB, a 77.8% reduction compared to YOLOv8n; and the model size is compressed to 1.3MB, only 24.5% of YOLOv12n and one to two orders of magnitude smaller than traditional detection models (e.g., SSD: 35MB, Faster R-CNN: 165MB). In terms of real-time performance, the model achieves a detection speed (FPS) of 88 frames per second, a 15.8% improvement over YOLOv7-tiny, making it the best of all models. While maintaining a high level of detection accuracy, this model successfully solves the problems of insufficient parameter compression and excessive consumption of computing resources in existing lightweight models, providing a better solution for edge device deployment.
[0131] Table 4: Comparative experiments of different models
[0132]
[0133] Ablation experiment:
[0134] To evaluate the effectiveness of C3k2-FG, HSFPN, and SRD in the YOLO-FSL model in improving the bobbin variety recognition model, this example designed eight control experiments under the same experimental environment (hardware, environment configuration, hyperparameters, and dataset partitioning). The ablation experiment results of the YOLO-FSL model are shown in Table 5, where √ indicates that the module is enabled and × indicates that the module is not enabled. The experimental results are analyzed as follows:
[0135] Table 5: Ablation experiments
[0136]
[0137] Table 5 shows that the introduction of the C3k2-FG module improves model accuracy, with precision (P) and recall (R) increasing by 0.6% and 0.4%, respectively. The module also offers advantages in reducing computational burden, with parameter count and floating-point operations (FLOPs) decreasing by 6.6% and 6.3%, respectively. However, inference speed decreases to 66 frames per second (FPS). This indicates that the C3k2-FG module improves detection accuracy by enhancing bobbin feature extraction, but at the expense of computational efficiency. In contrast, the HSFPN module demonstrates significant model compression capabilities. When used alone, the parameter count is reduced to 1.81M, FLOPs to 5.5G, and inference speed increased by 21.6%, reaching 90 FPS. This demonstrates that the HSFPN module, with its dual advantages of lightweightness and efficiency, effectively reduces the model's computational complexity while maintaining high inference speed. After introducing the SRD detection head, the model achieved inference speed of 101 FPS while maintaining 99.5% mAP50 detection accuracy, a 36.5% improvement over the baseline model. Despite the increased inference speed, the model size increased to 6.4MB, indicating that while the SRD detection head improves performance, it also increases the model's storage requirements.
[0138] In module combination experiments, the modules demonstrated good synergy. For the combination of the C3k2-FG and HSFPN modules, the model parameters were reduced to 1.65M, and FLOPs were reduced by 19.0% to 5.1G, while maintaining accuracy, with P and R reaching 99.7% and 99.9%, respectively. Compared to single modules, the combined model was more lightweight while maintaining high accuracy. Further combining the SRD detection head with the C3k2-FG+HSFPN combination resulted in the YOLO-FSL model, which exhibited the best balance: while both P and R were improved, the parameter count, FLOPs, and model size were reduced to 1.50M, 4.4G, and 3.5MB, respectively. Furthermore, the inference speed was increased by 10 frames per second.
[0139] In summary, each module demonstrates distinct advantages both individually and in combination. The C3k2-FG module enhances feature extraction capabilities, the HSFPN module significantly reduces model size, and the SRD detection head improves inference speed. When combined with multiple modules, the model achieves a better balance between accuracy, computational efficiency, and storage requirements, making it particularly suitable for deployment on resource-constrained edge devices.
[0140] This experiment performed 500 rounds of sparse training based on the trained weights of the YOLO-FSL model. The initial learning rate for model training was set to 0.01. To investigate the impact of the regularization coefficient on model performance, the experiment first set a small regularization coefficient λ and observed the distribution of the batch normalization (BN) layer. The value of λ was then gradually increased. Specifically, the values of λ were set to 0.01, 0.1, 0.2, and 0.3. The changes in the BN distribution were monitored using TensorBoard, demonstrating the effect of sparse training under different λ values. The relevant comparison plots are shown in Figures 13(a), 13(b), 13(c), and 13(d).
[0141] As shown in Figures 13(a), 13(b), 13(c), and 13(d), as the number of training rounds increases, the γ coefficient in the batch normalization layer gradually approaches 0. At the same time, as the regularization coefficient λ increases, the rate at which the γ coefficient approaches 0 accelerates. The data in Table 6 show that the model's detection accuracy decreases as λ increases. When λ is 0.01, detection accuracy reaches its highest level, almost matching that of the original model. However, when λ increases from 0.2 to 0.3, model performance deteriorates significantly, with precision (P) and recall (R) decreasing by 2.3% and 0.8%, respectively. When λ is 0.2, the model is able to maintain high detection accuracy while improving training speed. Therefore, based on the above experimental results, 0.2 is selected as the optimal value for the regularization coefficient λ in this embodiment.
[0142] Table 6: Model performance with different λ
[0143]
[0144] Comparative experiments of different pruning rates and fine-tuning:
[0145] After sparse training, the bobbin variety recognition model needs to be pruned. In the experiment, the pruning rates were set to 20%, 40%, 50%, and 60%. Table 7 shows the model performance at different pruning rates. When the pruning rate is between 20% and 50%, the model performance decreases slightly, primarily due to the removal of unimportant channels. However, when the pruning rate is set to 60%, the pruning operation reduces the floating-point operations by 52.3% and remains unchanged. At this point, P and R represent a significant drop in model performance, indicating that the removal of key parameters makes it impossible to restore the model's accuracy through fine-tuning. To further identify the most optimal pruning rate, experiments with pruning rates of 51% and 52% were added.
[0146] Table 7 shows that as the pruning rate increases from 0% to 50%, the model's parameter count decreases from 1.50M to 0.34M, the floating-point operations decrease from 4.4G to 1.8G, and the model size is compressed to 37.1% of its original size. Notably, when the pruning rate does not exceed 50%, the model's mAP50 remains high at 99.5%, while the inference speed (FPS) increases from 84 to 88. This demonstrates that channel pruning effectively removes redundant parameters without significantly impacting the model's recognition performance. When the pruning rate increases to 51%, mAP50 experiences its first 0.7% decrease, but the model's inference speed reaches a peak of 89 FPS. Further increasing the pruning rate to 52%, precision (P) and recall (R) drop significantly by 12.9% and 29.9%, respectively, while mAP50 also drops significantly to 91.8%, indicating that the model is approaching a critical degradation point. When the pruning rate was further increased to 52.3%, the model suffered structural damage and the mAP50 dropped sharply to 67.2%. The model could not recover its performance through conventional fine-tuning.
[0147] Comprehensive experimental results show that a pruning rate of 50% is the optimal balance point for this experiment: compared with the original model, the number of model parameters is reduced by 77.3%, the computational complexity is reduced by 59.1%, while maintaining the mAP50 index of 99.5%, and the model size is compressed to 1.3MB, fully meeting the needs of lightweight deployment in industrial scenarios and improving the stability and efficiency of the model on resource-constrained devices. When the pruning rate is 50%, the channel comparison before and after pruning is as follows Figure 14 As shown, Figure 14 Medium grey represents the channel before pruning, and black represents the channel after pruning.
[0148] Table 7: Model pruning and fine-tuning experimental results
[0149]
[0150] In this embodiment, a lightweight, high-precision recognition model based on an improved YOLOv11, YOLO-FSLP, is proposed to address the problems of high model complexity and deployment difficulty in the task of identifying bobbin varieties in the textile industry. By designing the C3k2-FG module to optimize redundancy in feature extraction, introducing the HSFPN neck network to enhance the ability to express bobbin features, and re-parameterizing convolution and weight sharing mechanisms to reconstruct the detection head, the model's parameter count and computational complexity are reduced. Furthermore, the model is compressed using a channel pruning strategy, ultimately resulting in a compact, lightweight variety recognition model with a parameter count of 0.34M and a computational cost of 1.8G. While maintaining 99.5% mAP accuracy, the lightweight variety recognition model is reduced to 1.3MB in size, and the inference speed is increased by 14 frames per second. Experimental results demonstrate that this lightweight variety recognition model offers significant advantages in balancing lightweightness and accuracy, providing an effective solution for edge device deployment in industrial scenarios.
[0151] Looking ahead, further optimization of the YOLO-FSLP model can focus on the following two aspects: first, enhancing the recognition capability for more complex textile types and expanding its application scope; second, further improving the model's inference speed by integrating other deep learning techniques. Importantly, the lightweight variety recognition model of this embodiment already meets the needs of today's textile companies for cheese yarn variety recognition, providing strong support for the intelligent upgrade of textile companies.
[0152] Example 2
[0153] like Figure 15 As shown, a fiber yarn variety identification device involved in Example 2 of the present application includes:
[0154] A data set construction module 100 is used to obtain annotated pictures containing cheese yarns and then construct a data set;
[0155] The model construction module 200 is used to construct a YOLO-FSL network model, specifically comprising the following steps: using the YOLOv11n model as the base model, adopting the lightweight feature extraction module C3k2-FG to replace the original C3K2 module, using the HS-FPN network as the neck network, and constructing an SRD detection head based on a weight sharing mechanism;
[0156] A training module 300 is used to train the YOLO-FSL network model using the data set to obtain a yarn variety recognition model;
[0157] A model compression module 400 is used to compress the variety identification model to obtain a lightweight variety identification model;
[0158] The identification module 500 is used to identify the yarn variety using the lightweight variety identification model.
[0159] It should be noted that, for other specific implementations of the fiber yarn variety identification device in this embodiment, reference can be made to the specific implementations of the above-mentioned fiber yarn variety identification method, and will not be repeated here to avoid redundancy.
[0160] Example 3
[0161] A computer-readable storage medium according to embodiment 3 of the present application, wherein the computer-readable storage medium stores program code for execution by a device, the program code including steps for executing the method in any one of the implementations in embodiment 1 of the present application;
[0162] Among them, the computer-readable storage medium can be a read-only memory (ROM), a static storage device, a dynamic storage device or a random access memory (RAM); the computer-readable storage medium can store program code, and when the program stored in the computer-readable storage medium is executed by the processor, the processor is used to execute the steps of the method in any one of the implementation methods in Example 1 of the present application.
[0163] Example 4
[0164] like Figure 16 As shown, an electronic device involved in Example 4 of the present application includes a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements the method in any one of the implementations in Example 1 of the present application;
[0165] Among them, the processor can adopt a general central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), a graphics processing unit (GPU) or one or more integrated circuits to execute relevant programs to implement the method in any one of the implementation methods in Example 1 of the present application.
[0166] The processor may also be an integrated circuit electronic device with signal processing capabilities. In the implementation process, each step of the method in any one of the implementations in Example 1 of the present application may be completed by hardware integrated logic circuits in the processor or software instructions.
[0167] The above-mentioned processor can also be a general-purpose processor, a digital signal processor, an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. The various methods, steps, and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the embodiments of the present application can be directly embodied as being executed by a hardware decoding processor, or being executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium mature in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and, in combination with its hardware, completes the functions required to be executed by the units included in the data processing device of the embodiment of the present application, or executes the method in any one of the implementation modes in Example 1 of the present application.
[0168] The above are only preferred specific implementations of this application; however, the scope of protection of this application is not limited thereto. Any person skilled in the art who, within the technical scope disclosed in this application, makes equivalent substitutions or modifications based on the technical solutions and improved concepts of this application shall be covered by the scope of protection of this application.
Claims
1. A fiber yarn variety identification method, characterized in that: include: Obtain annotated images of cheese yarn and construct a dataset; Constructing the YOLO-FSL network model includes the following steps: using the YOLOv11n model as the base model, using the lightweight feature extraction module C3k2-FG to replace the original C3K2 module, using the HS-FPN network as the neck network, and constructing an SRD detection head based on the weight sharing mechanism; Using the data set to train the YOLO-FSL network model to obtain a yarn variety recognition model; compressing the variety identification model to obtain a lightweight variety identification model; The lightweight variety recognition model is used to identify yarn varieties.
2. The fiber yarn variety identification method according to claim 1, characterized in that: The construction of the lightweight feature extraction module C3k2-FG includes: based on the Fasterblock residual structure composed of partial convolution PConv and standard convolution SConv, constructing a residual module FGblock by introducing the convolution gated linear unit mechanism, and replacing the original Bottleneck structure with the residual module FGblock to obtain the lightweight feature extraction module C3k2-FG.
3. The fiber yarn variety identification method according to claim 1, characterized in that: The HS-FPN network includes: Feature selection module, which uses the channel attention module to select the input feature map Processing, performing global average pooling and global maximum pooling in parallel, and fusing the resulting features, using the Sigmoid activation function to generate channel attention weights , to achieve adaptive selection of channel dimension, where C, H and W represent the number of channels, the height and width of the feature map respectively; Feature fusion module, which uses high-level features as weights to filter the necessary semantic information contained in low-scale features. Up-sampled by 3×3 transposed convolution with a stride of 2 to obtain ; Low-level features Dimension alignment is performed by bilinear interpolation to obtain The CA module converts high-level features into attention weights and filters low-level features. The filtered low-scale features are fused with high-level features to enhance the feature expression of the model. The formula for the fusion process of filtered low-scale features and high-level features is as follows: ; ; Among them, f b is the high-level feature of the input; T conv represents a 3×3 transposed convolution with a stride of 2, which is used to upsample high-level features; BL represents bilinear interpolation, which is used to assist in feature size alignment; f att is the high-level feature after upsampling box alignment; f s is the low-level feature of the input; CA is the Channel Attention module, which is used to generate attention weights; f out is the feature of the final fusion output.
4. The fiber yarn variety identification method according to claim 1, characterized in that: The SRD detection head consists of three groups of detection heads, each of which uses P3, P4 and P5 level features for target recognition. The received P3, P4 and P5 level input features are processed by 3x3 reparameterized convolution respectively. The features are aggregated using the Conv_GN module with 3x3 reparameterized convolution and 1x1 convolution kernel. The features extracted by shared convolution are input to the classification head and regression head, and the Scale layer is added to scale the features at different levels.
5. The fiber yarn variety identification method according to claim 1, characterized in that: The variety recognition model is compressed to obtain a lightweight variety recognition model, including: sparse training, model pruning and model fine-tuning.
6. The fiber yarn variety identification method according to claim 5, characterized in that: The sparse training includes normalizing the input of each neuron in each layer of the neural network in each batch. The normalization formula is: ; ; in, represents the input of the BN layer, and represents the output of the BN layer, Represents the average value of the activation function input on the BN layer, Represents the standard deviation of the activation function input on the BN layer, is the scaling factor, is the deviation factor; is the data after batch normalization, is a small positive number; The L1 norm regularization constraint of the scaling factor γ is introduced into the objective function. The network weights and the scaling factor are updated synchronously through a joint optimization strategy. The gradient penalty generated by the constraint term forces the γ parameter to shrink toward the origin of the coordinate system to complete the channel sparse operation. The loss function formula for model sparse training is: ; ; in, represents the loss term of network training, Indicates the sparse processing of L1 regularization on the scaling factor, x is the input of the training, y is the target of the training, are the training weights in the network, represents the penalty term for sparse training, is the sparsity regularization coefficient.
7. The fiber yarn variety identification method according to claim 6, characterized in that: The model pruning includes: removing channels with γ close to 0 that are screened out after sparse training; the model fine-tuning includes: adaptively adjusting the weight distribution of retained parameters so that the remaining network can effectively compensate for the feature expression capabilities of the removed channels.
8. A fiber yarn variety identification device, characterized in that: include: The dataset construction module is used to obtain the annotated images containing cheese yarn and then construct the dataset; The model construction module is used to build the YOLO-FSL network model. The specific steps include: taking the YOLOv11n model as the base model, using the lightweight feature extraction module C3k2-FG to replace the original C3K2 module, using the HS-FPN network as the neck network, and building an SRD detection head based on the weight sharing mechanism; A training module, configured to train the YOLO-FSL network model using the data set to obtain a yarn variety recognition model; A model compression module, used for compressing the variety identification model to obtain a lightweight variety identification model; The identification module is used to identify the yarn variety using the lightweight variety identification model.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores program codes for execution by a device, wherein the program codes include steps for executing the method according to any one of claims 1 to 7.
10. An electronic device, characterized in that: The electronic device includes a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements the method according to any one of claims 1 to 7.
Citation Information
Cited By
Unmanned aerial vehicle infrared small target detection method based on PGF-RTDETR
CN121191040A