Visual identification method for appearance defect detection of polyester yarn disc
By using an improved YOLOv1 model, combined with feature fusion and attention mechanisms, the problem of detecting small defects on polyester yarn discs was solved, achieving efficient and accurate automated detection and reducing costs.
Patent Information
- Application Number
- CN202511755562.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-27
- Publication Date
- 2026-02-06
AI Technical Summary
Existing technologies are unable to effectively identify small defects such as thread run-in, waste thread run-in, and snagging on polyester yarn reels, resulting in low detection efficiency, poor accuracy, and high manual inspection costs.
An improved YOLOv1 model was used to construct an SP-YOLO neural network, which introduced a scale sequence feature fusion module, a triple feature encoder module, and a channel and position attention mechanism module. The network was trained by combining EIoU loss, cls loss, and dfl loss to achieve automatic identification and detection of appearance defects in polyester yarn discs.
It improves the accuracy and efficiency of detecting appearance defects in polyester yarn discs, reduces the false negative rate, adapts to the detection of yarn discs with different thicknesses and texture densities, does not require major modifications to the model architecture, and reduces industrial implementation costs.
Smart Images

Figure CN121482011A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to a visual recognition method for detecting defects in the appearance of polyester yarn discs, belonging to the technical field of industrial defect detection, and specifically relates to the application of deep learning and target detection model in the detection of defects in the appearance of polyester yarn discs. The present application constructs a SP-YOLO neural network based on an improved YOLOv11 model, introduces a scale sequence feature fusion module, a triple feature encoder module and a channel and position attention mechanism module in the feature enhancement module of the model, and realizes the automatic recognition and detection of defects in the appearance of polyester yarn discs. BACKGROUND
[0002] As a key tool for carrying and protecting polyester filaments, polyester yarn discs play an important role in the manufacturing of textiles, clothing and home supplies. In actual production and packaging, due to factors such as poor production environment cleanliness, unstable equipment operation and non-standard manual operation, defects such as poor forming, oil stains, paper tube damage and hooking hair often occur on the surface of polyester yarn discs. These defects not only seriously affect the product quality of polyester yarn discs, but also may cause problems such as broken lines and uneven dyeing in the subsequent processing, increasing production costs and reducing production efficiency.
[0003] Currently, the industry mainly uses manual detection to identify defects in the appearance of polyester yarn discs. However, manual detection has many drawbacks. On the one hand, manual detection is inefficient and cannot meet the needs of large-scale industrial production; on the other hand, detection personnel are easily affected by fatigue, subjective factors and other factors, resulting in poor stability and accuracy of the detection results, especially for defects such as hooking hair.
[0004] The detection of defects in the appearance of polyester yarn discs based on machine vision has the advantages of being fast, accurate and stable, and can effectively make up for the shortcomings of manual detection. The research and application of this technology has important practical significance for improving the production quality of polyester yarn discs, reducing production costs and promoting the intelligent development of the chemical fiber industry.
[0005] The detection of relatively obvious defects such as yarn disc forming abnormalities, surface oil stains and paper tube damage is relatively easy, but the detection effect of defects such as tail filament inclusion, waste filament inclusion and hooking hair is poor. The detection of these three types of defects is difficult because their color is consistent with the yarn disc, they are distributed on the upper and lower end surfaces of the yarn disc, and they are small and close to the end surface of the yarn disc, so the contrast of this type of defect in the image is low, and the algorithm has difficulty in distinguishing it from the texture on the surface of the yarn disc, greatly increasing the difficulty of image processing.
[0006] In view of the problem of "difficult detection" of defects in the appearance of yarn discs, research is conducted on the method of detecting defects in the appearance of yarn discs. The goal of the research is to realize high-accuracy detection of defects in the appearance of yarn discs by combining machine vision technology and deep learning algorithms. SUMMARY
[0007] The present application relates to a kind of SP-YOLO-based appearance defect automatic detection method of yarn disc, unlike existing yarn disc manual detection and traditional image processing method, propose a kind of appearance defect detection method of combining SP-YOLO framework of terylene yarn disc.The method can realize the automation detection of industrial production site while maintaining high detection precision, is expected to reduce labor cost, improve detection efficiency and reliability.
[0008] (1) data set making
[0009] The present application constructs the appearance defect image data set of terylene yarn disc, data set includes 13256 images, covers tail silk, waste silk, three kinds of defect samples such as hooking hair, and is labeled, data set is divided according to training set, verification set and test set, for model training and evaluation.
[0010] (2) SP-YOLO model construction
[0011] In order to further improve the detection ability of yarn disc appearance defect and thus reduce the missed detection rate, the present application introduces scale sequence feature fusion module, triple feature encoder module and channel and position attention mechanism module in the Neck part based on YOLOv11 framework, to adapt to the luster and texture characteristics of terylene yarn disc:
[0012] ① scale sequence feature fusion SSFF module
[0013] By applying size alignment to different scale feature maps, and then using 3D convolution for cross-scale fusion, unified feature expression for size defects is realized. This module effectively solves the missed detection problem caused by large size difference of yarn disc appearance defects.
[0014] ② triple feature encoder TFE module
[0015] Large, medium and small three kinds of feature maps are processed respectively: large feature retains defect edge gradient, medium feature maintains local semantics, and small feature contains global distribution. Then, through channel splicing, detail enhancement features are formed to ensure that defects will not be lost in the down-sampling process.
[0016] ③ channel and position attention mechanism CPAM module
[0017] Firstly, channel attention is used to filter important feature channels related to defects and suppress redundant channels corresponding to background yarn disc texture; then, position attention is used to accurately locate the defect area and enhance the perception ability of spatial distribution of defects. This module effectively solves the false detection problem under the interference of complex yarn disc surface texture.
[0018] (3) defect detection and position prediction
[0019] The detection branch predicts the multi-scale features output by Backbone and Neck, generating multiple candidate anchor boxes. Each anchor box contains the following information: defect coordinates: x, y, w, h; defect category: tail ribbon ingress, waste ribbon ingress, hook and burr; confidence level.
[0020] During training, this invention uses EIoU loss as the localization loss, combined with cls loss and dfl loss, and guides model training through weighted fusion of the three to achieve collaborative optimization of bounding box coordinate regression, classification and feature distribution learning.
[0021] (4) Test result output and industrial adaptation
[0022] Visualize the defect detection results output by the Head unit, specifically including:
[0023] By overlaying detection boxes, category, and confidence information onto the original image, the location of defects in the yarn reel can be visually displayed, facilitating manual verification and integration with the quality inspection process.
[0024] Through the above steps, the present invention can achieve automatic detection of appearance defects of polyester yarn discs while maintaining high precision, and the output results are concise and clear, making it suitable for industrial applications.
[0025] Compared with the prior art, the present invention has the following obvious advantages and beneficial effects:
[0026] Compared with traditional manual inspection methods, the SP-YOLO model for automatic detection of yarn disc appearance defects proposed in this invention, based on deep learning, introduces the SSFF module for deep fusion of multi-scale features, the TFE module for preserving defect details, and the CPAM module for suppressing background interference on the basis of the YOLOv11 basic model. It can effectively identify defect areas such as tail thread ingress, waste thread ingress, and snagging defects in polyester yarn discs with high accuracy. It can also be adapted to the detection of appearance defects in polyester yarn discs with different thicknesses and texture densities without requiring significant modifications to the model architecture, thus reducing the cost of industrial implementation. Attached Figure Description
[0027] Figure 1 Defect type diagram
[0028] Figure 2 Visual inspection results of appearance defects in polyester yarn reels
[0029] Figure 3 Architecture diagram of a polyester yarn reel appearance defect detection model based on SP-YOLO
[0030] Figure 4 TFE module schematic diagram
[0031] Figure 5CPAM module schematic diagram
[0032] Figure 6 Flowchart of a visual recognition method for detecting appearance defects in polyester yarn reels Detailed Implementation
[0033] Based on the above description, the following is a specific implementation process example, but the scope of protection of this patent is not limited to this implementation process.
[0034] First, images of polyester yarn reels are acquired, and a defect-annotated dataset is constructed. Second, an SP-YOLO neural network model is built. This model, based on YOLOv11, introduces a scale sequence feature fusion module, a triple feature encoder module, and a channel and position attention mechanism module in the Neck part. The SP-YOLO network model is trained using the annotated dataset, and the optimal detection performance weights are saved. The polyester yarn reel image to be detected is input into the trained model, and the defect detection results are output. Details are as follows:
[0035] Step 1: Building the dataset
[0036] Images of yarn disc appearance defects were captured using a yarn disc image acquisition prototype to construct a dataset of polyester yarn disc appearance defects. The dataset was then labeled using Labelimg software, and the labeled files were saved in txt file format adapted to YOLO series algorithms. The dataset was divided into training set, validation set, and test set.
[0037] Step 2: Construct a neural network model
[0038] Based on YOLOv11, three innovative modules, SSFF, TFE and CPAM, are introduced to generate the SP-YOLO neural network model. SP-YOLO consists of three main parts: feature extraction module, feature enhancement module and head detection module. The innovative modules are all introduced in the feature enhancement module.
[0039] Step 2.1: Construct the feature extraction module Backbone
[0040] Using CSPDarknet53 as the feature extraction unit, multi-scale feature extraction is performed on the preprocessed 4D tensor to output a multi-granularity feature map adapted for defect detection. The specific process is as follows:
[0041] Step 2.1.1: Constructing the Feature Extraction Architecture
[0042] CSPDarknet53 consists of one Conv-BN-SiLU unit with a 3×3 kernel and a stride of 2; five C3 modules containing three convolutional layers each; and one global average pooling layer. It transforms a 640×640 input image into five feature maps of different resolutions through progressive downsampling with a stride of 2: 640×640 with 64 channels (P1), 320×320 with 128 channels (P2), 160×160 with 256 channels (P3), 80×80 with 512 channels (P4), and 40×40 with 1024 channels (P5). The convolutional layers in the Conv-BN-SiLU unit and C3 module contain learnable weights and biases, and the BN layers contain learnable scaling and offset coefficients. These parameters are used for adaptive optimization of the feature extraction process.
[0043] Step 2.1.2: Core Feature Filtering
[0044] Based on the appearance defect features of polyester yarn discs, P3 contains defect details and P4 / P5 contain global semantics of large defects. P3, P4, and P5 are selected as inputs for subsequent feature fusion units, while P1 and P2, which have weak semantic information, are discarded to reduce computational complexity.
[0045] Step 2.2: Construct the feature enhancement module Neck
[0046] Based on the original Neck section, an innovative module, "Scale Sequence Feature Fusion SSFF - Triple Feature Encoder TFE - Channel and Position Attention Mechanism CPAM", is introduced to optimize P3, P4, and P5. The SSFF module deeply fuses multi-scale features, the TFE module preserves defect details, and the CPAM module suppresses background interference.
[0047] Step 2.2.1: Introduce the Scale Sequence Feature Fusion (SSFF) module
[0048] Specific process:
[0049] Step 2.2.1.1: Feature Size Alignment
[0050] The nearest neighbor interpolation method is used to adapt the step size to the feature map resolution difference, and the resolution of P4 and P5 is uniformly adjusted to 160×160 of P3. The adjusted P3, P4, and P5 are transformed from 3D tensors of H×W×C to 4D tensors of D×H×W×C through tensor dimension expansion operation, where D is the scale dimension with a value of 3, corresponding to the 3 feature map scales; W is the width, H is the height, and C is the number of channels of the image.
[0051] Step 2.2.1.2: 3D Convolutional Fusion
[0052] Three 4D tensors are concatenated along the scale dimension and input into a 3×3×3 kernel. The output consists of a 3D convolutional layer with 256 channels, a 3D batch normalization layer, and a SiLU activation function. The function expression is as follows:
[0053] SiLU(x)=x·σ(x)
[0054] Where x is the input value of the function, and σ(x) is the Sigmoid function, defined as:
[0055]
[0056] Used to "gated" the input x, that is, to dynamically adjust the weight of the input;
[0057] Output feature map F that integrates multi-scale defect features SSFF The feature map has a resolution of 160×160 and 256 channels;
[0058] 3D convolutional layers contain learnable 3×3×3 convolutional kernel weights and biases, and 3D batch normalization layers contain learnable scaling and offset coefficients.
[0059] Step 2.2.2: Introduce the Triple Feature Encoder (TFE) module
[0060] Specific process:
[0061] Step 2.2.2.1: Feature Size Decomposition
[0062] P3 is defined as a "large-size feature" containing defect edge details, P4 is defined as a "medium-size feature" containing local defect semantics, and P5 is defined as a "small-size feature" containing global defect distribution.
[0063] Step 2.2.2.2: Targeted Feature Processing
[0064] Large feature P3: Adaptive max pooling and adaptive average pooling are applied to the same spatial size as P4, and the results of the two are added together;
[0065] Small-size feature P5: Upsampled to the same spatial size as P4 using nearest neighbor interpolation to minimize the smoothing damage to details during the upsampling process;
[0066] Medium-sized feature P4: Does not change the size, and is directly used as the reference scale for fusion;
[0067] Step 2.2.2.3: Channel Dimension Stitching
[0068] The processed large, medium, and small feature maps are concatenated along the channel dimension C to output a detail-enhanced feature map F. TFEThe feature map has a resolution of 160×160 and 768 channels;
[0069] Step 2.2.3: Introduce the Channel and Position Attention Mechanism (CPAM) module
[0070] Specific process:
[0071] Step 2.2.3.1: Channel Attention Filtering
[0072] Enter F TFE Perform a global average pooling operation on each channel to obtain the channel statistics vector;
[0073] Using kernel size A 1D convolutional layer with 768 output channels captures local interactions between channels;
[0074] The channel attention weight vector is generated by the Sigmoid function, with a dimension of 768, and is related to F. TFE By multiplying each channel sequentially, we obtain the channel-optimized feature map F. CA The feature map has a resolution of 160×160 and 768 channels;
[0075] 1D convolutional layers contain learnable weights and biases;
[0076] Step 2.2.3.2: Location Attention Localization
[0077] Enter F CA With F SSFF The element-wise summation feature is used to perform average pooling along the width W and height H axes, respectively, as shown in the formula:
[0078]
[0079] Where E is the superimposed feature map, i,j are the row and column indices of the output feature map, used to locate the spatial position of the output features, p w p h The results are obtained by pooling the width and height axes, respectively, preserving the spatial location information of the defects.
[0080] p w With p h The layers are concatenated along the channel dimension, and a 1×1 convolutional layer with 768 input and output channels and a sigmoid activation function are used to generate a positional attention weight map F. PA The image has a resolution of 160×160 and 768 channels.
[0081] Each 1×1 convolutional layer contains learnable weights and biases for adaptive positional attention weights, enabling effective differentiation between defect features and the background.
[0082] Step 2.2.3.3: Attention Fusion
[0083] F CA With F PA Perform the Hadamard product element-wise multiplication to output the final optimized feature map F. CPAM The image has a resolution of 160×160 and 768 channels.
[0084] Step 2.3: Construct the Head Detection Module
[0085] Specific process:
[0086] Step 2.3.1: Location Prediction
[0087] Defect detection branch: for F CPAM The Backbone outputs P4 and P5, after channel adjustment, generate prediction anchor boxes at three scales, corresponding to small, medium, and large defects. Each anchor box contains three parts: "defect coordinates (x, y, w, h), defect category, and confidence level"; Step 2.3.2: Loss function optimization
[0088] Using EIoU loss as the box loss, the formula is:
[0089]
[0090]
[0091] ρ(w,w gt )=|ww gt |
[0092] ρ(h,h gt )=|hh gt |
[0093] Where ρ represents the Euclidean distance, IoU is the intersection-union ratio of the predicted bounding box and the ground truth bounding box, and b, b gt Here, represents the center coordinates of the predicted bounding box and the ground truth bounding box, respectively, and w and h represent the width and height of the predicted bounding box, respectively. gt h gt These are the width and height of the actual bounding box, respectively. c h c These are the width and height of the smallest bounding rectangle of the predicted bounding box and the ground truth bounding box, respectively.
[0094] Using cls loss as the classification loss, the formula is:
[0095]
[0096] Where N is the number of samples; C is the total number of categories; p ic y represents the class probability predicted by the model. ic For real labels, when y ic=1, which means that the i-th sample belongs to class c; otherwise, it is 0.
[0097] Distributed focus loss (DFL) is used as an auxiliary bounding box regression to refine the bounding box coordinate regression process, allowing the model to learn the probability distribution of the coordinates. The formula is:
[0098]
[0099] t0, t1: The discrete interval range corresponding to the bounding box coordinates; the continuous coordinates are divided into multiple discrete intervals, and the position of the true coordinates is represented by a probability distribution;
[0100] p k The probability of the true distribution, i.e., when the true coordinates fall in the k-th interval, p k =1, otherwise 0;
[0101] The interval probability distribution predicted by the model;
[0102] The three loss functions do not work independently, but rather work together in a weighted fusion to guide the model parameter update; each of them has its own specific optimization objective: EIoU loss focuses on localization accuracy, cls loss ensures classification accuracy, and dfl loss refines bounding box coordinate regression.
[0103] The three factors are summed by pre-setting weights to obtain the "total loss" of model training; the gradient is calculated based on the total loss value, and all parameters of the model are updated at once to avoid optimization conflicts in different dimensions;
[0104] The weights in this invention are fixed as follows: EIoU: 7.5; cls: 0.5; dfl: 1.5. The effects of different weight ratios are as follows: EIoU: the larger the weight, the more the model focuses on optimizing the "localization accuracy between the predicted box and the ground truth box"; cls: corresponds to classification loss, the larger the weight, the more the model focuses on optimizing the "classification accuracy of the target category"; dfl: corresponds to distribution focus loss, the larger the weight, the more the model focuses on "refined regression of bounding box coordinates".
[0105] The regression branch convolutional layer, classification branch convolutional layer, and DFL-supporting convolutional layer of the defect detection head all contain learnable weights and biases, which are adapted to EIoU loss, cls loss, and DFL loss, respectively.
[0106] The detection method described above is characterized in that:
[0107] Step 3: Neural Network Training
[0108] Step 3.1: Hardware Platform Setup
[0109] Training was conducted on an experimental platform with an NVIDIA RTX 4070 Tis graphics card and Python 3.9 as the programming language; the detection task was trained using the partitioned training and validation sets.
[0110] Step 3.2: Hyperparameter Settings
[0111] During training, 16 images were input each time, and the input image size was uniformly scaled to 640×640 pixels; the model training was iterated for 200 rounds, and a validation mechanism was enabled; for the loss function, EIoU loss was used as the localization loss, cls loss was used as the classification loss, and DFL loss was combined to assist training; the initial learning rate was 0.01, the final learning rate was 0.01, the momentum was set to 0.937, and the weight decay coefficient was 0.0005;
[0112] Step 3.3: Training Iteration and Parameter Update
[0113] Loss function-driven parameter optimization: During training, the model calculates the "loss between the predicted result and the true label". EIoU loss optimizes localization, cls loss optimizes classification, and dfl loss optimizes coordinate precision. Backpropagation is used to adjust network weights. The decrease in loss value is essentially a reduction in the model's "prediction error", making the predicted box fit the true box better and the category judgment more accurate.
[0114] As the loss function is optimized, especially the validation set loss decreases, the model's prediction performance on the validation set will gradually improve: the number of effectively matched prediction boxes (TP) increases, while the number of false positives (FP) and false negatives (FN) decreases, thereby improving Precision, Recall, and mAP. 50-95 rise;
[0115] Step 3.3.1: Forward Propagation and Loss Calculation
[0116] The model receives a batch of training set samples, extracts features through the backbone network, fuses features through the neck network, and outputs two prediction results from the detection head: bounding box distribution and class confidence. The prediction results are then decoded to obtain the actual bounding box coordinates and class probabilities.
[0117] Calculate the component losses:
[0118] EIOU loss: The localization loss is obtained by calculating the IoU between the predicted bounding box and the true bounding box, the center distance error, and the aspect ratio error.
[0119] DFL loss: Calculates the coordinate regression granularity loss based on the difference between the bounding box distribution prediction results and the true coordinates;
[0120] CLS loss: Calculate the difference between the predicted category and the true label to quantify the category judgment error; sum the results according to the set weights to obtain the total loss, which quantifies the overall error between the predicted result and the true label.
[0121] Step 3.3.2: Backpropagation to calculate gradient
[0122] The backpropagation of the total loss is invoked, and the partial derivative of the total loss with respect to each learnable parameter, i.e., the gradient, is automatically calculated based on the chain rule. The gradient starts from the detection head and backpropagates layer by layer to the neck network and the backbone network. The gradient value of each parameter reflects the degree of influence of that parameter on the total loss. The larger the absolute value of the gradient, the more significant the influence. The framework automatically accumulates the gradient of batch samples to prepare for parameter updates.
[0123] Step 3.3.3: SGD Optimizer Performs Parameter Update. The optimizer adjusts the parameters according to the gradient values of each parameter and the SGD update rule. The core formula is:
[0124]
[0125] θ t The model parameter set at iteration t is a vector composed of all learnable parameters. The learnable parameters of the model described in this patent cover the inherent parameters of all trainable layers in the backbone network, feature enhancement module, and defect detection head. Specifically, these include: the weights and biases of the Conv-BN-SiLU units and C3 modules in the CSPDarknet53 backbone network, as well as the scaling and offset coefficients of the corresponding BN layers; the weights and biases of the 3D convolutional layers and the scaling and offset coefficients of the 3D batch normalization layer in the SSFF module of the feature enhancement module; the weights and biases of the 1D convolutional layers and 1×1 convolutional layers in the CPAM module; and the weights and biases of the regression branch convolutional layers used for defect coordinate prediction, the classification branch convolutional layers used for category judgment, and the DFL convolutional layers used for coordinate refinement regression in the defect detection head. These parameters respectively serve feature extraction, multi-scale feature fusion, attention weight learning, defect coordinate regression, category recognition, and coordinate distribution optimization.
[0126] θ t+1 The set of model parameters updated after the (t+1)th iteration.
[0127] η: Learning rate, a positive hyperparameter that controls the magnitude of parameter updates; b: Batch size, the number of training samples used to calculate the gradient in a single iteration.
[0128] The loss function J in parameter θ t Sample (x) i ,y i The gradient vector of the parameter θ
[0129] The average gradient vector of all b samples in a batch
[0130] v t The momentum auxiliary vector at iteration t is used to record the historical parameter update trend. The specific numerical calculation of the momentum auxiliary vector is obtained through initialization and recursive calculation in each iteration: Before training begins, the momentum auxiliary vector v0 is initialized to a vector of all zeros, with its dimension completely consistent with the dimension of the model's learnable parameters θ; v0 is recursively calculated in each iteration. t+1 The value of the momentum auxiliary vector v in each iteration, starting from the first iteration. t+1 They will all use the momentum value v from the previous round. t +Gradient information for the current batch, calculated using a fixed formula. Calculations show that each dimension of the vector is calculated independently according to this logic, ultimately yielding v that is consistent with the parameter dimensions. t+1 Numerical value; v t It is an intermediate value that connects the preceding and following steps; the value of v calculated in each round. t+1 It will be automatically cached and used as v in the next iteration. t
[0131] v t+1 The momentum auxiliary vector updated after the (t+1)th iteration.
[0132] γ: Momentum coefficient, ranging from [0,1), used to control the degree of influence of historical update trends on the current update.
[0133] x i The input feature vector of the i-th training sample
[0134] y i : The true label corresponding to the i-th training sample
[0135] J(θ t ;x i ,y i ): Loss function, used to quantize the parameter θ t The following model applies to samples (x) i ,y i The gradient is cleared by performing a gradient clearing operation on the prediction error of the current batch of samples to avoid the accumulation of gradients between the current batch and the next batch of samples, and to ensure that the gradient is calculated independently for each batch of training.
[0136] Step 3.5: Save the optimal model
[0137] Step 3.5.1: Calculate the evaluation metrics after each training round. Before calculating any metric, it is essential to first determine whether the predicted bounding box "hits" the ground truth bounding box using the Intersection over Union (IoU). This is the core prerequisite for object detection metrics.
[0138] IoU definition: the area of the intersection of the predicted bounding box and the area of the union of the predicted bounding box and the ground truth bounding box, ranging from 0 to 1, with the closer to 1 indicating a higher degree of overlap;
[0139] Formula: IoU=(A∩B) / (A∪B), where A is the predicted bounding box area and B is the ground truth bounding box area; Matching threshold: The default is IoU≥0.7, if satisfied, it is considered "valid match" or prediction is successful; if not satisfied, it is considered "invalid prediction" or false alarm.
[0140] Matching rules: A ground truth bounding box can only match one predicted bounding box with the highest confidence level to avoid double counting; predicted bounding boxes that do not match ground truth bounding boxes are considered "false positives"; ground truth bounding boxes that are not matched by predicted bounding boxes are considered "missed detections".
[0141] Core indicator calculation method:
[0142] 1. Precision: "Of the predicted positive examples, how many were actually correct?"
[0143] Core definition: The ratio of all true positive (TP) boxes that are "effectively matched predicted boxes" to all true positive (TP) boxes that are "predicted as positive by the model" plus false positive (FP) boxes;
[0144] Calculation formula:
[0145] Precision = TP / (TP + FP)
[0146] TP: A predicted bounding box with IoU ≥ threshold and correct category prediction; a true positive prediction is correct.
[0147] FP: IoU < threshold, or a predicted bounding box with an incorrect category prediction, resulting in a false positive.
[0148] In layman's terms: Precision measures a model's ability to "report fewer false alarms." The higher the precision, the fewer false alarms. Example: If a model predicts 10 "defect boxes," and 7 of them match real defects, then TP = 7, and 3 are false alarms, then FP = 3. Therefore, Precision = 7 / (7+3) = 70%.
[0149] 2. Recall: "Of all the true positive examples, how many did the model find?"
[0150] Core definition: The proportion of all "effectively matched predicted boxes" to all "real positive boxes in the dataset";
[0151] Calculation formula:
[0152] Recall = TP / (TP + FN)
[0153] FN: The ground truth bounding box that was not matched by any predicted bounding box;
[0154] 3.mAP 50 "Average precision across all categories when IoU = 0.5"
[0155] mAP 50 It is "AP for each category" 50 "average value", while AP 50 It is the area under the Precision-Recall curve for a single category, calculated in 3 steps:
[0156] Step 1: Calculate AP for a single category 50
[0157] Filter the predicted bounding boxes for this category: retrieve all boxes that the model predicts to be in the current category, and sort them from highest to lowest confidence.
[0158] Calculate cumulative TP and FP for each box: Starting with the box with the highest confidence, check each box one by one to see if it matches the true box of that category, and accumulate the number of TP and FP.
[0159] Calculate point-by-point Precision and Recall: For each prediction box, use the currently accumulated TP and FP to calculate the corresponding Precision and Recall, forming a series of (Recall, Precision) coordinate points;
[0160] Plot the PR curve and calculate its area: Connect all coordinate points with Recall on the horizontal axis and Precision on the vertical axis to form the PR curve; AP 50 It refers to the area under this curve, ranging from 0 to 1; the larger the area, the better the performance.
[0161] Step 2: Calculate mAP 50
[0162] AP for all categories in the dataset 50 Add them together, then divide by the total number of categories, that is: mAP 50 =(AP 50 Category 1+AP 50 Category 2+...+AP 50 _Category N) / N
[0163] In layman's terms: mAP (Maximum Applicability) is the model's comprehensive detection capability across all categories. 50 The higher the value, the more stable the model's performance under normal positioning accuracy.
[0164] 4.mAP 50-95 "Average precision across all categories under multiple IoU thresholds"
[0165] mAP 50-95 It is more than mAP 50A more stringent comprehensive indicator, the core of which is "covering different positioning accuracy requirements", is calculated in three steps:
[0166] Step 1: Determine the IoU threshold range. The IoU threshold range is from 0.5 to 0.95, with values incremented by 0.05, for a total of 10 thresholds.
[0167] [0.5,0.55,0.6,0.65,0.7,0.75,0.8,0.85,0.9,0.95];
[0168] Step 2: Calculate mAP for each threshold
[0169] For each IoU threshold, such as 0.5, 0.55...0.95, calculate based on mAP. 50 The calculation logic calculates the mAP at the threshold, i.e., mAP. 55 mAP 60 ...mAP 95 ;
[0170] Step 3: Calculate mAP 50-95
[0171] Add the mAP corresponding to the 10 thresholds together, then divide by 10, that is:
[0172] mAP 50-95 =(mAP 50 +mAP 55 +...+mAP 95 ) / 10
[0173] In layman's terms: mAP measures the overall performance of a model under positioning requirements ranging from lenient to strict. 50-95 The higher the value, the more stable the model's positioning accuracy and the stronger its adaptability to different scenarios.
[0174] Step 3.5.2: Save the optimal model framework based on evaluation metrics. This will save the core metric mAP of the current epoch. 50-95 Compared with the "historical best index", a model is considered a "new best model" if it meets the following two conditions:
[0175] Core condition: Current mAP 50-95 Historical high mAP 50-95 The increase must exceed the threshold, which defaults to 0.001, or 0.1%.
[0176] Example: Historical high mAP 50-95 =0.78, the current epoch is 0.781, an increase of 0.001 satisfies the condition; if the increase is only 0.0005 and does not reach the threshold, it is not considered optimal.
[0177] Auxiliary condition: When the core indicator improves, the auxiliary indicator mAP also improves. 50 Recall did not decrease significantly, and the framework allows for small fluctuations by default to avoid misjudging "unbalanced models" as optimal.
[0178] Counterexample: Current mAP 50-95 The accuracy improved from 0.78 to 0.782, but recall dropped from 0.85 to 0.7, and the false negative rate increased sharply. The framework may think that the model "overly pursues localization accuracy at the expense of recall" and may not update the optimal model.
[0179] If the current model is determined to be the "historical best", the framework will perform the following operations:
[0180] Save the best model overwrite: Save the model weights of the current epoch, including the backbone network, detector head, and loss function parameters, as the best model, overwriting the previous "historical best model";
[0181] Record optimal information: Mark the current epoch and all corresponding metrics in the log for easy follow-up;
[0182] The optimal model is "dynamically covered". It is only updated when a "new optimal model" appears. If it is not optimal, it will not be modified, ensuring that the model with the "strongest overall performance throughout the entire training process" is saved in the end.
[0183] When the loss function is optimized to the point where "the validation set metric reaches its historical high", such as mAP 50-95 The framework will only save the current model as the optimal model when it exceeds the previous optimal value; if the loss function continues to decrease, but the validation set metric starts to decline, the optimal model will not be updated.
[0184] Step 4: Defect Identification
[0185] The SP-YOLO network model and its trained optimal weights are used to detect defects in the appearance of polyester yarn reels. Test set images or any images to be detected are input into the SP-YOLO model as the dataset. The model performs inference on the input dataset and visualizes the detection results, generating bounding boxes for yarn reel appearance defects on the original images, labeling defect categories and confidence levels. The model's detection results are output as follows: Figure 2 As shown.
[0186] The results of this example are shown below:
[0187] Table 1 Comparison of detection metrics between YOLOv11 and SP-YOLO
[0188]
[0189] Compared to the baseline YOLOv11, the overall accuracy (P) of SP-YOLO improved from 0.817 to 0.820 after this modification, and the mAP increased. 50-95 The improvement from 0.493 to 0.508 indicates more robust localization and detail retention at high IoU thresholds. Among defect types, scrap ribbon ingress shows the most significant improvement, with P increasing from 0.852 to 0.857, R from 0.824 to 0.838, and mAP... 50 From 0.886 to 0.894, mAP 50-95 The increase from 0.578 to 0.616 indicates that the change is more effective in preserving edge and outline details in scenes with strong texture interference; the tail ribbon ingress and hook-and-pick fur are also improved in mAP. 50-95 A slight improvement. Overall, this change largely maintains the same mAP. 50 At the same time, it significantly improved mAP 50-95 This study validated the effectiveness of multi-scale alignment and detail preservation strategies for defect detection in polyester yarn coils. The core requirement in this scenario is to prioritize avoiding missed defect detections. The cost of defective products flowing downstream due to missed detections far outweighs the minor deviations at the edges of the prediction box. mAP... 50 The focus is precisely on the key objective of "accurately locating the defect," rather than demanding higher overlap in bounding box details. Furthermore, defects in polyester yarn discs are often characterized by irregular edges and low contrast with the background; a high IoU threshold can easily mask the model's actual detection capability due to biases in bounding box details, thus affecting mAP. 50 On the contrary, it can more realistically reflect the model's effective ability to identify such defects, and its improvement directly corresponds to the reduction of the missed detection rate in production, which has clear practical application value.
Claims
1. A visual recognition method for detecting appearance defects in polyester yarn reels, characterized in that: First, images of polyester yarn discs are acquired and a defect-labeled dataset is constructed. Second, an SP-YOLO neural network model is constructed. Based on YOLOv11, this model introduces a scale sequence feature fusion module, a triple feature encoder module, and a channel and position attention mechanism module in the Neck part. The SP-YOLO network model is trained using a labeled dataset, and the optimal detection performance weights are saved. The image of the polyester yarn disc to be detected is input into the trained model, and the defect detection result is output. This achieves the detection of appearance defects in the polyester yarn disc.
2. The detection method according to claim 1, characterized in that: Step 1: Building the dataset Images of yarn disc appearance defects were captured using a yarn disc image acquisition prototype to construct a dataset of polyester yarn disc appearance defects. The dataset was then labeled using Labelimg software, and the labeled files were saved in txt file format adapted to YOLO series algorithms. The dataset was divided into training set, validation set, and test set.
3. The detection method according to claim 1, characterized in that: Step 2: Construct a neural network model Based on YOLOv11, three innovative modules, SSFF, TFE and CPAM, are introduced to generate the SP-YOLO neural network model. SP-YOLO consists of three main parts: feature extraction module, feature enhancement module and head detection module. The innovative modules are all introduced in the feature enhancement module. Step 2.1: Construct the feature extraction module Backbone Using CSPDarknet53 as the feature extraction unit, multi-scale feature extraction is performed on the preprocessed 4D tensor to output a multi-granularity feature map adapted for defect detection. The specific process is as follows: Step 2.1.1: Constructing the Feature Extraction Architecture CSPDarknet53 consists of one Conv-BN-SiLU unit with a 3×3 kernel and a stride of 2; five C3 modules containing three convolutional layers each; and one global average pooling layer. It transforms a 640×640 input image into five feature maps of different resolutions through progressive downsampling with a stride of 2: 640×640 with 64 channels (P1), 320×320 with 128 channels (P2), 160×160 with 256 channels (P3), 80×80 with 512 channels (P4), and 40×40 with 1024 channels (P5). The convolutional layers in the Conv-BN-SiLU unit and C3 module contain learnable weights and biases, and the BN layers contain learnable scaling and offset coefficients. These parameters are used for adaptive optimization of the feature extraction process. Step 2.1.2: Core Feature Filtering Based on the appearance defect features of polyester yarn discs, P3 contains defect details and P4 / P5 contain global semantics of large defects. P3, P4, and P5 are selected as inputs for subsequent feature fusion units, while P1 and P2, which have weak semantic information, are discarded to reduce computational complexity. Step 2.2: Construct the feature enhancement module Neck Based on the original Neck section, an innovative module, "Scale Sequence Feature Fusion SSFF - Triple Feature Encoder TFE - Channel and Position Attention Mechanism CPAM", is introduced to optimize P3, P4, and P5. The SSFF module deeply fuses multi-scale features, the TFE module preserves defect details, and the CPAM module suppresses background interference. Step 2.2.1: Introduce the Scale Sequence Feature Fusion (SSFF) module Specific process: Step 2.2.1.1: Feature Size Alignment The nearest neighbor interpolation method is used to adapt the step size to the feature map resolution difference, and the resolution of P4 and P5 is uniformly adjusted to 160×160 of P3. The adjusted P3, P4, and P5 are transformed from 3D tensors of H×W×C to 4D tensors of D×H×W×C through tensor dimension expansion operation, where D is the scale dimension with a value of 3, corresponding to the 3 feature map scales; W is the width, H is the height, and C is the number of channels of the image. Step 2.2.1.2: 3D Convolutional Fusion Three 4D tensors are concatenated along the scale dimension and input into a 3×3×3 kernel. The output consists of a 3D convolutional layer with 256 channels, a 3D batch normalization layer, and a SiLU activation function. The function expression is as follows: SiLU(x)=x·σ(x) Where x is the input value of the function, and σ(x) is the Sigmoid function, defined as: Used to "gated" the input x, that is, to dynamically adjust the weight of the input; Output feature map F that integrates multi-scale defect features SSFF The feature map has a resolution of 160×160 and 256 channels; 3D convolutional layers contain learnable 3×3×3 convolutional kernel weights and biases, and 3D batch normalization layers contain learnable scaling and offset coefficients. Step 2.2.2: Introduce the Triple Feature Encoder (TFE) module Specific process: Step 2.2.2.1: Feature Size Decomposition P3 is defined as "large-size feature", containing defect edge details; P4 is defined as "medium-size feature", containing local defect semantics; and P5 is defined as "small-size feature", containing global defect distribution. Step 2.2.2.2: Targeted Feature Processing Large feature P3: Adaptive max pooling and adaptive average pooling are applied to the same spatial size as P4, and the results of the two are added together; Small-size feature P5: Upsampled to the same spatial size as P4 using nearest neighbor interpolation to minimize the smoothing damage to details during the upsampling process; Medium-sized feature P4: Does not change the size, and is directly used as the reference scale for fusion; Step 2.2.2.3: Channel Dimension Stitching The processed large, medium, and small feature maps are concatenated along the channel dimension C to output a detail-enhanced feature map F. TFE The feature map has a resolution of 160×160 and 768 channels; Step 2.2.3: Introduce the Channel and Position Attention Mechanism (CPAM) module Specific process: Step 2.2.3.1: Channel Attention Filtering Enter F TFE Perform a global average pooling operation on each channel to obtain the channel statistics vector; Using kernel size A 1D convolutional layer with 768 output channels captures local interactions between channels; The channel attention weight vector is generated by the Sigmoid function, with a dimension of 768, and is related to F. TFE By multiplying each channel sequentially, we obtain the channel-optimized feature map F. CA The feature map has a resolution of 160×160 and 768 channels; 1D convolutional layers contain learnable weights and biases; Step 2.2.3.2: Location Attention Localization Enter F CA With F SSFF The element-wise summation feature is used to perform average pooling along the width W and height H axes, respectively, as shown in the formula: Where E is the superimposed feature map, i,j are the row and column indices of the output feature map, used to locate the spatial position of the output features, p w p h The results are obtained by pooling the width and height axes, respectively, preserving the spatial location information of the defects. p w With p h The layers are concatenated along the channel dimension, and a 1×1 convolutional layer with 768 input and output channels and a sigmoid activation function are used to generate a positional attention weight map F. PA The image has a resolution of 160×160 and 768 channels. Each 1×1 convolutional layer contains learnable weights and biases for adaptive positional attention weights, enabling effective differentiation between defect features and the background. Step 2.2.3.3: Attention Fusion F CA With F PA Perform the Hadamard product element-wise multiplication to output the final optimized feature map F. CPAM The image has a resolution of 160×160 and 768 channels. Step 2.3: Construct the Head Detection Module Specific process: Step 2.3.1: Location Prediction Defect detection branch: for F CPAM The Backbone outputs P4 and P5, after channel adjustment, generate predicted anchor boxes at three scales, corresponding to small, medium, and large defects. Each anchor box contains three parts: defect coordinates (x, y, w, h), defect category, and confidence level. Step 2.3.2: Loss function optimization uses EIoU loss as the box loss, with the following formula: ρ(w,w gt )=|w-w gt | ρ(h,h gt )=|h-h gt | Where ρ represents the Euclidean distance, IoU is the intersection-union ratio of the predicted bounding box and the ground truth bounding box, and b, b gt Here, represents the center coordinates of the predicted bounding box and the ground truth bounding box, respectively, and w and h represent the width and height of the predicted bounding box, respectively. gt h gt These are the width and height of the actual bounding box, respectively. c h c These are the width and height of the smallest bounding rectangle of the predicted bounding box and the ground truth bounding box, respectively. Using cls loss as the classification loss, the formula is: Where N is the number of samples; C is the total number of categories; p ic y represents the class probability predicted by the model. ic For real labels, when y ic =1, which means that the i-th sample belongs to class c; otherwise, it is 0. Distributed focus loss (DFL) is used as an auxiliary bounding box regression to refine the bounding box coordinate regression process, allowing the model to learn the probability distribution of the coordinates. The formula is: t0, t1: The discrete interval range corresponding to the bounding box coordinates; the continuous coordinates are divided into multiple discrete intervals, and the position of the true coordinates is represented by a probability distribution; p k The probability of the true distribution, i.e., when the true coordinates fall in the k-th interval, p k =1, otherwise 0; The interval probability distribution predicted by the model; The three loss functions do not work independently, but rather work together in a weighted fusion to guide the model parameter update; each of them has its own specific optimization objective: EIoU loss focuses on localization accuracy, cls loss ensures classification accuracy, and dfl loss refines bounding box coordinate regression. The three factors are summed by pre-setting weights to obtain the "total loss" of model training; the gradient is calculated based on the total loss value, and all parameters of the model are updated at once to avoid optimization conflicts in different dimensions; The weights in this invention are fixed as follows: EIoU: 7.5; cls: 0.5; dfl: 1.
5. The effects of different weight ratios are as follows: EIoU: the larger the weight, the more the model focuses on optimizing the "localization accuracy between the predicted box and the ground truth box"; cls: corresponds to classification loss, the larger the weight, the more the model focuses on optimizing the "classification accuracy of the target category"; dfl: corresponds to distribution focus loss, the larger the weight, the more the model focuses on "refined regression of bounding box coordinates". The regression branch convolutional layer, classification branch convolutional layer, and DFL-supporting convolutional layer of the defect detection head all contain learnable weights and biases, which are adapted to EIoU loss, cls loss, and DFL loss, respectively.
4. The detection method according to claim 1, characterized in that: Step 3: Neural Network Training Step 3.1: Hardware Platform Setup Training was conducted on an experimental platform with an NVIDIA RTX 4070 Tis graphics card and Python 3.9 as the programming language; the detection task was trained using the partitioned training and validation sets. Step 3.2: Hyperparameter Settings During training, 16 images were input each time, and the input image size was uniformly scaled to 640×640 pixels; the model training was iterated for 200 rounds, and a validation mechanism was enabled; for the loss function, EIoU loss was used as the localization loss, cls loss was used as the classification loss, and DFL loss was combined to assist training; the initial learning rate was 0.01, the final learning rate was 0.01, the momentum was set to 0.937, and the weight decay coefficient was 0.0005; Step 3.3: Training Iteration and Parameter Update Loss function-driven parameter optimization: During training, the model calculates the "loss between the predicted result and the true label". EIoU loss optimizes localization, cls loss optimizes classification, and dfl loss optimizes coordinate precision. Backpropagation is used to adjust network weights. The decrease in loss value is essentially a reduction in the model's "prediction error", making the predicted box fit the true box better and the category judgment more accurate. As the loss function is optimized, especially the validation set loss decreases, the model's prediction performance on the validation set will gradually improve: the number of effectively matched prediction boxes (TP) increases, while the number of false positives (FP) and false negatives (FN) decreases, thereby improving Precision, Recall, and mAP. 50-95 rise; Step 3.3.1: Forward Propagation and Loss Calculation The model receives a batch of training set samples, extracts features through the backbone network, fuses features through the neck network, and outputs two prediction results from the detection head: bounding box distribution and class confidence. The prediction results are then decoded to obtain the actual bounding box coordinates and class probabilities. Calculate the component losses: EIOU loss: The localization loss is obtained by calculating the IoU between the predicted bounding box and the true bounding box, the center distance error, and the aspect ratio error. DFL loss: Calculates the coordinate regression granularity loss based on the difference between the bounding box distribution prediction results and the true coordinates; CLS loss: Calculate the difference between the predicted category and the true label to quantify the category judgment error; sum the results according to the set weights to obtain the total loss, which quantifies the overall error between the predicted result and the true label. Step 3.3.2: Backpropagation to calculate gradient The backpropagation of the total loss is invoked, and the partial derivative of the total loss with respect to each learnable parameter, i.e., the gradient, is automatically calculated based on the chain rule. The gradient starts from the detection head and backpropagates layer by layer to the neck network and the backbone network. The gradient value of each parameter reflects the degree of influence of that parameter on the total loss. The larger the absolute value of the gradient, the more significant the influence. The framework automatically accumulates the gradient of batch samples to prepare for parameter updates. Step 3.3.3: SGD optimizer performs parameter updates The optimizer adjusts the parameters according to the gradient values of each parameter and the SGD update rule. The core formula is: θ t The model parameter set at iteration t is a vector composed of all learnable parameters. The learnable parameters of the model described in this patent cover the inherent parameters of all trainable layers in the backbone network, feature enhancement module, and defect detection head. Specifically, these include: the weights and biases of the Conv-BN-SiLU units and C3 modules in the CSPDarknet53 backbone network, as well as the scaling and offset coefficients of the corresponding BN layers; the weights and biases of the 3D convolutional layers and the scaling and offset coefficients of the 3D batch normalization layer in the SSFF module of the feature enhancement module; the weights and biases of the 1D convolutional layers and 1×1 convolutional layers in the CPAM module; and the weights and biases of the regression branch convolutional layers used for defect coordinate prediction, the classification branch convolutional layers used for category judgment, and the DFL convolutional layers used for coordinate refinement regression in the defect detection head. These parameters respectively serve feature extraction, multi-scale feature fusion, attention weight learning, defect coordinate regression, category recognition, and coordinate distribution optimization. θ t+1 The set of model parameters updated after the (t+1)th iteration. η: Learning rate, a positive hyperparameter that controls the magnitude of parameter updates. b: Batch size, i.e., the number of training samples used to calculate the gradient in a single iteration. The loss function J in parameter θ t Sample (x) i ,y i The gradient vector of the parameter θ The average gradient vector of all b samples in a batch v t The momentum auxiliary vector at iteration t is used to record the historical parameter update trend. The specific numerical calculation of the momentum auxiliary vector is obtained through initialization and recursive calculation in each iteration: Before training begins, the momentum auxiliary vector v0 is initialized to a vector of all zeros, with its dimension completely consistent with the dimension of the model's learnable parameters θ; v0 is recursively calculated in each iteration. t+1 The value of the momentum auxiliary vector v in each iteration, starting from the first iteration. t+1 They will all use the momentum value v from the previous round. t +Gradient information for the current batch, calculated using a fixed formula. Calculations show that each dimension of the vector is calculated independently according to this logic, ultimately yielding v that is consistent with the parameter dimensions. t+1 Numerical value; v t It is an intermediate value that connects the preceding and following steps; the value of v calculated in each round. t+1 It will be automatically cached and used as v in the next iteration. t v t+1 The momentum auxiliary vector updated after the (t+1)th iteration. γ: Momentum coefficient, ranging from [0,1), used to control the degree of influence of historical update trends on the current update. x i The input feature vector of the i-th training sample y i : The true label corresponding to the i-th training sample J(θ t ;x i ,y i ): Loss function, used to quantize the parameter θ t The following model applies to samples (x) i ,y i Prediction error Perform gradient zeroing to prevent the gradient of the next batch of samples from accumulating with that of the current batch, and ensure that the gradient is calculated independently for each batch of training. Step 3.5: Save the optimal model Step 3.5.1: Calculate the evaluation metrics after each round of training. Before calculating any metrics, it is essential to first determine whether the predicted bounding box "hits" the ground truth bounding box using the Intersection over Union (IoU). This is the core premise of object detection metrics. IoU definition: the area of the intersection of the predicted bounding box and the area of the union of the predicted bounding box and the ground truth bounding box, ranging from 0 to 1, with the closer to 1 indicating a higher degree of overlap; Formula: IoU=(A∩B) / (A∪B), where A is the predicted bounding box area and B is the ground truth bounding box area; Matching threshold: The default is IoU≥0.7, if it is met, it is considered "valid match" and that is, the prediction is successful; if it is not met, it is considered "invalid prediction" and that is, a false alarm. Matching rules: A ground truth bounding box can only match one predicted bounding box with the highest confidence level to avoid double counting; predicted bounding boxes that do not match ground truth bounding boxes are considered "false positives"; ground truth bounding boxes that are not matched by predicted bounding boxes are considered "missed detections". Core indicator calculation method:
1. Precision: "Of the predicted positive examples, how many were actually correct?" Core definition: The ratio of all true positive (TP) boxes that are "effectively matched predicted boxes" to all true positive (TP) boxes that are "predicted as positive by the model" plus false positive (FP) boxes; Calculation formula: Precision = TP / (TP + FP) TP: A predicted bounding box with IoU ≥ threshold and correct category prediction; a true positive prediction is correct. FP: IoU < threshold, or a predicted bounding box with an incorrect category prediction, resulting in a false positive. In layman's terms: Precision measures a model's ability to "report fewer false alarms." The higher the precision, the fewer false alarms. Example: If a model predicts 10 "defect boxes," and 7 of them match real defects, then TP = 7, and 3 are false alarms, then FP = 3. Therefore, Precision = 7 / (7+3) = 70%.
2. Recall: "Of all the true positive examples, how many were found by the model?" Core definition: The proportion of all "effectively matched predicted boxes" to all "real positive boxes in the dataset"; Calculation formula: Recall = TP / (TP + FN) FN: The ground truth bounding box that was not matched by any predicted bounding box; 3.mAP 50 "Average precision across all categories when IoU = 0.5" mAP 50 It is "AP for each category" 50 "average value", while AP 50 It is the "area under the Precision-Recall curve" for a single category, calculated in 3 steps: Step 1: Calculate AP for a single category 50 Filter the predicted bounding boxes for this category: retrieve all boxes that the model predicts to be in the current category, and sort them from highest to lowest confidence. Calculate cumulative TP and FP for each box: Starting with the box with the highest confidence, check each box one by one to see if it matches the true box of that category, and accumulate the number of TP and FP. Calculate point-by-point Precision and Recall: For each prediction box, use the currently accumulated TP and FP to calculate the corresponding Precision and Recall, forming a series of (Recall, Precision) coordinate points; Plot the PR curve and calculate its area: Connect all coordinate points with Recall on the horizontal axis and Precision on the vertical axis to form the PR curve; AP 50 It refers to the area under this curve, ranging from 0 to 1; the larger the area, the better the performance. Step 2: Calculate mAP 50 AP for all categories in the dataset 50 Add them together, then divide by the total number of categories, that is: mAP 50 =(AP 50 Category 1+AP 50 Category 2+...+AP 50 _Category N) / N In layman's terms: mAP (Maximum Applicability) is the model's comprehensive detection capability across all categories. 50 The higher the value, the more stable the model's performance under normal positioning accuracy. 4.mAP 50-95 "Average precision across all categories under multiple IoU thresholds" mAP 50-95 It is more than mAP 50 A more stringent comprehensive indicator, the core of which is "covering different positioning accuracy requirements", is calculated in three steps: Step 1: Determine the IoU threshold range The IoU threshold is set from 0.5 to 0.95, with a value taken in increments of 0.05, for a total of 10 thresholds: [0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, 0.9, 0.95]. Step 2: Calculate mAP for each threshold For each IoU threshold, such as 0.5, 0.55...0.95, calculate based on mAP. 50 The calculation logic calculates the mAP at the threshold, i.e., mAP. 55 mAP 60 ...mAP 95 ; Step 3: Calculate mAP 50-95 Add the mAP corresponding to the 10 thresholds, then divide by 10 to get: mAP 50-95 =(mAP 50 +mAP 55 +...+mAP 95 ) / 10 In layman's terms: mAP measures the overall performance of a model under positioning requirements ranging from lenient to strict. 50-95 The higher the value, the more stable the model's positioning accuracy and the stronger its adaptability to different scenarios. Step 3.5.2: Save the optimal model based on the evaluation indicators. The framework will use the core metric mAP of the current epoch. 50-95 Compared with the "historical best index", a model is considered a "new best model" if it meets the following two conditions: Core condition: Current mAP 50-95 Historical high mAP 50-95 The increase must exceed the threshold, which defaults to 0.001, or 0.1%. Example: Historical high mAP 50-95 =0.78, the current epoch is 0.781, an increase of 0.001 satisfies the condition; if the increase is only 0.0005 and does not reach the threshold, it is not considered optimal. Auxiliary condition: When the core indicator improves, the auxiliary indicator mAP also improves. 50 Recall did not decrease significantly, and the framework allows for small fluctuations by default to avoid misjudging "unbalanced models" as optimal. Counterexample: Current mAP 50-95 The accuracy improved from 0.78 to 0.782, but recall dropped from 0.85 to 0.7, and the false negative rate increased sharply. The framework may think that the model "overly pursues localization accuracy at the expense of recall" and may not update the optimal model. If the current model is determined to be the "historical best", the framework will perform the following operations: Save the best model overwrite: Save the model weights of the current epoch, including the backbone network, detector head, and loss function parameters, as the best model, overwriting the previous "historical best model"; Record optimal information: Mark the current epoch and all corresponding metrics in the log for easy follow-up; The optimal model is "dynamically covered" and is only updated when a "new optimal model" appears. If it is not optimal, it will not be modified, ensuring that the model that has the "strongest overall performance throughout the entire training process" is saved in the end. The framework will save the current model as the optimal model only when the loss function is optimized to the point where the validation set metric reaches its historical high. If the loss function continues to decrease but the validation set metric begins to fall, the optimal model will not be updated.
5. The detection method according to claim 1, characterized in that: Step 4: Defect Identification The SP-YOLO network model and the optimal model trained are used to detect defects in the appearance of polyester yarn discs. The test set images or any images to be detected are input into the SP-YOLO model as the dataset. The model will infer the input dataset and visualize the detection results, generating yarn disc appearance defect boxes, labeling defect categories and confidence levels on the original images.