Vamp defect detection method and system based on deep learning
By improving the YOLOv11 algorithm, the CCAS-YOLO model is constructed and combined with the feature pyramid network and loss function, the problems of slow manual detection speed and inconsistent standards in the existing technology are solved, and efficient and accurate upper defect detection is achieved.
Patent Information
- Application Number
- CN202510116870.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-06-20
AI Technical Summary
In the prior art, upper defect detection relies on manual labor, and there are problems such as slow detection speed, inconsistent standards, and high working intensity, which affects the production efficiency and quality of shoemaking.
Using the upper defect detection method based on deep learning, the CCAS-YOLO model is built by improving the network framework of the YOLOv11 algorithm, combining the C3k2-CaFormer-CGLU module, ASF-YOLO and SDI to build a feature pyramid network, and integrating the Focaler-Shape-IoU loss function to improve detection efficiency and accuracy.
It realizes automatic identification and classification of upper defects, improves detection efficiency and accuracy, reduces calculation volume, and adapts to the upper inspection needs of different batches and specifications.
Smart Images

Figure CN120182176A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method and system for detecting shoe upper defects, and in particular to a method and system for detecting shoe upper defects based on deep learning, belonging to the technical field of object detection. Background Art
[0002] In recent years, with the rapid development of artificial intelligence and machine learning technologies, the field of shoe upper defect detection needs to gradually move away from manual detection and transfer to intelligent detection. Deep learning algorithms train models to enable the models to automatically identify defect features on shoe uppers. This method can achieve automatic identification and classification of shoe upper defects, further improving the detection efficiency and accuracy. And an adaptive algorithm is used to dynamically adjust the detection parameters to meet the detection requirements of shoe uppers of different batches and specifications. This also helps to improve the flexibility and accuracy of detection.
[0003] Most shoe manufacturing enterprises are in the semi-automation stage mainly relying on manual labor, and the detection link is completed by manual detection. However, manual detection has problems such as slow detection speed, inconsistent detection standards, high work intensity, and inconsistent detection means, which greatly affect the production efficiency and quality of shoe manufacturing. Summary of the Invention
[0004] Object of the Invention: The object of the present invention is to provide a method and system for detecting shoe upper defects based on deep learning that can improve the detection efficiency.
[0005] Technical Solution: A method for detecting shoe upper defects based on deep learning according to the present invention includes:
[0006] (1) Collect shoe upper defect pictures, classify and label them according to four defect types of cracking, glue overflow, breakage, and stain, and construct a data set;
[0007] (2) Improve the network framework based on the YOLOv11 algorithm to obtain an improved CCAS-YOLO model; the improvement includes: constructing a C3k2-CaFormer-CGLU module to replace the C3k2 module of the original model, combining ASF-YOLO and SDI to construct a feature pyramid network, and fusing and creating a loss function Focaler-Shape-IoU to replace the CIoU loss function of the original model;
[0008] (3) Use the improved CCAS-YOLO model to detect shoe upper defects in the constructed data set.
[0009] Furthermore, the step (1) further includes processing the collected shoe upper defect pictures, and the preprocessing includes grayscale transformation and image filtering;
[0010] The gray-scale transformation weakens the background, enhances the dynamic range, and enhances the target features by redistributing the pixel values of the image. Specifically, a linear function is used as the gray-scale mapping relationship between the original gray-scale f(x) and the transformed gray-scale g(x), and the mathematical expression is:
[0011]
[0012] where Q1 and Q2 are gray-scale thresholds, with Q1 < Q2, and Z1 and Z2 are the new gray-scales in the output image corresponding to the gray-scale value interval [Q1, Q2] of the input image;
[0013] For the image filtering, the mean filtering method is adopted. Specifically, given an image h(x) with a pixel size of M×M, the gray-scale mean of several pixel points within a preset neighborhood around the point (x, y) is taken as the gray-scale of this point in the enhanced image, and then the above similar operation is performed on the M×M pixel points to construct a new image k(x), and the mathematical expression is:
[0014]
[0015] where x is the horizontal coordinate of the image, y is the vertical coordinate of the image, m is the width of the neighborhood, and n is the height of the neighborhood.
[0016] Furthermore, constructing the C3k2-CaFormer-CGLU module in step (2) to replace the C3k2 module of the original model includes: adopting a 4-stage framework, specifying the use of convolution as the token mixer in the first two stages, and using the attention mechanism in the last two stages to construct CAFormer, and then adding Convolutional GLU (CGLU) to combine convolution and the gating mechanism.
[0017] Furthermore, combining ASF-YOLO and SDI to construct a feature pyramid network in step (2) includes: using the Neck of the weighted bidirectional feature pyramid network ASF-YOLO in combination with the multi-level feature fusion module SDI to form a new Neck network structure;
[0018] The ASF-YOLO model is designed with a scale sequence feature fusion (SSFF) module and a triple feature encoder (TFE) module to fuse the multi-scale feature maps extracted from the backbone in the path aggregation network (PANet) structure, and then a channel and position attention mechanism (CPAM) is designed to integrate the feature information from the SSFF and TFE modules;
[0019] The SSFF combines the global semantic information of images at different scales by normalizing, upsampling, and concatenating multi-scale features into 3D convolutions; the TFE module contains small, medium, and large feature maps to capture the fine spatial information of small objects at different scales.
[0020] In the SDI module, using the hierarchical feature maps generated by the encoder, first apply the spatial and channel attention mechanisms to the feature f of each level i, and the formula is as follows:
[0021]
[0022] where f i 0 represents the feature of the i-th layer, and f i 1 represents the processed feature map of the i-th layer, and represent the parameters of the i-th layer spatial and the i-th layer channel attention respectively;
[0023] Apply a 1×1 convolution to reduce the number of channels of f to c, where c is a hyperparameter, and the resulting feature map is denoted as where H i 、W i and c represent the width, height, and number of channels of f respectively; then send the feature map f i 2 to the decoder. At each decoder level, use f i 2 as the target reference, and then adjust the size of the feature map of each j level to make it have the same resolution as f i 2 The formula is as follows:
[0024]
[0025] where D, I, and U represent adaptive average pooling, identity mapping, and bilinear interpolation respectively, and interpolate f i 2 to the resolution of H i ×W i and 1 ≤ i, j ≤ M;
[0026] Subsequently, apply a 3×3 convolution to smooth each resized feature map The formula is as follows:
[0027]
[0028] where θ ij represents the parameter of the smoothing convolution, represents the j-th smoothed feature map of the i-th layer;
[0029] After adjusting all the feature maps of the i-th layer to the same resolution, the element-wise Hadamard product is applied to all the smoothed feature maps to enhance the features of the i-th layer. The formula is as follows:
[0030]
[0031] Furthermore, the Focaler-Shape-IoU loss function that fuses and creates in step (2) replaces the CIoU loss function of the original model, including: fusing the Shape-IoU and Focaler-IoU techniques, and calculating the loss by linearly mapping the interval to focus on the shape and scale of the bounding box itself. Specifically:
[0032] Let L Focaler-shape-IoU The loss function is defined as:
[0033] L Focaler-Shape-IoU = L ShapeIoU + IoU - L FocalarIoU
[0034] where L Shape-IoU is defined as follows:
[0035] L Shape-IoU = 1 - IoU + distance Shape + 0.5 × Ω Shape
[0036]
[0037]
[0038] where scale is the scale factor, determined according to the proportion of the upper shoe surface defect, ww and hh are the weight coefficients in the horizontal and vertical directions respectively, x c and y c are the horizontal and vertical coordinates of the anchor box, and are the horizontal and vertical coordinates of the target box, w and h are the width and height of the anchor box, w gt and h gt are the width and height of the target box, and c is a constant used to standardize the distance;
[0039] L Focaler-IoU is defined as follows:
[0040]
[0041] L Focaler-IoU = 1 - IoU focaler
[0042] where [d, u] ∈ [0, 1].
[0043] Based on the same inventive concept, the present invention also provides a shoe upper defect detection system based on deep learning, including:
[0044] An initialization module, configured to collect shoe upper defect pictures, classify and label them according to four types of defect types: cracking, glue overflow, breakage, and stains, and construct a data set;
[0045] A model improvement module, configured to improve based on the network framework of the YOLOv11 algorithm to obtain an improved CCAS-YOLO model; the improvement includes: constructing a C3k2-CaFormer-CGLU module to replace the C3k2 module of the original model, combining ASF-YOLO and SDI to construct a feature pyramid network, and fusing and creating a loss function Focaler-Shape-IoU to replace the CIoU loss function of the original model;
[0046] A detection module, configured to use the improved CCAS-YOLO model to detect shoe upper defects in the constructed data set.
[0047] Furthermore, the initialization module further includes processing the collected shoe upper defect pictures, and the preprocessing includes grayscale transformation and image filtering;
[0048] For the grayscale transformation, by redistributing the pixel values of the picture, weakening the background, enhancing the dynamic range, and enhancing the target features. Specifically, taking the linear function as the grayscale mapping relationship between the original grayscale f(x) and the transformed grayscale g(x), the mathematical expression is:
[0049]
[0050] where Q1 and Q2 are grayscale thresholds, where Q1 < Q2, and Z1 and Z2 are the new grayscales corresponding to the input image grayscale value interval [Q1, Q2] in the output image;
[0051] For the image filtering, the method of mean filtering is adopted. Specifically, given an image h(x) with a pixel size of M×M, taking the grayscale mean of several pixel points within a preset neighborhood around the point (x, y) as the grayscale of this point in the enhanced image, and then performing the above similar operations on M×M pixel points to construct a new image k(x), the mathematical expression is:
[0052]
[0053] where x is the horizontal coordinate of the image, y is the vertical coordinate of the image, m is the width of the neighborhood, and n is the height of the neighborhood.
[0054] Furthermore, the model improvement module constructs a C3k2-CaFormer-CGLU module to replace the C3k2 module of the original model, including: adopting a 4-stage framework, specifying to use convolution as the token mixer in the first two stages and the attention mechanism in the last two stages to construct CAFormer, and then adding Convolutional GLU (CGLU) to combine convolution and the gating mechanism.
[0055] Furthermore, the model improvement module combines ASF-YOLO and SDI to construct a feature pyramid network, including: combining the Neck of the weighted bidirectional feature pyramid network ASF-YOLO with the multi-level feature fusion module SDI to form a new Neck network structure;
[0056] The ASF-YOLO model is designed with a Scale Sequence Feature Fusion (SSFF) module and a Triple Feature Encoder (TFE) module to fuse the multi-scale feature maps extracted from the backbone in the Path Aggregation Network (PANet) structure, and then designs a Channel and Position Attention Mechanism (CPAM) to integrate the feature information from the SSFF and TFE modules;
[0057] The SSFF combines the global semantic information of images at different scales by normalizing, upsampling, and concatenating the multi-scale features into 3D convolution; the TFE module contains small, medium, and large feature maps to capture the fine spatial information of small objects at different scales
[0058] In the SDI module, using the hierarchical feature maps generated by the encoder, first apply the spatial and channel attention mechanisms to the feature f of each level i, and the formula is as follows:
[0059]
[0060] where f i 0 represents the feature of the i-th layer, and f i 1 represents the processed feature map of the i-th layer, and respectively represent the parameters of the i-th layer spatial and the i-th layer channel attention;
[0061] Apply a 1×1 convolution to reduce the number of channels of f to c, where c is a hyperparameter, and the resulting feature map is denoted as where H i 、W i and c respectively represent the width, height, and number of channels of f; then send the feature map f i 2 to the decoder, and at each decoder level, use f i2 As a target reference, the feature map size of each j level is subsequently adjusted to be the same as that of f i 2 with the same resolution, and the formula is as follows:
[0062]
[0063] where D, I, and U represent adaptive average pooling, identity mapping, and bilinear interpolation respectively, and f i 2 is interpolated to the resolution of H i ×W i , and 1 ≤ i, j ≤ M;
[0064] Subsequently, a 3×3 convolution is applied to smooth each resized feature map The formula is as follows:
[0065]
[0066] where θ ij represents the parameter of the smoothing convolution, represents the j-th smoothed feature map of the i-th layer;
[0067] After adjusting all the feature maps of the i-th layer to the same resolution, an element-wise Hadamard product is applied to all the smoothed feature maps to enhance the features of the i-th layer, and the formula is as follows:
[0068]
[0069] Furthermore, the model improvement module fuses and creates a loss function Focaler-Shape-IoU to replace the CIoU loss function of the original model, including: fusing Shape-IoU and Focaler-IoU techniques, and calculating the loss by linearly mapping the interval to focus on the shape and scale of the bounding box itself. Specifically:
[0070] Define the L Focaler-shape-IoU loss function as:
[0071] L Focaler-Shape-IoU = L ShapeIoU + IoU - L FocalarIoU
[0072] where L Shape-IoU is defined as follows:
[0073] L Shape-IoU = 1 - IoU + distance Shape + 0.5 × Ω Shape
[0074]
[0075]
[0076] Among them, scale is the scale factor, determined according to the proportion of the upper surface defect, ww and hh are the weight coefficients in the horizontal and vertical directions respectively, and x c and y c are the horizontal and vertical coordinates of the anchor box, and are the horizontal and vertical coordinates of the target box, w and h are the width and height of the anchor box, w gt and h gt are the width and height of the target box, and c is a constant used to standardize the distance;
[0077] L Focaler-IoU is defined as follows:
[0078]
[0079] L Focaler-IoU = 1 - IoU focaler
[0080] where [d, u] ∈ [0, 1].
[0081] Beneficial effects: Compared with the prior art, the present invention constructs an improved CCAS - YOLO model. By constructing a new module C3k2 - CaFormer - CGLU to replace the C3k2 module of the original model, combining ASF - YOLO and SDI to construct a new feature pyramid network, and fusing and creating a new loss function Focaler - Shape - IoU to replace the CIoU loss function of the original model, while ensuring higher accuracy than other models, the computational amount does not increase significantly, improving the detection effect and efficiency of upper surface defects. BRIEF DESCRIPTION OF THE DRAWINGS
[0082] Figure 1 is the flowchart of the method according to the embodiment of the present invention;
[0083] Figure 2 is the schematic diagram of the improved CCAS - YOLO model according to the embodiment of the present invention;
[0084] Figure 3 is the schematic diagram of the CaFormer structure according to the embodiment of the present invention;
[0085] Figure 4 is the schematic diagram of the CaFormer block in the CaFormer structure according to the embodiment of the present invention;
[0086] Figure 5 is the schematic diagram of the CGLU structure according to the embodiment of the present invention;
[0087] Figure 6 This is the ASF-YOLO framework diagram of the embodiment of the present invention;
[0088] Figure 7 This is the schematic diagram of the TFE structure of the embodiment of the present invention;
[0089] Figure 8 This is the schematic diagram of the CPAM structure of the embodiment of the present invention;
[0090] Figure 9 This is the schematic diagram of the SDI structure of the embodiment of the present invention;
[0091] Figure 10 This is the schematic diagram of the Focaler-Shape-IoU structure of the embodiment of the present invention. Detailed implementation manners
[0092] In order to enable those skilled in the art of this technology to better understand the solutions of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application.
[0093] As shown in the attached Figure 1 figure, the method for detecting upper surface defects based on deep learning in this embodiment includes:
[0094] S001: Collect upper surface defect pictures, classify and label them according to four types of defect types: cracking, glue overflow, breakage, and stain, and construct a data set;
[0095] S002: Improve the network framework based on the YOLOv11 algorithm to obtain an improved CCAS-YOLO model; the improvements include: constructing a C3k2-CaFormer-CGLU module to replace the C3k2 module of the original model, constructing a feature pyramid network by combining ASF-YOLO and SDI, and fusing and creating a loss function Focaler-Shape-IoU to replace the CIoU loss function of the original model;
[0096] S003: Use the improved CCAS-YOLO model to detect upper surface defects in the constructed data set.
[0097] Specifically: In step S001, a camera is used to take pictures to collect pictures of shoe upper defects in various directions. These pictures are classified and processed according to four defect types: cracking, glue overflow, damage, and stains. First, the pictures are preprocessed using gray-scale transformation and image filtering, and then labeled using the labelimg data annotation software. Gray-scale transformation mainly achieves this by redistributing the pixel values of the image, weakening the background, enhancing the dynamic range, and making the target features more obvious. The linear transformation of the image uses a linear function as the gray-scale mapping relationship between the original gray-scale f(x) and the transformed gray-scale g(x). Its mathematical expression is:
[0098]
[0099] Usually, noise interference will occur during the acquisition of shoe upper defect information, resulting in unsatisfactory results. These noises will interfere with the feature extraction information of shoe upper defects. Therefore, these noises must be removed to improve the quality of the image. Common filtering methods mainly include mean filtering, median filtering, etc. Mean filtering mainly performs linear filtering within a certain domain, and the effect has an averaging property. Given an image h(x) with a pixel size of M×M, the gray-scale mean of some pixel points (excluding the point (x, y)) within a preset domain around the point (x, y) is taken as the gray-scale of this point in the enhanced image, and then the above similar operations are performed on the M×M pixel points to construct a new image k(x). Its mathematical expression is:
[0100]
[0101] Among them, x is the horizontal coordinate of the image, y is the vertical coordinate of the image, m, n represent the size of the neighborhood, m is the width of the neighborhood, and n is the height of the neighborhood.
[0102] In step S002, the network framework of YOLOv11 algorithm is improved to obtain the improved CCAS-YOLO model. The CCAS-YOLO model is as Figure 2 ;
[0103] S0021 Construct a brand-new module C3k2-CaFormer-CGLU to replace the C3k2 module of the original model: Since the computational complexity of self-attention is quadratic with respect to the number of tokens, using ordinary self-attention in the first two stages would be cumbersome as there are many tokens in these two stages. In contrast, convolution is a local operation with a computational complexity linear with respect to the token length. Our new model adopts a 4-stage framework and specifies using convolution as the token mixer in the first two stages and the attention mechanism in the last two stages to construct CAFormer. Then, CGLU (Convolutional GLU) is added later, which combines convolution and gating mechanism and can selectively pass information channels, improving the effectiveness and flexibility of feature extraction. CAFormer is as Figure 3 and Figure 4 shown, and CGLU is as Figure 5 shown;
[0104] S0022 Combine the Neck of the weighted bidirectional feature pyramid network ASF-YOLO with the multi-level feature fusion module (SDI) to form a brand-new Neck network structure. The ASF-YOLO model improves the accuracy and speed in processing images by combining spatial and scale features. The main idea of the SDI module is to enhance the semantic information and detailed information in the image by integrating the hierarchical feature maps generated by the encoder;
[0105] S00221 The ASF-YOLO model designs a scale sequence feature fusion (SSFF) module and a triple feature encoder (TFE) module to fuse the multi-scale feature maps extracted from the backbone in the path aggregation network (PANet) structure. SSFF combines the global semantic information of images at different scales by normalizing, upsampling, and concatenating multi-scale features into 3D convolution. Therefore, it can effectively process objects of different sizes, orientations, and aspect ratios in the scale space representation to improve object separation. TFE contains small, medium, and large feature maps to capture the fine spatial information of small objects at different scales. These overcome the limitations of FPN in YOLOv11, that is, the correlation between pyramid feature maps cannot be fully utilized through simple summation and concatenation operations. Then, a channel and position attention mechanism (CPAM) is designed to integrate the feature information from the SSFF and TFE modules. This module allows the model to adaptively adjust its attention to the channels and spatial positions related to small targets at different scales. The ASF-YOLO framework diagram is as Figure 6 shown, the TFE structure is as Figure 7 shown, and the CPAM structure is as Figure 8 shown;
[0106] S00222 The SDI module, as Figure 9As shown
[0107] Using the hierarchical feature maps generated by the encoder, we first apply spatial and channel attention mechanisms to the features f at each level i. This process enables the features to integrate local spatial information and global channel information, and the formula is as follows:
[0108]
[0109] where f i 1 represents the feature map after processing at the i-th layer, and represent the parameters of spatial and channel attention at the i-th layer respectively. In addition, we apply 1×1 convolution to reduce the number of channels of f to c, where c is a hyperparameter. The resulting feature map is denoted as where H i , W i and c represent the width, height, and number of channels of f respectively. Then the refined feature map is sent to the decoder. At each decoder level, we use f i 2 as the target reference. Then, we resize the feature map at each j level to have the same resolution as f i 2 The formula is as follows:
[0110]
[0111] where D, I, and U represent adaptive average pooling, identity mapping, and bilinear interpolation respectively, interpolating f i 2 to the resolution of H i ×W i , and 1 ≤ i, j ≤ M.
[0112] After that, 3×3 convolution is applied to smooth each resized feature map The formula is as follows:
[0113]
[0114] where θ ij represents the parameter of the smoothing convolution, represents the j-th smoothed feature map at the i-th layer. After adjusting all the feature maps at the i-th layer to the same resolution, we apply the element-wise Hadamard product to all the resized feature maps to enhance the features at the i-th layer, making them have more semantic information and finer details at the same time. The formula is as follows:
[0115]
[0116] S023 combines and creates a new loss function, the Focaler-Shape-IoU loss function, as Figure 10 shown: By integrating the advantages of Shape-IoU and Focaler-IoU technologies, the shape and scale of the bounding box itself are focused through linear interval mapping to calculate the loss. The specific formula is as follows:
[0117] L Focaler-shape-IoU The loss function is defined as
[0118] L Focaler-Shape-IoU = L ShapeIoU + IoU - L FocalarIoU
[0119] where L Shape-IoU is defined as
[0120] L Shape-IoU = 1 - IoU + distance Shape + 0.5 × Ω Shape
[0121]
[0122]
[0123] where scale is the scale factor, related to the proportion of the upper surface defect, ww and hh are the weight coefficients in the horizontal and vertical directions respectively, and their values are related to the shape of the upper surface defect.
[0124] where L Focaler-IoU is defined as
[0125]
[0126] L Focaler-IoU = 1 - IoU focaler
[0127] where [d, u] ∈ [0, 1], by adjusting the values of d and u, focusing on the regression samples of the upper surface defect to make the bounding box regression more accurate and the detection of the upper surface defect more accurate.
[0128] In step S003, the improved model is used to detect the upper surface defect of the data set processed in step S001;
[0129] This embodiment further includes step S004. To verify the effectiveness of the CCAS-YOLO model, the detection structures of various mainstream object detection networks are reproduced on the upper surface defect data set and compared with our model. The comparison results are shown in Table 1.
[0130] Table 1 Comparison table of detection performance of upper surface defect data set
[0131]
[0132] This method selects the corresponding deep learning model and uses the stochastic gradient descent method for training. The learning rate is 0.01, the batch size is set to 8, the number of epochs is set to 300, and the weight decay is set to 0.0005. After completing the model training, the detection model metrics are obtained. The optimal detection model is selected based on performance evaluation metrics such as accuracy, recall, Map50, Params (M), and GFLOPs. Then, the optimal detection model is selected based on the performance evaluation metrics, and the shoe upper defects are detected using the object detection head.
[0133] The precision expression is:
[0134]
[0135] The recall expression is:
[0136]
[0137] The accuracy expression is:
[0138]
[0139] In the formula, TP is the number of positive samples correctly identified as positive samples, FN is the number of positive samples misidentified as negative samples, and FP is the number of negative samples misidentified as positive samples.
[0140] mAP represents the mean of APs for all classes in the entire dataset, and the calculation formula is as follows.
[0141]
[0142] In the formula, AP is the area under the PR curve; mAP is the average of APs calculated after calculating AP for each class, and N is the total number of each class.
[0143] The experimental environment of this method is based on the PyTorch deep learning framework and runs on the Windows operating system. The NVIDIA GeForce RTX 4060 GPU is used for computing acceleration, and the detection model metrics are obtained after completing the model training. Experimental environment parameters: initial learning rate 0.01, number of training rounds 200, optimizer SGD, momentum coefficient 0.937, weight decay 0.0005, and input image resolution 640x640.
[0144] Table 1 compares the performance of different models on the upper defect dataset, including precision, recall, mAP, etc. These models show different performances on these metrics. Compared with other YOLO models, Map is the highest, indicating that the optimization of this method is effective in improving the detection accuracy of upper defects.
[0145] Based on the same inventive concept, this embodiment also provides a deep learning-based upper defect detection system, including:
[0146] An initialization module, configured to collect upper defect pictures, classify and label them according to four types of defect types: cracking, glue overflow, breakage, and stain, and construct a dataset;
[0147] A model improvement module, configured to improve based on the network framework of the YOLOv11 algorithm to obtain an improved CCAS-YOLO model; the improvements include: constructing a C3k2-CaFormer-CGLU module to replace the C3k2 module of the original model, combining ASF-YOLO and SDI to construct a feature pyramid network, and fusing and creating a loss function Focaler-Shape-IoU to replace the CIoU loss function of the original model;
[0148] A detection module, configured to use the improved CCAS-YOLO model to detect upper defects in the constructed dataset.
[0149] Furthermore, the initialization module further includes processing the collected upper defect pictures, and the preprocessing includes grayscale transformation and image filtering;
[0150] For the grayscale transformation, by redistributing the pixel values of the picture, weakening the background, enhancing the dynamic range, and enhancing the target features. Specifically, taking the linear function as the grayscale mapping relationship between the original grayscale f(x) and the transformed grayscale g(x), the mathematical expression is:
[0151]
[0152] where Q1 and Q2 are grayscale thresholds, where Q1 < Q2, and Z1 and Z2 are the new grayscales corresponding to the input image grayscale value interval [Q1, Q2] in the output image;
[0153] For the image filtering, the method of mean filtering is adopted. Specifically, given an image h(x) with a pixel size of M×M, taking the grayscale mean of several pixel points within a preset neighborhood around the point (x, y) as the grayscale of this point in the enhanced image, and then performing the above similar operations on M×M pixel points to construct a new image k(x), the mathematical expression is:
[0154]
[0155] Among them, x is the horizontal coordinate of the image, y is the vertical coordinate of the image, m is the width of the neighborhood, and n is the height of the neighborhood.
[0156] Furthermore, the model improvement module constructs a C3k2-CaFormer-CGLU module to replace the C3k2 module of the original model, including: adopting a 4-stage framework, specifying to use convolution as the token mixer in the first two stages and the attention mechanism in the last two stages to construct CAFormer, and then adding Convolutional GLU (CGLU) to combine convolution and the gating mechanism.
[0157] Furthermore, the model improvement module combines ASF-YOLO and SDI to construct a feature pyramid network, including: combining the Neck of the weighted bidirectional feature pyramid network ASF-YOLO with the multi-level feature fusion module SDI to form a new Neck network structure;
[0158] The ASF-YOLO model is designed with a scale sequence feature fusion (SSFF) module and a triple feature encoder (TFE) module to fuse the multi-scale feature maps extracted from the backbone in the path aggregation network (PANet) structure, and then designs a channel and position attention mechanism (CPAM) to integrate the feature information from the SSFF and TFE modules;
[0159] The SSFF combines the global semantic information of images at different scales by normalizing, upsampling, and concatenating the multi-scale features into 3D convolution; the TFE module contains small, medium, and large feature maps to capture the fine spatial information of small objects at different scales
[0160] In the SDI module, using the hierarchical feature maps generated by the encoder, first apply the spatial and channel attention mechanisms to the feature f of each level i, and the formula is as follows:
[0161]
[0162] where f i 0 represents the feature of the i-th layer, and f i 1 represents the processed feature map of the i-th layer, and respectively represent the parameters of the i-th layer spatial and the i-th layer channel attention;
[0163] Apply a 1×1 convolution to reduce the number of channels of f to c, where c is a hyperparameter, and the resulting feature map is denoted as where H i 、W iWidth, height, and number of channels of f are represented by a, b, and c respectively; then the feature map f i 2 is sent to the decoder. At each decoder layer, f i 2 is used as the target reference, and then the feature map size of each j layer is adjusted to be the same as that of f i 2 The resolution is the same, and the formula is as follows:
[0164]
[0165] where D, I, and U represent adaptive average pooling, identity mapping, and bilinear interpolation respectively. Interpolate f i 2 to the resolution of H i ×W i , and 1 ≤ i, j ≤ M;
[0166] Subsequently, a 3×3 convolution is applied to smooth each resized feature map The formula is as follows:
[0167]
[0168] where θ ij represents the parameter of the smoothing convolution, represents the j-th smoothed feature map of the i-th layer;
[0169] After adjusting all the feature maps of the i-th layer to the same resolution, an element-wise Hadamard product is applied to all the smoothed feature maps to enhance the features of the i-th layer. The formula is as follows:
[0170]
[0171] Furthermore, the model improvement module fuses and creates a loss function Focaler-Shape-IoU to replace the CIoU loss function of the original model, including: fusing Shape-IoU and Focaler-IoU techniques, and calculating the loss by linearly mapping the interval to focus on the shape and scale of the bounding box itself. Specifically:
[0172] Define the L Focaler-shape-IoU loss function as:
[0173] L Focaler-Shape-IoU = L ShapeIoU + IoU - L FocalarIoU
[0174] where L Shape-IoU is defined as follows:
[0175] LShape-IoU = 1 - IoU + distance Shape + 0.5 × Ω Shape
[0176]
[0177] Among them, scale is the scale factor, determined according to the proportion of the upper shoe surface defect, ww and hh are the weight coefficients in the horizontal and vertical directions respectively, x c and y c are the horizontal and vertical coordinates of the anchor box, and are the horizontal and vertical coordinates of the target box, w and h are the width and height of the anchor box, w gt and h gt are the width and height of the target box, c is a constant for normalizing the distance;
[0178] L Focaler-IoU is defined as follows:
[0179]
[0180] L Focaler-IoU = 1 - IoU focaler
[0181] Among them, [d, u] ∈ [0, 1].
Claims
1. A method for detecting shoe upper defects based on deep learning, characterized in that: Including: (1) Collect pictures of shoe upper defects, classify and label them according to four defect types: cracking, glue overflow, damage and stains, and construct a data set; (2) Improve the network framework based on the YOLOv11 algorithm to obtain the improved CCAS-YOLO model; The improvements include: constructing a C3k2-CaFormer-CGLU module to replace the C3k2 module of the original model, combining ASF-YOLO and SDI to construct a feature pyramid network, and fusing and creating a loss function Focaler-Shape-IoU to replace the CIoU loss function of the original model; (3) Use the improved CCAS-YOLO model to detect shoe upper defects in the constructed data set.
2. The method for detecting upper defects based on deep learning according to claim 1, characterized in that: Step (1) also includes processing the collected pictures of shoe upper defects, and the preprocessing includes grayscale transformation and image filtering; For the grayscale transformation, by redistributing the pixel values of the picture, the background is weakened, the dynamic range is enhanced, and the target features are enhanced. Specifically, a linear function is used as the grayscale mapping relationship between the original grayscale f(x) and the transformed grayscale g(x), and the mathematical expression is: where Q1 and Q2 are grayscale thresholds, where Q1 < Q2, and Z1 and Z2 are the new grayscales corresponding to the input image grayscale value interval [Q1, Q2] in the output image; For the image filtering, the mean filtering method is adopted. Specifically, given an image h(x) with a pixel size of M×M, the grayscale mean of several pixel points within a preset neighborhood around the point (x, y) is taken as the grayscale of this point in the enhanced image, and then the above similar operation is performed on the M×M pixel points to construct a new image k(x), and the mathematical expression is: where x is the horizontal coordinate of the image, y is the vertical coordinate of the image, m is the width of the neighborhood, and n is the height of the neighborhood.
3. The method for detecting upper defects based on deep learning according to claim 1, characterized in that: In step (2), constructing the C3k2-CaFormer-CGLU module to replace the C3k2 module of the original model includes: adopting a 4-stage framework, specifying to use convolution as the token mixer in the first two stages and the attention mechanism in the last two stages to construct CAFormer, and then adding Convolutional GLU (CGLU) to combine convolution and the gating mechanism.
4. The method for detecting upper defects based on deep learning according to claim 1, characterized in that: In step (2), combining ASF-YOLO and SDI to construct a feature pyramid network includes: combining the Neck of the weighted bidirectional feature pyramid network ASF-YOLO with the multi-level feature fusion module SDI to form a new Neck network structure; The ASF-YOLO model is designed with a scale sequence feature fusion (SSFF) module and a triple feature encoder (TFE) module to fuse the multi-scale feature maps extracted from the backbone in the path aggregation network (PANet) structure, and then a channel and position attention mechanism (CPAM) is designed to integrate the feature information from the SSFF and TFE modules; The SSFF combines the global semantic information of images at different scales by normalizing, upsampling, and concatenating multi-scale features into 3D convolutions; the TFE module contains small, medium, and large feature maps to capture the fine spatial information of small objects at different scales. In the SDI module, using the hierarchical feature maps generated by the encoder, a spatial and channel attention mechanism is first applied to the feature f at each level i, as follows: Among them, f i 0 represents the features of the i-th layer, f i 1 represents the feature map after processing of the i-th layer, and Represent the parameters of the i-th layer spatial and i-th layer channel attention respectively; Applying 1×1 convolution reduces the number of channels of f to c, where c is a hyperparameter, and the resulting feature map is represented as Among them, H i , W i and c represent the width, height and number of channels of f respectively; then the feature map f i 2 Sent to the decoder, at each decoder level, using f i 2 As the target reference, the feature map size of each j level is then adjusted to be consistent with f i 2 With the same resolution, the formula is as follows: Among them, D, I and U represent adaptive average pooling, identity mapping and bilinear interpolation respectively. i 2 Interpolate to H i ×W i The resolution of , and 1≤i,j≤M; A 3×3 convolution is then applied to smooth each resized feature map. The formula is as follows: Among them, θ ij represents the parameters of smoothed convolution, Represents the j-th smoothed feature map of the i-th layer; After adjusting all the feature maps at the i-th layer to the same resolution, an element-wise Hadamard product is applied to all the smoothed feature maps to enhance the features at the i-th layer, as follows:
5. The method for detecting upper defects based on deep learning according to claim 1, characterized in that: In step (2), the fused and created loss function Focaler-Shape-IoU replaces the CIoU loss function of the original model, including: fusing the Shape-IoU and Focaler-IoU techniques, and calculating the loss by linearly mapping the interval to focus on the shape and scale of the bounding box itself, specifically: L Focaler-shape-IoU The loss function is defined as: L Focaler-Shape-IoU =L ShapeIoU +IoU-L FocalarIoU Among them, L Shape-IoU The definition is as follows: L Shape-IoU =1-IoU+distance Shape +0.5×Ω Shape Among them, scale is the proportional factor, which is determined according to the proportion of the upper defects, ww and hh are the weight coefficients in the horizontal and vertical directions respectively, and x c With y c is the horizontal and vertical coordinates of the anchor box, and is the horizontal and vertical coordinates of the target box, w and h are the width and height of the anchor box, w gt With h gt is the width and height of the target box, c is a constant used to standardize the distance; L Focaler-IoU The definition is as follows: L Focaler-IoU =1-IoU focaler where [d, u] ∈ [0, 1].
6. A shoe upper defect detection system based on deep learning, characterized in that: Including: An initialization module for collecting upper shoe surface defect pictures, classifying and labeling them according to four defect types: cracking, glue overflow, breakage, and stain, and constructing a data set; A model improvement module for improving based on the network framework of the YOLOv11 algorithm to obtain the improved CCAS-YOLO model; the improvements include: constructing a C3k2-CaFormer-CGLU module to replace the C3k2 module of the original model, combining ASF-YOLO and SDI to construct a feature pyramid network, and fusing and creating a loss function Focaler-Shape-IoU to replace the CIoU loss function of the original model; A detection module for using the improved CCAS-YOLO model to detect upper shoe surface defects in the constructed data set.
7. The shoe upper defect detection system based on deep learning according to claim 6, characterized in that: The initialization module also includes processing the collected upper shoe surface defect pictures, and the preprocessing includes grayscale transformation and image filtering; For the grayscale transformation, by redistributing the pixel values of the picture, the background is weakened, the dynamic range is enhanced, and the target features are enhanced. Specifically, a linear function is used as the grayscale mapping relationship between the original grayscale f(x) and the transformed grayscale g(x), and the mathematical expression is: where Q1 and Q2 are grayscale thresholds, where Q1 < Q2, and Z1 and Z2 are the new grayscales corresponding to the input image grayscale value interval [Q1, Q2] in the output image; For the image filtering, the mean filtering method is adopted. Specifically, given an image h(x) with a pixel size of M × M, the grayscale mean of several pixel points in a preset neighborhood around the point (x, y) is taken as the grayscale of this point in the enhanced image, and then the above similar operation is performed on the M × M pixel points to construct a new image k(x), and the mathematical expression is: where x is the horizontal coordinate of the image, y is the vertical coordinate of the image, m is the width of the neighborhood, and n is the height of the neighborhood.
8. The shoe upper defect detection system based on deep learning according to claim 6, characterized in that: The model improvement module constructs a C3k2-CaFormer-CGLU module to replace the C3k2 module of the original model, including: adopting a 4-stage framework, specifying the use of convolution as a token mixer in the first two stages, using the attention mechanism in the last two stages, building CAFormer, and then adding Convolutional GLU (CGLU) to combine convolution and gating mechanisms.
9. The shoe upper defect detection system based on deep learning according to claim 6, characterized in that: The model improvement module combines ASF-YOLO and SDI to build a feature pyramid network, including: using the Neck of the weighted bidirectional feature pyramid network ASF-YOLO and combining it with the multi-level feature fusion module SDI to form a new Neck network structure; The ASF-YOLO model is designed with a scale sequence feature fusion (SSFF) module and a triple feature encoder (TFE) module to fuse the multi-scale feature maps extracted from the backbone in a path aggregation network (PANet) structure, and then a channel and position attention mechanism (CPAM) is designed to integrate the feature information from the SSFF and TFE modules; The SSFF combines the global semantic information of images of different scales by normalizing, upsampling and concatenating multi-scale features into 3D convolutions; the TFE module contains small, medium and large feature maps to capture the fine spatial information of small objects at different scales. In the SDI module, using the hierarchical feature map generated by the encoder, we first apply the spatial and channel attention mechanism to the feature f of each level i, as follows: Among them, f i 0 represents the features of the i-th layer, f i 1 represents the feature map after processing of the i-th layer, and Represent the parameters of the i-th layer spatial and i-th layer channel attention respectively; Applying 1×1 convolution reduces the number of channels of f to c, where c is a hyperparameter, and the resulting feature map is represented as Among them, H i , W i and c represent the width, height and number of channels of f respectively; then the feature map f i 2 Sent to the decoder, at each decoder level, using f i 2 As the target reference, the feature map size of each j level is then adjusted to be consistent with f i 2 With the same resolution, the formula is as follows: Among them, D, I and U represent adaptive average pooling, identity mapping and bilinear interpolation respectively. i 2 Interpolate to H i ×W i The resolution of , and 1≤i,j≤M; A 3×3 convolution is then applied to smooth each resized feature map. The formula is as follows: Among them, θ ij represents the parameters of smoothed convolution, Represents the j-th smoothed feature map of the i-th layer; After adjusting all the feature maps of the i-th layer to the same resolution, the element-wise Hadamard product is applied to all the smoothed feature maps to enhance the features of the i-th layer. The formula is as follows:
10. The shoe upper defect detection system based on deep learning according to claim 6, characterized in that: The model improvement module integrates and creates the loss function Focaler-Shape-IoU to replace the CIoU loss function of the original model, including: integrating Shape-IoU and Focaler-IoU technologies, and calculating the loss by focusing on the shape and scale of the bounding box itself through linear interval mapping, specifically: L Focaler-shape-IoU The loss function is defined as: L Focaler-Shape-IoU =L ShapeIoU +IoU-L FocalarIoU Among them, L Shape-IoU The definition is as follows: L Shape-IoU =1-IoU+distance Shape +0.5×Ω Shape Among them, scale is the proportional factor, which is determined according to the proportion of the upper defects, ww and hh are the weight coefficients in the horizontal and vertical directions respectively, and x c With y c is the horizontal and vertical coordinates of the anchor box, and is the horizontal and vertical coordinates of the target box, w and h are the width and height of the anchor box, w gt With h gt is the width and height of the target box, c is a constant used to standardize the distance; L Focaler-IoU The definition is as follows: L Focaler-IoU =1-IoU focaler where [d,u]∈[0,1].
Citation Information
Cited By
Target detection model and method based on deep learning
CN121095726A