Insole blanking material detection method based on improved YOLOv5

By improving the YOLOv5 network model, combining FasterNet block, SE attention mechanism and Focal-EIoU loss function, the problems of inefficient and high error rate of traditional insole punching materials are solved, and more efficient, accurate and real-time detection effects are achieved.

CN120014231APending Publication Date: 2025-05-16ZHEJIANG SCI-TECH UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510075996.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Traditional insole punching material detection methods are inefficient and have high error rates, especially when dealing with complex contours or mass production, which can easily lead to waste of materials or inconsistent product quality.

Method used

The insole punching material detection method based on improved YOLOv5 is adopted. By using FasterNet block to replace the Bottleneckmu structure in the C3 module of the YOLOv5 network model, and the SE attention mechanism is introduced in the feature channel, and the Focal-EIoU loss function is used instead of the original CIoU loss function, the LightCut-YOLOv5 network model is obtained.

Benefits of technology

It improves the accuracy and real-time detection, enhances the model's adaptability to complex textures and lighting changes, reduces false positive rates, and reduces hardware resource dependence, and is suitable for embedded devices with limited resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014231A_ABST
    Figure CN120014231A_ABST
Patent Text Reader

Abstract

The invention provides an improved YOLOv5-based insole blanking material detection method, which comprises the following steps of: generating a new insole blanking material image by using data enhancement to expand the number of data sets and obtain more data samples, then marking an insole blanking material target, and dividing a training set, a test set and a verification set according to a preset proportion, so as to improve the detection accuracy of the insole blanking material. Inputting the training set to the input end of the LightCut-YOLOv5 network, and realizing faster and more efficient feature extraction and fusion of the feature map through a C3F module, an SE module and an SPPF module in the backbone network, so that the model can dynamically adjust the importance of different features, more efficiently and accurately capture key features in the feature map, and the accuracy of feature extraction is improved. And the target detection precision and speed are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a method for detecting insole blanking materials based on improved YOLOv5. Background Art

[0002] Insole punching material detection is a crucial part of the insole manufacturing process. Its purpose is to accurately identify and locate the position and shape of the material during the cutting process to ensure cutting accuracy and reduce material waste. In the traditional insole punching process, it mainly relies on manual visual inspection and manual adjustment of cutting machine parameters. This method is inefficient and has a high error rate, especially when dealing with complex contours or mass production, which can easily lead to material waste or inconsistent product quality. In recent years, with the rapid development of industrial automation technology, machine vision technology has been gradually introduced into cutting machines to replace manual inspection. However, traditional machine vision methods are mostly based on template matching or simple feature extraction algorithms. Due to the robustness of the algorithm, they perform poorly when dealing with different materials, colors, textures and lighting changes. In addition, traditional methods are still unable to cope with insole punching scenarios that require high detection accuracy and real-time performance. Summary of the invention

[0003] In view of this, the purpose of the present invention is to provide an insole blanking material detection method based on improved YOLOv5, which can effectively, accurately and quickly perform preliminary positioning of the insole blanking material.

[0004] In order to solve the above technical problems, the technical solution of the present invention is: a method for detecting insole blanking materials based on improved YOLOv5, in which the Bottleneckmu structure is replaced by FasterNet block in the C3 module of the YOLOv5 network model; and the SE attention mechanism is introduced in the feature channel of the YOLOv5 model; the Focal-EIoU loss function is used to replace the original CIoU loss function to obtain the LightCut-YOLOv5 network model, and the following steps are also included:

[0005] S1, collect insole cutting material images through the camera and build a data set;

[0006] S2. Using data augmentation to expand the number of insole cutting material images in the data set, preprocessing the expanded data set, annotating it, and dividing it into a training set, a test set, and a validation set according to a preset ratio;

[0007] S3. Build the LightCut-YOLOv5 network model;

[0008] S4, input the training set in step S2 into the LightCut-YOLOv5 network model, and then output the coordinates of the insole punching material after detection by the LightCut-YOLOv5 network model;

[0009] S4.1, input the training set in step S2 into the C3F module of the backbone network, the input image features are first convolved through the channel attention module, and then convolved through the spatial attention module to obtain the output feature map of the output backbone network;

[0010] S4.2, passing the output feature map of the backbone network through the small target detection layer, fusing the output feature map of the top layer with the output feature map of the second layer in the backbone network, and outputting a fused feature map;

[0011] S4.3, detecting the fused feature map and outputting the coordinates of the detected insole material boundary box;

[0012] S5. Evaluate the LightCut-YOLOv5 network model in step S4, repeat steps S1 to S4, determine the insole punching material image input model, input the insole punching material image to be located into the LightCut-YOLOv5 network model for detection, and output the detected insole material boundary box coordinates.

[0013] To implement the above technical solution, new insole punching material images are generated by using data enhancement, thereby expanding the number of data sets, which can effectively alleviate the problem of lack of insole punching material images and obtain more data samples; by dividing the expanded data samples into training sets, test sets and validation sets in proportion, it is convenient to input the training set into the LightCut-YOLOv5 network model in the future, so as to detect the insole punching material images and output the insole material boundary box coordinates; the C3 module in the YOLOv5 network model is lightweighted by FasterNet block, and FasterNet Block is an efficient and lightweight network structure. By simplifying the computational complexity and reducing the number of parameters, the model's inference time and memory usage are greatly reduced. The real-time performance of the model is improved: the improved YOLOv5 network can detect insole punching materials in real time in a high-speed production line environment. The lightweight design reduces the demand for high-performance hardware and facilitates deployment in resource-limited embedded devices or industrial equipment. The SE attention mechanism is introduced. The SE attention mechanism dynamically weights the importance of feature channels, enabling the network to more effectively capture key features while suppressing redundant information. The feature extraction capability is enhanced. In scenarios where the insole material has complex textures and more details, the SE attention mechanism helps the model focus on key areas more accurately. Improve the robustness of detection; improve the adaptability of the model and enhance the network's adaptability to complex environmental factors such as lighting changes and material differences; replace the original CIoU with the Focal-EIoU loss function. While considering the overlap between the predicted box and the true box, Focal-EIoU introduces a penalty for the difference in aspect ratio, and reduces the weight of easy-to-classify samples through the focus mechanism, focusing on optimizing the detection of difficult samples; improve positioning accuracy, more accurately handle the difference in aspect ratio of the predicted box, and adapt to the diverse shapes of insoles; optimize edge conditions, better solve difficult-to-detect situations such as small defects and complex contours, and improve overall detection performance; reduce false alarm rate, by focusing on difficult-to-detect samples and reducing over-optimization of simple samples, thereby reducing false alarms and missed detections.

[0014] As a preferred solution of the present invention, in step S2, the insole image data is enhanced by a Mosaic function.

[0015] By implementing the above technical solution, multiple images can be randomly combined to increase the diversity and robustness of samples in the data set.

[0016] As a preferred solution of the present invention, the preprocessing in step S2 includes correction and normalization processing.

[0017] Implementing the above technical solution makes it easier for images to meet the requirements of the input LightCut-YOLOv5 network.

[0018] As a preferred embodiment of the present invention, the labeling method in step S2 includes using LabelImg labeling software to label each target in the insole punching material in the image, and dividing the training set, test set and validation set in a ratio of 8:1:1.

[0019] Implementing the above technical solution makes it easier to label and create a database, while also ensuring that the training set has a sufficient amount of data.

[0020] As a preferred solution of the present invention, the LightCut-YOLOv5 network model in step S3 includes an input end, a backbone network, a neck network and an output end;

[0021] The input end passes through the backbone network and the neck network respectively and then outputs through the output end;

[0022] The backbone network includes Conv module, C3F module, SE module and SPPF module;

[0023] The Conv module includes a convolutional layer, a batch normalization layer, and a ReLu activation function;

[0024] The C3F module includes a first path and a second path, the first path is composed of a Conv module and more than two Faster_Block modules, and the second path is composed of a Pconv module and a PWconv;

[0025] The SPPF module includes the Conv module and the Maxpool module;

[0026] The CBAM module includes a channel attention module and a spatial attention module;

[0027] The neck network contains multiple C3F modules for further feature extraction and fusion.

[0028] Implementing the above technical solution can significantly improve the model's attention and sensitivity to the insole punching material, thereby better extracting fusion features.

[0029] As a preferred solution of the present invention, the Focal-EloU loss function adopts the FocalLoss idea to give higher weights to difficult-to-detect samples, and combines the EloU method to make the aspect ratio of the predicted box more matched with the actual object, so as to achieve more accurate bounding box regression; its formula is as follows:

[0030]

[0031] Among them, IOU = |A∩B| / |A∪B|, γ is the control calculation value, w c and h care the width and height of the minimum bounding box, b and b gt represents the center coordinates of the predicted box and the real box, p(.) represents the Euclidean distance of the center point coordinates, and w gt and h gt is the width and height of the real box, w and h are the width and height of the predicted box.

[0032] To implement the above technical solution, higher weights are given to samples that are difficult to detect, and the EloU method is used to make the aspect ratio of the predicted box more closely match the actual object, so as to achieve more accurate bounding box regression. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 It is a flow chart of the working method of the present invention;

[0034] Figure 2 This is a schematic diagram of the CF3 module structure in the present invention;

[0035] Figure 3 It is a schematic diagram of the structure of the LightCut-YOLOv5 network in the present invention. DETAILED DESCRIPTION

[0036] The specific implementation modes of the present invention are further described below in conjunction with the accompanying drawings to make the technical solutions of the present invention easier to understand and grasp.

[0037] A method for detecting insole blanking materials based on improved YOLOv5 includes a YOLOv5 network model. In the C3 module of the YOLOv5 network model, a Bottleneckmu structure is replaced by a FasterNet block to achieve lightweight. A SE attention mechanism is introduced into the feature channel of the YOLOv5 model to enhance the model's attention to key features by adaptively learning channel feature weights. A Focal-EloU loss function is used to replace the original CloU loss function to improve the bounding box prediction performance and the aspect ratio difference between the predicted box and the actual box. The YOLOv5 network model is obtained, and the improved YOLOv5 network model is a LightCut-YOLOv5 network model.

[0038] In addition, according to the image characteristics of insole cutting materials, an image preprocessing method is integrated, including background removal, edge enhancement and contrast adjustment.

[0039] The specific steps include:

[0040] S1, collect insole cutting material images through the camera and build a data set;

[0041] S2. Using data augmentation to expand the number of insole cutting material images in the data set, preprocessing the expanded data set, annotating it, and dividing it into a training set, a test set, and a validation set according to a preset ratio;

[0042] S3. Build the LightCut-YOLOv5 network model;

[0043] S4, input the training set in step S2 into the LightCut-YOLOv5 network model, and then output the coordinates of the insole punching material after detection by the LightCut-YOLOv5 network model;

[0044] S4.1, input the training set in step S2 into the C3F module of the backbone network, the input image features are first convolved through the channel attention module, and then convolved through the spatial attention module to obtain the output feature map of the output backbone network;

[0045] S4.2, passing the output feature map of the backbone network through the small target detection layer, fusing the output feature map of the top layer with the output feature map of the second layer in the backbone network, and outputting a fused feature map;

[0046] S4.3, detecting the fused feature map and outputting the coordinates of the detected insole material boundary box;

[0047] S5. Evaluate the LightCut-YOLOv5 network model in step S4, repeat steps S1 to S4, determine the insole punching material image input model, input the insole punching material image to be located into the LightCut-YOLOv5 network model for detection, and output the detected insole material boundary box coordinates.

[0048] The LightCut-YOLOv5 network model in step S3 includes an input end, a backbone network, a neck network, and an output end. The input end passes through the backbone network and the neck network respectively and then outputs through the output end.

[0049] Among them, the backbone network includes Conv module, C3F module, SE module and SPPF module.

[0050] The Conv module includes a convolutional layer, a batch normalization layer, and a ReLu activation function.

[0051] The C3F module includes a first path and a second path. The first path consists of a Conv module and more than two Faster_Block modules; the second path consists of a Pconv module and a PWconv.

[0052] The SPPF module includes a Conv module and a Maxpool module.

[0053] The CBAM module includes a channel attention module and a spatial attention module.

[0054] Multiple C3F modules are included in the neck network for further feature extraction and fusion.

[0055] In step S2, data enhancement is performed through the Mosaic function, and then the number of images of insole punching materials in the data set is expanded by data enhancement. The expanded data set is corrected and normalized, and then Labelmg annotation software is used to annotate each target in the insole punching material in the image, and the training set, test set, and validation set are divided into training set, test set, and validation set in a ratio of 8:1:1.

[0056] The FasterNet structure is used to replace the C3 module Bottleneck in the YOLOv5 network model, ensuring that the number of model parameters and calculations are reduced by simplifying the convolution operation and feature extraction process, thereby improving the real-time detection performance.

[0057] The SE attention mechanism is embedded in the network channel; channel weighting improves the model's sensitivity to small defects and edge information, thereby enhancing the detection accuracy of insole cutting materials.

[0058] The Focal-EloU loss function is used to replace the original CloU loss function, giving higher weights to samples that are difficult to detect, and combined with the EloU method to make the aspect ratio of the predicted box more consistent with the actual object, so as to achieve more accurate bounding box regression. The formula is as follows:

[0059]

[0060] Among them, IOU = |A∩B| / |A∪B|, γ is the control calculation value, w c and h c are the width and height of the minimum bounding box, b and b gt represents the center coordinates of the predicted box and the real box, p(.) represents the Euclidean distance of the center point coordinates, and w gt and h gt is the width and height of the real box, w and h are the width and height of the predicted box.

[0061] The insole blanking material detection model of the LightCut-YOLOv5 network model constructed by training is as follows: the network is deployed using the Pytorch deep learning framework, based on the Window operating system, and written in Python. The server computer configuration used in the model experiment is: NVIDIA RTX A5000 graphics card, pytorch version 1.10.0, cuda version 11.4, cpu is Intel(R) Xeon(R) Platinum 8352V CPU@2.10GHz, and the machine has 256GB of RAM. The number of iterations of model training is set to 200, the batch size is set to 64, the IOU threshold is 0.5, and the initial learning rate is 0.01. In order to simulate the actual reasoning time of the industrial computer as much as possible, the onnx model, image reading, and image processing use ordinary computers and only use cpu reasoning. The ordinary computer configuration is as follows: Intel(R) Core(TM) i5-13500H, with 16GB of RAM.

[0062] The insole blanking material detection model of the trained LightCut-YOLOv5 network model is used to detect and locate the target on the validation set, and the detection effect of the model is evaluated. Specifically, the average accuracy mAP@.5, the number of parameters (106), the amount of calculation (GFLOPs), and the inference time (ms) of the actual onnx model are used to evaluate the detection effect and performance of the model. Among them, AP is the area precision (Precision, P) under the recall rate Recall and precision rate Precision curve: It indicates the proportion of the number of correct targets detected by the model to the total number of targets, reflecting the accuracy of the model in target detection. The formula is as follows:

[0063]

[0064] Among them, TP is the number of correct detections; FP is the number of false detections; TP+FP represents the set of true boxes; FN means that the correct object is detected as a false object; TP+FN is all the labeled true samples; K is the total number of categories, and APi is the AP value of the i-th category. And the ablation experiment verifies the positioning of each module for the yolov5 model. The ablation experiment results are as follows:

[0065] Table 1 Ablation experiment results

[0066]

[0067] From the ablation experiment results table, we can see that replacing the C3 module with the C3F module can reduce the amount of parameters and calculations, but the accuracy has a certain degree of decline, but it does not have much impact on the subsequent positioning steps to be done in this article. Although the subsequent addition of the SE module brings a certain amount of parameters, it has a more obvious improvement in the overall accuracy. The last added Focal_EIOU brings a slight improvement in the selection box positioning. In the end, compared with the original yolov5 model, the accuracy dropped by 0.3%. In actual engineering, the image processing time requirements are relatively high. This article compares the inference time of the deep learning model through an ordinary computer. After the lightweight model, the inference time can be significantly reduced, about 7%. Although the accuracy has dropped to a certain extent, it is generally within an acceptable range.

[0068] Of course, the above are only typical examples of the present invention. In addition, the present invention may also have many other specific implementations. All technical solutions formed by equivalent replacement or equivalent transformation fall within the scope of protection required by the present invention.

Claims

1. A method for detecting insole blanking materials based on improved YOLOv5, characterized in that: In the C3 module of the YOLOv5 network model, the Bottleneckmu structure is replaced by FasterNet block; and the SE attention mechanism is introduced in the feature channel of the YOLOv5 model; the Focal-EIoU loss function is used to replace the original CIoU loss function to obtain the LightCut-YOLOv5 network model, which also includes the following steps: S1, collect insole cutting material images through the camera and build a data set; S2. Using data augmentation to expand the number of insole cutting material images in the data set, preprocessing the expanded data set, annotating it, and dividing it into a training set, a test set, and a validation set according to a preset ratio; S3. Build the LightCut-YOLOv5 network model; S4, input the training set in step S2 into the LightCut-YOLOv5 network model, and then output the coordinates of the insole punching material after detection by the LightCut-YOLOv5 network model; S4.1, input the training set in step S2 into the C3F module of the backbone network, the input image features are first convolved through the channel attention module, and then convolved through the spatial attention module to obtain the output feature map of the output backbone network; S4.2, passing the output feature map of the backbone network through the small target detection layer, fusing the output feature map of the top layer with the output feature map of the second layer in the backbone network, and outputting a fused feature map; S4.3, detecting the fused feature map and outputting the coordinates of the detected insole material boundary box; S5. Evaluate the LightCut-YOLOv5 network model in step S4, repeat steps S1 to S4, determine the insole punching material image input model, input the insole punching material image to be located into the LightCut-YOLOv5 network model for detection, and output the detected insole material boundary box coordinates.

2. The insole blanking material detection method based on improved YOLOv5 according to claim 1 is characterized in that: In step S2, the insole image data is enhanced by a Mosaic function.

3. The insole blanking material detection method based on improved YOLOv5 according to claim 2 is characterized in that: The preprocessing in step S2 includes correction and normalization.

4. The insole blanking material detection method based on improved YOLOv5 according to claim 3 is characterized in that: The labeling method in step S2 includes using Label Img labeling software to label each target in the insole punching material in the image, and dividing the training set, test set and verification set according to the ratio of 8:1:

1.

5. The insole blanking material detection method based on improved YOLOv5 according to claim 1 is characterized in that: The LightCut-YOLOv5 network model in step S3 includes an input end, a backbone network, a neck network and an output end; The input end passes through the backbone network and the neck network respectively and then outputs through the output end; The backbone network includes Conv module, C3F module, SE module and SPPF module; The Conv module includes a convolutional layer, a batch normalization layer, and a ReLu activation function; The C3F module includes a first path and a second path, the first path is composed of a Conv module and more than two Faster_Block modules, and the second path is composed of a Pconv module and a PWconv; The SPPF module includes the Conv module and the Maxpool module; The CBAM module includes a channel attention module and a spatial attention module; The neck network contains multiple C3F modules for further feature extraction and fusion.

6. The insole blanking material detection method based on improved YOLOv5 according to claim 1 is characterized in that: The Focal-EIoU loss function adopts the FocalLoss idea to give higher weights to difficult-to-detect samples, and combines the EIoU method to make the aspect ratio of the predicted box more consistent with the actual object, so as to achieve more accurate bounding box regression; The formula is as follows: Among them, IOU = |A∩B| / |A∪B|, γ is the control calculation value, w c and h c are the width and height of the minimum bounding box, b and b gt represents the center coordinates of the predicted box and the real box, p(.) represents the Euclidean distance of the center point coordinates, and w gt and h gt is the width and height of the real box, w and h are the width and height of the predicted box.