Sole segmentation method based on improved YOLOv8n-seg
By improving the YOLOv8n-seg model, using BiFPN, LFConv and SCIA attention mechanism modules, the problems of large calculation and difficult deployment of the three-dimensional glue-coated trajectory extraction model in the existing technology are solved, and an efficient and lightweight sole segmentation algorithm is realized, which is suitable for mobile robot systems.
Patent Information
- Application Number
- CN202510254386.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-06
AI Technical Summary
When the prior art realizes automatic extraction of three-dimensional glue coating trajectory of soles, there are defects such as large calculation amount and parameters, slow model operation speed, and difficulty in deploying on mobile terminals. Especially under personalized customization, the shoe sample replacement speed is fast, and the glue coating trajectory is required to be extracted in real time.
By improving the YOLOv8n-seg model, replacing the neck frame with BiFPN, replacing the Conv module with LFConv, and adding the SCIA attention mechanism module before the LFConv module, the improved SoleSeg model is obtained, reducing the GFLOPs, parameter quantity and size of the model, so that it meets the requirements of practicality and lightweight in terms of segmentation accuracy and speed.
The improved SoleSeg model has achieved practical and lightweight requirements in terms of segmentation accuracy and speed. GFLOPs have decreased by 19.5%, the parameter volume has decreased by 47.1%, and the model size has decreased by 41.5%, making it easier to deploy on mobile terminals such as robot systems.
Smart Images

Figure CN120107592A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image segmentation, and in particular to a shoe sole segmentation method based on improved YOLOv8n-seg. Background Art
[0002] While facing the challenges of the global economic situation, the development of digital online shopping has brought about the rise of personalized customization needs, prompting shoe companies to continue to innovate. Intelligent manufacturing, 3D printing technology, and new material applications are changing the production methods of the shoe industry. While providing personalized products that meet consumer needs, it is necessary to further improve production efficiency and reduce costs. At present, many shoe manufacturers still use manual gluing to bond the soles and uppers, which is inefficient, unstable in gluing quality, and the glue is harmful to the health of workers. This has become one of the bottlenecks in the production of multi-variety and small batch shoes under personalized customization needs. The use of automatic gluing technology using machine vision provides an economical solution for this.
[0003] The automatic gluing technology of machine vision and the extraction of gluing tracks directly affect the overall gluing quality and production efficiency. Due to the fast replacement of shoe samples under personalized customization, machine vision must have the ability to extract gluing tracks in real time, especially for the three-dimensional gluing tracks of the soles. How to achieve this at a low cost is the key technology for automatic gluing. Accurately obtaining the two-dimensional contour of the soles helps to obtain the three-dimensional gluing tracks.
[0004] Deep learning models rely on a large amount of computation and parameters to achieve high accuracy, which directly affects the model's running speed and memory usage. Mobile devices have limited memory and processing resources, and a model that is too large will extend the response time or even fail to run. Therefore, it is particularly important to build an accurate and lightweight deep learning model that is easier to deploy on mobile devices such as robotic systems. Summary of the invention
[0005] In order to address the shortcomings of the prior art, the purpose of the present invention is to provide a sole segmentation method based on improved YOLOv8n-seg. The improved model based on this method has a 19.5% decrease in GFLOPs, a 47.1% reduction in parameters, and a 41.5% reduction in model size compared to the unimproved model, so that the method meets the requirements of practicality and lightweight in terms of segmentation accuracy and speed, and is easier to deploy on mobile terminals including but not limited to robotic systems.
[0006] Specifically, the present invention is achieved through the following technical solutions:
[0007] A shoe sole segmentation method based on improved YOLOv8n-seg includes the following steps:
[0008] S1) obtaining a shoe sole data set and dividing it into a training set and a validation set;
[0009] S2) replacing the neck framework of the YOLOv8n-seg model with a BiFPN framework, replacing the Conv module with a LFConv module, and adding a SCIA attention mechanism module before the LFConv module to obtain an improved SoleSeg model;
[0010] S3) using the training set and the validation set to train and validate the SoleSeg model to obtain a trained sole segmentation model;
[0011] S4) Inputting the to-be-segmented sole image into a trained sole segmentation model to obtain a segmentation detection result.
[0012] Furthermore, the sole dataset in step S1) is obtained by mixing a public sole dataset on the Internet and a sole picture set taken in different environments.
[0013] Specifically, mixing the sole image sets taken in different environments can further obtain more actual data, overcome the problem of insufficient number of data sets, and improve the accuracy of the model.
[0014] Preferably, the online shoe sole public data set is obtained by downloading from the Internet. The different environments include but are not limited to different backgrounds and lights.
[0015] Furthermore, the sole image set is labeled before mixing; including but not limited to polygonal labeling of the sole images; specifically, the labeling includes but is not limited to manual labeling, such as using LabelMe to perform polygonal labeling on the sole part in the image, that is, labeling the sole part in the image.
[0016] Furthermore, after mixing, the sole dataset is enhanced by rotation, scaling and blurring.
[0017] Preferably, the specific manner of data enhancement is image brightening, image darkening, image rotation, image flipping, image scaling, image blurring, image deformation and adding noise.
[0018] Furthermore, the sole data set is divided into the training set and the validation set in a ratio of 8:2.
[0019] Furthermore, the weight calculation process in the BiFPN framework in step S2) is as follows:
[0020]
[0021] in, represents the intermediate transition features of the top-down path of the i-th layer; represents the input feature map of the i-th layer; w 1 and w2 is the weight parameter of the current layer input and the next layer input; ∈ is a hyperparameter to prevent the gradient from disappearing; It represents the final output feature of the i-th layer from bottom to top; the Convolution operation is to perform convolution on the feature map after weighted summation.
[0022] Furthermore, the LFConv module in step S2) includes two parallel DWConv modules for independently processing input features, a concat module for fully fusing information from different convolutional layers, a GN module for group normalization, and an activation function SiLu.
[0023] Furthermore, the SCIA attention mechanism module in step S2) is a three-branch structure, in which the first branch introduces the coordinate encoding mechanism of CA, performs global average pooling on the input feature map in the height direction and the width direction respectively, and obtains the feature vectors in the two directions. Then, the dependency relationship and spatial position information between channels are captured through the convolution layer, the information in the two directions is spliced, and the attention weight is obtained by the activation function; the second and third branches capture the feature associations between the channel dimension and the spatial dimension respectively; the size is restored through permute operation and Z-Pool processing, and then convolution and normalization; finally, the results of the three branches are fused to form an attention tensor across channels and space.
[0024] Furthermore, the first branch introduces the coordinate encoding mechanism of CA, and performs global average pooling on the input feature map in the height direction and width direction respectively. The output of the c-th channel at the height h and the output of the c-th channel at the width w can be expressed as follows:
[0025]
[0026] in, is the average value of the cth channel at height h with respect to width w; is the average value of the cth channel over the width w and the height h; x c (h, i) is the eigenvalue of the cth channel at height h and width i, x c (j, w) is the eigenvalue of the cth channel at height j and width w;
[0027] Concat is used to capture the dependency and spatial position information between channels, and a 1x1 convolution operation is performed to adjust the number of channels of the feature map. Batch normalization and activation function processing are performed to separate the features in the width and height directions along the spatial dimension. 1x1 convolution is applied to the two separate feature layers to restore the feature size. After the Sigmoid activation function is applied, the attention scores in the width and height dimensions are obtained. Finally, the original input feature map is multiplied by the attention scores in the width and height directions to obtain the output of the first branch.
[0028] Furthermore, the second branch captures information about the channel dimension. The input feature first undergoes a permute operation, rotates 90 degrees counterclockwise along the W axis, performs a Z-Pool process on the H dimension, restores the size through convolution and normalization, undergoes activation function processing, and finally undergoes a permuter operation, rotates 90 degrees clockwise along the W axis to restore the same scale as the input feature to obtain the output of the second branch.
[0029] Furthermore, the third branch captures information of the spatial dimension. The input feature first undergoes a permute operation, rotates 90 degrees counterclockwise along the H axis, performs a Z-Pool process on the W dimension, restores the size through convolution and normalization, undergoes activation function processing, and finally undergoes a permuter operation, rotates 90 degrees clockwise along the H axis to restore the same scale as the input feature to obtain the output of the third branch.
[0030] Furthermore, the Z-Pool process is to perform maximum pooling and average pooling on the values of all channels at each spatial position of the input feature, and connect the two new channels to form a new tensor. For example, after the tensor with a shape of (C×H×W) is processed by Z-pool, a new tensor with a shape of (2×H×W) will be obtained. The calculation process of Z-Pool is shown in the following formula:
[0031] Z-pool(χ)=[MaxPool 0d (χ), AvgPool 0d (χ)].
[0032] Furthermore, the 0th to 9th layers in the backbone network of the SoleSeg model are connected to LFConv, LFConv, C2f, LFConv, C2f, LFConv, C2f, LFConv, C2f, and SPPF respectively; the 10th to 30th layers in the neck network are Conv, Conv, Conv, Upsample, Concat, C2f, Upsample, Concat, C2f, SCIA, LFConv, Concat, C2f, SCIA, LFConv, Concat, C2f, SCIA, LFConv, Concat, and C2f respectively; the output end of the 10th layer SPPF is connected to The input of the 11th layer Conv is connected, the output of the 9th layer SPPF is connected to the input of the 10th layer Conv, the output of the 6th layer C2f is connected to the input of the 11th layer Conv, the output of the 4th layer C2f is connected to the input of the 12th layer Conv, the output of the 10th layer Conv is connected to the input of the 13th layer Upsample, the outputs of the 11th layer Conv and the 13th layer Upsample are connected to the input of the 14th layer Concat, the output of the 14th layer Concat is connected to the input of the 15th layer C2f, the output of the 15th layer C2f is connected to the input of the 16th layer Upsample, the 12th layer Conv and the 16th layer Upsa The output of mple is connected to the input of the 17th layer Concat, the output of the 17th layer Concat is connected to the input of the 18th layer C2f, the output of the 18th layer C2f is connected to the input of the 19th layer SCIA, the output of the 19th layer SCIA is connected to the input of the 20th layer LFConv, the outputs of the 12th layer Conv, the 18th layer C2f, and the 20th layer LFConv are connected to the input of the 21st layer Concat, the output of the 21st layer Concat is connected to the input of the 22nd layer C2f, the output of the 22nd layer C2f is connected to the input of the 23rd layer SCIA, and the output of the 23rd layer SCIA is connected to the input of the 24th layer LFConv The output ends of the 11th layer Conv, the 15th layer C2f, and the 24th layer LFConv are connected to the input end of the 25th layer Concat, the output end of the 25th layer Concat is connected to the input end of the 26th layer C2f, the output end of the 26th layer C2f is connected to the input end of the 27th layer SCIA, the output end of the 27th layer SCIA is connected to the input end of the 28th layer LFConv, the output ends of the 10th layer Conv and the 28th layer LFConv are connected to the input end of the 29th layer Concat, the output end of the 29th layer Concat is connected to the input end of the 30th layer C2f, and the 22nd layer C2f, the 26th layer C2f, and the 30th layer C2f are each connected to a segmentation head Segment.
[0033] Beneficial effects:
[0034] 1. Although YOLOv8n-seg has good effects in terms of speed and accuracy, the structure of combining FPN and PAN adopted by YOLOv8n-seg in the neck network lacks the ability to model deep interactions across space and channel dimensions. The present invention replaces the neck framework of the YOLOv8n-seg model with the BiFPN architecture to improve the signal integration ability of the model; replaces all Conv modules in the backbone network of the YOLOv8n-seg model with LFConv modules, and replaces all Conv modules in front of Concat in the neck network of the YOLOv8n-seg model with LFConv modules to improve the convergence speed and feature representation ability of the model; adds a spatial channel linkage attention mechanism SCIA module in front of the newly added LFConv module in the neck network to enhance the model feature extraction ability, and finally obtains the improved SoleSeg model, so that compared with YOLOv8n-seg, the improved SoleSeg model of the present invention is not only high in precision, but also has a small number of parameters, fast model detection speed, and a small model, and is easier to deploy on the mobile terminal.
[0035] 2. The LFConv of the present invention introduces efficient depth separable convolution and group normalization. The former first performs channel-by-channel convolution and then performs point-by-point convolution, which can effectively reduce the amount of calculation and the number of parameters while maintaining the network feature extraction capability. The latter divides the feature channels of the sample into multiple groups and normalizes the channels within the groups respectively. Two parallel DWConv modules are used to independently process the input features, and then the information from different convolutional layers is fully integrated through concat. The spliced feature map is processed by GN, which can enhance the stability of model training while reducing the dependence on batch size during training. Finally, the commonly used Sigmoid activation function is replaced by the SiLu activation function, and its smooth gradient flow characteristics contribute to the efficient learning and rapid convergence of the model.
[0036] 3. The SCIA of the present invention adopts the three-branch structure of TripletAttention. The three-branch structure enables the model to perform deep feature interaction between channels and spatial dimensions, and the coordinate encoding mechanism can highlight the features of key areas in the image, improve the model's perception of complex scenes and low-contrast areas, and thus improve the segmentation accuracy of the model.
[0037] 4. Compared with the unimproved YOLOv8n-seg, the final SoleSeg model of the present invention has a 19.5% decrease in GFLOPs, a 47.1% reduction in parameters, and a 41.5% reduction in model size, so that the model and method meet the requirements of practicality and lightweight in terms of segmentation accuracy and speed. When realizing sole gluing, the improved model is easier to deploy in a robot system. During operation, the sole image can be obtained through a camera, and the model can be used for segmentation, and the result can be transmitted to the robotic arm to guide it for subsequent operations. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 The SoleSeg network structure diagram of the present invention;
[0039] Figure 2 The BiFPN architecture principle diagram of the present invention;
[0040] Figure 3 This is a structural diagram of the simplified fusion convolution module LFConv of the present invention;
[0041] Figure 4 This is a structural diagram of the spatial channel linkage attention mechanism module SCIA of the present invention;
[0042] Figure 5 This is a diagram showing the effect of the SoleSeg network of the present invention on sole segmentation. DETAILED DESCRIPTION
[0043] The present invention is described in detail below through specific implementation methods, but the scope of the present invention is not limited to the listed embodiments. In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0044] Embodiment 1:
[0045] A shoe sole segmentation method based on improved YOLOv8n-seg, the method comprising the following steps:
[0046] Step 1: Download the public shoe sole dataset online, take pictures of the shoe soles in different environments, manually annotate the pictures, and mix them with the public dataset. Then, perform data enhancement on the dataset by rotating, scaling, and blurring.
[0047] Specifically, the sole public dataset comes from the public dataset Suelas, which includes various types of shoes such as sports shoes, leather shoes, canvas shoes, sandals, etc.
[0048] Specifically, in order to verify the versatility of the proposed model in practice, three types of cameras, SYD-FK-V1, HBVCAM-W202011HD V33, and realsense d435i, were used to establish a dataset of sole images with different light intensities, including 132 images in low-light environments (low light will cause detail loss and shadows to be difficult to identify), 256 images in medium-light environments (common lighting conditions), and 136 images in high-light environments (high light will cause overexposure of the image).
[0049] Furthermore, LabelMe is used to perform polygon annotation on the sole part in the image, that is, the sole part is labeled in the image.
[0050] Furthermore, in order to improve the robustness and generalization ability of model training, 8 methods of OpenCV are used for offline data enhancement, including brightening, darkening, rotation, flipping, scaling, blurring, deformation and adding noise. Brightening and darkening are used to simulate scenes with different lighting conditions; rotation, flipping and scaling are used to simulate different shooting positions and angles; blurring, deformation and adding noise are used to simulate artifacts that may be introduced during image acquisition and processing. Finally, a dataset containing 3196 enhanced images is constructed.
[0051] Step 2: Process the labeled data set and divide it into a training set and a validation set according to an 8:2 ratio to prepare for subsequent model training;
[0052] Step 3: Improve the YOLOv8n-seg model, replace the neck frame with the BiFPN architecture to strengthen information integration, replace the original Conv module of the model with the streamlined fusion convolution module LFConv to reduce the model calculation and parameter amount, and add the spatial channel linkage attention mechanism SCIA module after the C2f module to focus on the channel information and spatial information of the feature layer, and obtain the improved SoleSeg model;
[0053] Specifically, the neck framework of the YOLOv8n-seg model is replaced with the BiFPN architecture to improve the signal integration ability of the model; the Conv modules in the backbone network of the YOLOv8n-seg model are replaced with LFConv modules, and the Conv modules before Concat in the neck network of the YOLOv8n-seg model are replaced with LFConv modules to improve the convergence speed and feature representation ability of the model; the spatial channel linkage attention mechanism SCIA module is added in front of the newly added LFConv module in the neck network to enhance the feature extraction ability of the model, and finally the improved SoleSeg model is obtained; So In the backbone network of the leSeg model, the connections from layer 0 to layer 9 are LFConv, LFConv, C2f, LFConv, C2f, LFConv, C2f, LFConv, C2f, and SPPF respectively; in the neck network, the connections from layer 10 to layer 30 are Conv, Conv, Conv, Upsample, Concat, C2f, Upsample, Concat, C2f, SCIA, LFConv, Concat, C2f, SCIA, LFConv, Concat, C2f, SCIA, LFConv, Concat, C2f, SCIA, LFConv, Concat, and C2f respectively;The output of the 10th layer SPPF is connected to the input of the 11th layer Conv, the output of the 9th layer SPPF is connected to the input of the 10th layer Conv, the output of the 6th layer C2f is connected to the input of the 11th layer Conv, the output of the 4th layer C2f is connected to the input of the 12th layer Conv, the output of the 10th layer Conv is connected to the input of the 13th layer Upsample, the outputs of the 11th layer Conv and the 13th layer Upsample are connected to the input of the 14th layer Concat, and the 14th layer C The output of oncat is connected to the input of the 15th layer C2f, the output of the 15th layer C2f is connected to the input of the 16th layer Upsample, the outputs of the 12th layer Conv and the 16th layer Upsample are connected to the input of the 17th layer Concat, the output of the 17th layer Concat is connected to the input of the 18th layer C2f, the output of the 18th layer C2f is connected to the input of the 19th layer SCIA, the output of the 19th layer SCIA is connected to the input of the 20th layer LFConv, the 12th layer The outputs of the 18th layer C2f and the 20th layer LFConv are connected to the input of the 21st layer Concat, the output of the 21st layer Concat is connected to the input of the 22nd layer C2f, the output of the 22nd layer C2f is connected to the input of the 23rd layer SCIA, the output of the 23rd layer SCIA is connected to the input of the 24th layer LFConv, the outputs of the 11th layer Conv, the 15th layer C2f and the 24th layer LFConv are connected to the input of the 25th layer Concat, the 25th layer Conc The output of at is connected to the input of the 26th layer C2f, the output of the 26th layer C2f is connected to the input of the 27th layer SCIA, the output of the 27th layer SCIA is connected to the input of the 28th layer LFConv, the output of the 10th layer Conv and the 28th layer LFConv is connected to the input of the 29th layer Concat, the output of the 29th layer Concat is connected to the input of the 30th layer C2f, and the 22nd layer C2f, the 26th layer C2f, and the 30th layer C2f are each connected to a segmentation head Segment. ;
[0054] Specifically, although YOLOv8n-seg has good results in speed and accuracy, the FPN and PAN structure used in the neck network of YOLOv8n-seg lacks the ability to model deep interactions across spatial and channel dimensions. Therefore, its neck framework is replaced with the BiFPN architecture. The BiFPN architecture dynamically adjusts the fusion weights according to the importance of features at different levels through weighted feature fusion to enhance the information integration ability of the model. The weight calculation process in BiFPN is shown in the following formula:
[0055]
[0056] in, represents the intermediate transition features of the top-down path of the i-th layer; represents the input feature map of the i-th layer; w 1 and w 2 is the weight parameter of the current layer input and the next layer input; ∈ is a hyperparameter to prevent the gradient from disappearing; It represents the final output feature of the i-th layer from bottom to top; the Convolution operation is to perform convolution on the feature map after weighted summation.
[0057] The original Conv module of the YOLOv8n-seg model is replaced with a streamlined fused convolution module LFConv. LFConv is a new lightweight convolution module that uses two parallel depth-wise separable convolution modules DWConv to independently process input features, then fully fuses information from different convolutional layers through concat, and then processes it through group normalization GN to enhance the stability of model training and reduce dependence on batch size during training. Finally, the SiLu activation function is used to help the model learn efficiently and converge quickly.
[0058] In order to improve the segmentation performance of YOLOv8n-seg, a spatial channel linkage attention mechanism SCIA module is added to its network to focus on the channel information and spatial information of the feature layer. The SCIA module adopts a three-branch structure.
[0059] The first branch introduces the coordinate encoding mechanism of CA, and performs global average pooling on the input feature map in the height direction and width direction respectively. The output of the c-th channel at the height h and the output of the c-th channel at the width w can be expressed as follows:
[0060]
[0061] in, is the average value of the cth channel at height h with respect to width w; is the average value of the cth channel over the width w and the height h; x c (h, i) is the eigenvalue of the cth channel at height h and width i, x c (j, w) is the eigenvalue of the cth channel at height j and width w;
[0062] Concat is used to capture the dependency and spatial position information between channels, and a 1x1 convolution operation is performed to adjust the number of channels of the feature map. Batch normalization and activation function processing are performed to separate the features in the width and height directions along the spatial dimension. 1x1 convolution is applied to the two separate feature layers to restore the feature size. After the Sigmoid activation function is applied, the attention scores in the width and height dimensions are obtained. Finally, the original input feature map is multiplied by the attention scores in the width and height directions to obtain the output of the first branch.
[0063] The second branch captures information about the channel dimension. The input feature first undergoes a permute operation, rotates 90 degrees counterclockwise along the W axis, performs a Z-Pool operation on the H dimension, restores the size through convolution and normalization, and then undergoes activation function processing. Finally, it undergoes a permuter operation, rotates 90 degrees clockwise along the W axis to restore the same scale as the input feature, and obtains the output of the second branch.
[0064] The third branch captures information of the spatial dimension. The input feature first undergoes a permute operation, rotates 90 degrees counterclockwise along the H axis, performs a Z-Pool process on the W dimension, restores the size through convolution and normalization, undergoes activation function processing, and finally undergoes a permuter operation, rotates 90 degrees clockwise along the H axis to restore the same scale as the input feature to obtain the output of the third branch.
[0065] The Z-Pool process is to perform maximum pooling and average pooling on the values of all channels at each spatial position of the input feature, and connect the two new channels to form a new tensor. For example, after the tensor with a shape of (C×H×W) is processed by Z-pool, a new tensor with a shape of (2×H×W) will be obtained. The calculation process of Z-Pool is shown in the following formula:
[0066] Z-pool(χ)=[MaxPool 0d (χ), AvgPool 0d (χ)].
[0067] The results of the three branches are fused to form an attention tensor across channels and spaces.
[0068] Through the above method, we can obtain a sole segmentation algorithm with small model parameters and good accuracy and speed.
[0069] Step 4: Use the training set and validation set to train and validate the SoleSeg model to obtain a trained sole segmentation model;
[0070] Step 5: Get the sole image to be segmented, input the trained sole segmentation model to obtain the segmentation detection result (such as Figure 5 shown).
[0071] Comparison of model results:
[0072] Evaluation indicators:
[0073] Precision (P): The proportion of samples predicted by the model as positive to those actually positive. Precision reflects the accuracy of the model's prediction results. A high precision means fewer false positives. When the model predicts the sole area, most of them are correct.
[0074] Recall (R): The proportion of samples that are actually positive that the model correctly predicts as positive. A high recall rate means that the model can identify as many sole areas as possible.
[0075] Mean Average Precision (mAP): A comprehensive evaluation of the model's performance at different intersection-over-union thresholds, related to P and R. A high mAP indicates that the model is both accurate and comprehensive in the segmentation task.
[0076] The value after mAP is the threshold of the intersection-over-union ratio. 0.5 means that the overlap between the prediction and the truth is at least 50% to be considered correct. mAP0.5:0.95 refers to the average precision in the interval of 0.5 to 0.95, where the intersection-over-union ratio increases from 0.5 by 0.05 to 0.95.
[0077] GFLOPs: A metric used to measure computing performance, indicating the model's ability to complete floating-point operations when processing data. The smaller the GFlops value, the lighter the model is and the faster it runs.
[0078] Number of parameters (params): The total number of all trainable variables in the model, reflecting the complexity and storage space of the model.
[0079] Model size: The storage space of the model file. In resource-constrained environments, smaller models are easier to deploy.
[0080] 1. Comparison of YOLOv8-seg network models
[0081] The data set in step 2 of Example 1 was used to test and compare the n, s, m, l, and x models respectively. The results of the five models after testing are shown in Table 1 below:
[0082] Table 1 Comparison of YOLOv8-seg network models
[0083] Model P R mAP0.5 mAP0.5: 0.95 GFLOPs param size YOLOv8n-sea 0.961 0.886 0.925 0.831 12.8 3.4M 6.5 yolov8s-seg 0967 0.899 0.931 0.851 42.9 11.8M 22.7 yolov8m-seg 0.967 0.898 0.937 0.872 110.2 27.2M 52.2 yolov8l-seg 0.973 0.896 0.936 0.871 220.5 45.9M 87.9 yolov8x-seg 0.975 0.899 0.937 0.876 334.9 71.8M 137
[0084] As shown in Table 1, considering the segmentation accuracy, computational complexity, parameter amount and size comprehensively, the YOLOv8n-seg model was finally selected to ensure the segmentation accuracy while having the smallest computational complexity, parameter amount and size.
[0085] 2. Comparison of Attention Modules
[0086] In order to study the effectiveness of the proposed SCAI attention mechanism, four popular attention modules were added to the neck network in the same way, and different models were trained and tested. The results are shown in Table 2 below:
[0087] Table 2 Comparison results of attention modules
[0088] Model P R mAP0.5 mAP0.5: 0.95 GFLOPs params size yolov8n-seg 0.961 0.886 0.925 0.831 12.8 3.4M 6.5 +CBAM 0.966 0.898 0.942 0.845 12.8 3.4M 6.6 +SE 0.972 0.890 0.938 0.843 12.4 3.3M 6.5 +CA 0.967 0.893 0.940 0.842 12.4 3.3M 6.5 +Triplet Attention 0.972 0.881 0.935 0.844 12.4 3.3M 6.5 +SCIA 0.978 0.907 0.952 0.846 12.9 3.4M 6.5
[0089] In Table 2, the first row is yolov8n-seg, the second row is yolov8n-seg+CBAM, the third row is yolov8n-seg+SE, and so on for the following rows.
[0090] 3. Comparison of lightweight modules
[0091] In order to verify the advantages of the LFConv module in improving model efficiency and maintaining performance, a comparison test was conducted on various lightweight modules. The comparison results are shown in Table 3 below:
[0092] Table 3 Comparison results of lightweight modules
[0093] Model P R mAP0.5 mAP0.5∶0.95 GFLOPs params size yolov8n--seg--SCIA 0.978 0.907 0.942 0.846 12.9 3.4MM 6.5 +DWR 0.965 0.888 0.934 0.846 13.7 3.2M 12.4 +efficientViT 0.967 0.892 0.933 0.833 13.9 4.6M 9.4 +fasternet 0.964 0.884 0.936 0.827 15.2 4.7M 9.3 +slimeck 0.961 0.892 0.927 0.825 11.5 3.2M 6.4 +DWConv 0.943 0.879 0.927 0.821 11.5 2.8M 5.4 +LFConv 0.958 0.888 0.925 0.822 11.3 2.8M 5.7
[0094] In Table 3, the first row is yolov8n-seg+SCIA, the second row is yolov8n-seg+SCIA+DWR, the third row is yolov8n-seg+SCIA+efficientViT, and so on for the following rows.
[0095] 4. Ablation Experiment Comparison
[0096] In order to verify the effectiveness of the improved module, an ablation experiment was conducted and the results are shown in Table 4 below:
[0097] Table 4. Comparison results of ablation experiments
[0098] ① ② ③ ④ P R mP0.5 mAP0.5: 0.95 GFLOPs params size √ 0.961 0.886 0.925 0.831 12.8 3.4M 6.5 √ √ 0.978 0.907 0.952 0.846 12.9 3.4M 6.5 √ √ √ 0.958 0.888 0.925 0.822 11.3 2.8M 5.7 √ √ √ √ 0.962 0.882 0.927 0.83 10.3 1.8M 3.8
[0099] The “√” in Table 4 indicates the activation of the corresponding method or module, where ① is yolov8n-seg, ② is SCIA, ③ is LFConv, and ④ is BiFPN.
[0100] It can be seen from Tables 1-4 that the final SoleSeg model of the present invention has a 19.5% decrease in GFLOPs, a 47.1% reduction in parameters, and a 41.5% reduction in model size compared to the unimproved YOLOv8n-seg. At the same time, compared with other attention modules and lightweight modules in the prior art, it has better precision and lightweight balance, so that the model and method meet the requirements of practicality and lightweight in terms of segmentation accuracy and speed. When realizing sole gluing, the improved model is easier to deploy in a robot system. During operation, the sole image can be obtained through the camera, and the model can be used for segmentation, and the result can be transmitted to the robotic arm to guide it for subsequent operations.
[0101] The embodiments of the present invention are described in detail above. The description of the above embodiments is only used to help understand the method of the present invention and its core concept, and does not limit the scope of implementation of the present invention. Therefore, all equivalent changes or modifications made according to the structure, characteristics and principles described in the patent scope of the present invention should be included in the scope of the patent application of the present invention. In summary, the content of this specification should not be understood as a limitation of the present invention.
Claims
1. A shoe sole segmentation method based on improved YOLOv8n-seg, characterized in that: The steps include: S1) obtaining a shoe sole data set and dividing it into a training set and a validation set; S2) replacing the neck framework of the YOLOv8n-seg model with a BiFPN framework, replacing the Conv module with a LFConv module, and adding a SCIA attention mechanism module before the LFConv module to obtain an improved SoleSeg model; S3) using the training set and the validation set to train and validate the SoleSeg model to obtain a trained sole segmentation model; S4) Inputting the to-be-segmented sole image into a trained sole segmentation model to obtain a segmentation detection result.
2. The method according to claim 1, characterized in that The shoe sole dataset in step S1) is obtained by mixing a public shoe sole dataset on the Internet and a shoe sole image set taken in different environments.
3. The method according to claim 1, characterized in that The weight calculation process in the BiFPN framework in step S2) is shown in the following formula: in, represents the intermediate transition features of the top-down path of the i-th layer; represents the input feature map of the i-th layer; w1 and w2 are the weight parameters of the current layer input and the next layer input; ∈ is a hyperparameter to prevent the gradient from disappearing; It represents the final output feature of the i-th layer from bottom to top; the Convolution operation is to perform convolution on the feature map after weighted summation.
4. The method according to claim 1, characterized in that The LFConv module in step S2) includes two parallel DWConv modules that independently process input features, a concat module that fully integrates information from different convolutional layers, a GN module that performs group normalization, and an activation function SiLu.
5. The method according to claim 1, characterized in that: The SCIA attention mechanism module in step S2) is a three-branch structure. The first branch introduces the coordinate encoding mechanism of CA, performs global average pooling on the input feature map in the height direction and the width direction respectively, obtains the feature vectors in the two directions, captures the dependency relationship and spatial position information between channels through the convolution layer, splices the information in the two directions, and obtains the attention weight by the activation function; The second and third branches capture the feature correlation between the channel dimension and the spatial dimension, respectively; Through permute operation and Z-Pool processing, and then convolution and normalization to restore the size; finally, the results of the three branches are fused to form an attention tensor across channels and space.
6. The method according to claim 5, characterized in that The first branch introduces the coordinate encoding mechanism of CA, and performs global average pooling on the input feature map in the height direction and width direction respectively. The output of the c-th channel at the height h and the output of the c-th channel at the width w can be expressed as follows: in, is the average value of the cth channel at height h with respect to width w; is the average value of the cth channel over the width w and the height h; x c (h, i) is the eigenvalue of the cth channel at height h and width i, x c (j, w) is the eigenvalue of the cth channel at height j and width w; Concat is used to capture the dependency and spatial position information between channels, and a 1x1 convolution operation is performed to adjust the number of channels of the feature map. Batch normalization and activation function processing are performed to separate the features in the width and height directions along the spatial dimension. 1x1 convolution is applied to the two separate feature layers to restore the feature size. After the Sigmoid activation function is applied, the attention scores in the width and height dimensions are obtained. Finally, the original input feature map is multiplied by the attention scores in the width and height directions to obtain the output of the first branch.
7. The method according to claim 5, characterized in that The second branch captures information about the channel dimension. The input feature first undergoes a permute operation, rotates 90 degrees counterclockwise along the W axis, performs a Z-Pool operation on the H dimension, restores the size through convolution and normalization, and then undergoes activation function processing. Finally, it undergoes a permuter operation, rotates 90 degrees clockwise along the W axis to restore the same scale as the input feature, and obtains the output of the second branch.
8. The method according to claim 5, characterized in that The third branch captures information of the spatial dimension. The input feature first undergoes a permute operation, rotates 90 degrees counterclockwise along the H axis, performs a Z-Pool process on the W dimension, restores the size through convolution and normalization, undergoes activation function processing, and finally undergoes a permuter operation, rotates 90 degrees clockwise along the H axis to restore the same scale as the input feature to obtain the output of the third branch.
9. The method according to claim 5, characterized in that The Z-Pool process is to perform maximum pooling and average pooling on the values of all channels at each spatial position of the input feature, and connect the two new channels to form a new tensor. The calculation process of Z-Pool is shown in the following formula: Z-pool(x)=[MaxPool 0d (x), AvgPool 0d (x)].
10. The method according to claim 1, characterized in that In the backbone network of the SoleSeg model, layers 0 to 9 are connected in sequence to LFConv, LFConv, C2f, LFConv, C2f, LFConv, C2f, LFConv, C2f, and SPPF; layers 10 to 30 in the neck network are Conv, Conv, Conv, Upsample, Concat, C2f, Upsample, Concat, C2f, SCIA, LFConv, Concat, C2f, SCIA, LFConv, Concat, C2f, SCIA, LFConv, Concat, and C2f; the output end of the 10th SPPF is connected to the 11th The input of the 9th layer Conv is connected, the output of the 9th layer SPPF is connected to the input of the 10th layer Conv, the output of the 6th layer C2f is connected to the input of the 11th layer Conv, the output of the 4th layer C2f is connected to the input of the 12th layer Conv, the output of the 10th layer Conv is connected to the input of the 13th layer Upsample, the outputs of the 11th layer Conv and the 13th layer Upsample are connected to the input of the 14th layer Concat, the output of the 14th layer Concat is connected to the input of the 15th layer C2f, the output of the 15th layer C2f is connected to the input of the 16th layer Upsample, the 12th layer Conv and the 16th layer Upsample are connected. The output of le is connected to the input of the 17th layer Concat, the output of the 17th layer Concat is connected to the input of the 18th layer C2f, the output of the 18th layer C2f is connected to the input of the 19th layer SCIA, the output of the 19th layer SCIA is connected to the input of the 20th layer LFConv, the outputs of the 12th layer Conv, the 18th layer C2f, and the 20th layer LFConv are connected to the input of the 21st layer Concat, the output of the 21st layer Concat is connected to the input of the 22nd layer C2f, the output of the 22nd layer C2f is connected to the input of the 23rd layer SCIA, the output of the 23rd layer SCIA is connected to the input of the 24th layer LFConv, The outputs of the 11th layer Conv, the 15th layer C2f, and the 24th layer LFConv are connected to the input of the 25th layer Concat, the output of the 25th layer Concat is connected to the input of the 26th layer C2f, the output of the 26th layer C2f is connected to the input of the 27th layer SCIA, the output of the 27th layer SCIA is connected to the input of the 28th layer LFConv, the outputs of the 10th layer Conv and the 28th layer LFConv are connected to the input of the 29th layer Concat, the output of the 29th layer Concat is connected to the input of the 30th layer C2f, and the 22nd layer C2f, the 26th layer C2f, and the 30th layer C2f are each connected to a segmentation head Segment.