A method for detecting a shoe upper lasting

CN122347712BActive Publication Date: 2026-09-25QUANZHOU HUASI INTELLIGENT TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610803455.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-05
Publication Date
2026-09-25
Estimated Expiration
2046-06-05

AI Technical Summary

Technical Problem

目前,套楦工序基本由人工完成,为了保证套楦的准确性,通常需要设置专门的人工岗在套楦后进行检查,不仅增加了人工投入成本,还使得最终的套楦质量依赖于人工的经验,效率也无法进一步提升

Benefits of technology

[0019]1、本发明首先依据已准确套楦的标准鞋面图像构建模板库,然后在加工现场获取训练集并对YOLO模型进行训练,使YOLO模型能够识别鞋面的左右脚信息以及提取特征点,实际检测时,获取待检测鞋面的鞋面图像,采用训练后的YOLO模型获取待检测鞋面的左右脚信息和特征点,并分割待检测鞋面的最外圈轮廓,以及根据各特征点分割出对应的局部特征区域,根据最外圈轮廓计算待检测鞋面的鞋码,进而根据鞋码和左右脚信息获取所匹配的模板,最后获取同一特征点对应的模板特征区域与局部特征区域,分别计算模板特征区域和局部特征区域的凸包,计算两凸包的归一化匹配度,当在[0,)之间归一化匹配度超过所有特征点数量的60%时,判定套楦准确,否则判定套楦出现歪斜。如此实现在实际检测时无需人工参与,有效降低人工参与;利用预训练后的YOLO模型,能够有效提升检测效率与检测质量,且YOLO模型可利用小样本进行训练,尤其适合本发明中鞋子样本数量少的工业场合;模板库中具有分别对应于各鞋款下各鞋码的左右脚的标准鞋面图像,使本发明能够适用于不同尺码与款式的鞋,适应性更强。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122347712B_ABST
    Figure CN122347712B_ABST
Patent Text Reader

Abstract

The application provides a kind of vamp lasting detection method, it is related to shoemaking field, comprising: utilize accurate lasting standard vamp image to construct template library;In processing site acquisition training set and to YOLO model are trained, so it can identify the left and right foot information of vamp and extract feature point;Acquire the vamp image of the shoe surface to be detected, adopt YOLO model to obtain the left and right foot information and feature point of the shoe surface to be detected, and segment the outermost circle contour of the shoe surface to be detected, and according to each feature point, the corresponding local feature area is segmented;According to the size of the shoe surface to be detected, the size of the shoe surface to be detected is calculated according to the size of the shoe surface to be detected, and the size of the shoe surface to be detected is calculated according to the size of the shoe surface to be detected, and the size of the shoe surface to be detected is calculated according to the size of the shoe surface to be detected, and the size of the shoe surface to be detected is calculated according to the size of the shoe surface to be detected, and the size of the shoe is calculated according to the size of the shoe surface to be detected, and the size of the size of the shoe surface to be detected is calculated according to the size of the size of the shoe surface to be detected, and the size of the size of the shoe is calculated according to the size of the size of the shoe surface to be detected, and the matching template is obtained according to the size of the size of the shoe surface to be detected, and the size of
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of shoemaking, and in particular to a method for testing shoe upper lasts. Background Technology

[0002] In shoe manufacturing production lines, before processes such as sanding and gluing the shoe uppers, the uppers need to be fitted into shoe lasts—a process known as lasting. The accuracy of lasting directly impacts the quality of subsequent processes. Currently, lasting is primarily done manually. To ensure accuracy, dedicated personnel are typically assigned to inspect the lasts after they've been fitted. This not only increases labor costs but also makes the final lasting quality dependent on manual experience, hindering further efficiency improvements. Furthermore, the sizes and styles of shoes produced on a single production line frequently change, adding further complexity to lasting and inspection. Summary of the Invention

[0003] The main objective of this invention is to propose a method for testing shoe upper lasts, which can effectively improve testing efficiency and quality, reduce labor costs, and has strong adaptability.

[0004] This invention is achieved through the following technical solution:

[0005] A method for testing shoe upper lasts includes the following steps:

[0006] Step S1: Construct a template library using standard shoe upper images with accurate lasting. The template library includes standard shoe upper images for the left and right feet corresponding to each shoe size under each shoe style. The standard shoe upper images are set with multiple feature points and template feature regions containing each feature point.

[0007] Step S2: Obtain a training set at the processing site and train the YOLO model so that the YOLO model can recognize the left and right foot information of the shoe upper and extract feature points. The training set includes shoe upper images of all shoe sizes under all shoe styles. The shoe upper images are labeled with left and right foot tags and multiple feature points. The shoe upper image refers to the image of the shoe upper that has been put on the shoe last.

[0008] Step S3: Obtain the image of the shoe upper to be detected. Use the pre-trained YOLO model to obtain the left and right foot information and feature points of the shoe upper to be detected, and segment the outermost contour of the shoe upper to be detected, and segment the corresponding local feature regions according to each feature point.

[0009] Step S4: Calculate the shoe size of the shoe upper to be detected based on the outermost contour obtained in step S3, and obtain the matching template based on the shoe size and left and right foot information;

[0010] Step S5: Obtain the template feature region and local feature region corresponding to the same feature point, calculate the convex hull of the template feature region and the local feature region respectively, and calculate the normalized matching degree of the two convex hulls. When [0, If the normalized matching degree between the features exceeds 60% of the total number of feature points, the last is considered to be accurate; otherwise, the last is considered to be skewed.

[0011] Furthermore, in step S1, the feature points marked on the standard shoe upper image are from areas with high contrast and different shapes and colors from other areas of the shoe upper. These areas include shoelace holes, shoe upper logos, and / or seams of different fabrics on the shoe upper.

[0012] Furthermore, in step S1, obtaining the template feature region specifically includes: a rectangular or circular frame centered on the feature point of the standard shoe upper image, with a distance of 5-10 pixels from the feature point.

[0013] Furthermore, the YOLO model includes a context broadcasting module, which comprises a global average pooling layer, a linear transformation layer, a dimension expansion layer, and a feature fusion layer connected in sequence. The global average pooling layer performs global pooling on the input features to extract global context features. The linear transformation layer performs dimension mapping and feature transformation on the global context features. The dimension expansion layer restores the global context features after the linear transformation layer to the same dimension as the input features. The feature fusion layer performs element-wise superposition of the output of the dimension expansion layer with the original input features, and finally outputs an enhanced feature tensor.

[0014] Furthermore, in step S4, according to the formula Predict shoe size ,in, The parameters a and b are calculated using the least squares method, x is the shoe length of the shoe upper to be tested obtained from the outermost contour, and n is the number of samples used for the least squares calculation. Let k be the actual shoe length of the k-th sample. For the kth sample, the actual shoe size is represented by a standard image of the shoe upper with an accurate last.

[0015] Furthermore, in step S5, the i-th normalized matching degree According to the formula The calculation is performed, where m is the number of pixels in the standard shoe upper image that form the convex hull of the template feature region, and m' is the number of pixels in the shoe upper image to be detected that form the convex hull of the local feature region. Let be the coordinates of the i-th pixel within the convex hull of the template feature region. The convex hull of the local feature region is the first The position coordinates of each pixel for and The matching function returns the angle value.

[0016] Furthermore, in step S1, the template library also includes a standard centerline of a standard shoe upper image; in step S3, the actual centerline of the shoe upper image to be detected is calculated based on the outermost contour; and in step S5, the deviation between the actual centerline and the standard centerline is compared. If the deviation is within a set threshold, the last is determined to be accurate; otherwise, the last is determined to be skewed.

[0017] Furthermore, in step S2, the left foot is labeled with label 1 and the right foot is labeled with label 0.

[0018] As can be seen from the above description of the present invention, compared with the prior art, the present invention has the following beneficial effects:

[0019] 1. This invention first constructs a template library based on a standard shoe upper image that has been accurately lasted. Then, a training set is obtained at the processing site, and the YOLO model is trained to enable the YOLO model to recognize the left and right foot information of the shoe upper and extract feature points. During actual detection, the shoe upper image to be detected is obtained, and the trained YOLO model is used to obtain the left and right foot information and feature points of the shoe upper to be detected. The outermost contour of the shoe upper to be detected is segmented, and the corresponding local feature regions are segmented according to each feature point. The shoe size of the shoe upper to be detected is calculated based on the outermost contour. Then, the matching template is obtained based on the shoe size and left and right foot information. Finally, the template feature region and local feature region corresponding to the same feature point are obtained. The convex hulls of the template feature region and local feature region are calculated respectively, and the normalized matching degree of the two convex hulls is calculated. When [0, When the normalized matching degree between the features exceeds 60% of the total number of feature points, the last matching is considered accurate; otherwise, the last matching is considered skewed. This eliminates the need for manual intervention during actual detection, effectively reducing human involvement. Utilizing a pre-trained YOLO model significantly improves detection efficiency and quality. Furthermore, the YOLO model can be trained with small samples, making it particularly suitable for industrial applications with limited shoe samples. The template library contains standard upper images of the left and right feet corresponding to different shoe sizes for each shoe style, enabling the invention to be applied to shoes of different sizes and styles, thus enhancing its adaptability.

[0020] 2. This invention adds the judgment of the centerline offset, making the final judgment result more accurate. Attached Figure Description

[0021] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0022] Figure 1 This is a flowchart of the present invention.

[0023] Figure 2 This is a schematic diagram of the standard shoe upper of the present invention.

[0024] Figure 3This is a schematic diagram of the outermost contour of the shoe upper to be tested according to the present invention.

[0025] Figure 4 This is a schematic diagram of a local feature area of ​​the shoe surface to be tested according to the present invention. Detailed Implementation

[0026] The present invention will be further described below through specific embodiments.

[0027] like Figures 1 to 4 As shown, the shoe last testing method includes the following steps:

[0028] Step S1: Construct a template library using standard shoe upper images that are accurately lasted. The template library includes standard shoe upper images for the left and right feet corresponding to each shoe size under each shoe style. The standard shoe upper images are set with multiple feature points, template feature areas containing each feature point, and a standard centerline.

[0029] In this embodiment, the same image acquisition structure is used for acquiring the standard shoe upper image, constructing the training set of shoe upper images, and the shoe upper image to be detected. Specifically, it includes a single-plane camera and dual strip light sources. The shoe last is inverted and fixed at the last-fitting station. The single-plane camera captures the front end of the shoe last from bottom to top, while the dual strip light sources illuminate the upper area of ​​the shoe upper from bottom to top with adjustable angle and brightness. The acquired image is of the front half of the shoe. The specific implementation of the image acquisition structure is based on existing technology.

[0030] The feature points marked on the standard shoe upper image are from areas with high contrast and different shapes and colors from other areas of the shoe upper. These areas include shoelace eyelets, shoe upper logos, and seams where different fabrics are joined on the shoe upper.

[0031] Obtaining the template feature region specifically includes: a rectangular or circular frame centered on the feature points of the standard shoe upper image, with a distance of 5-10 pixels from the feature points. The template feature region can be obtained manually or using tools such as Roboflow.

[0032] Step S2: For each shoe size of each shoe style, obtain training sets for the left and right shoe uppers. Use each training set to pre-train the YOLO model so that the YOLO model can recognize the left and right foot information of the shoe upper and extract feature points. The training set is formed by shoe upper images labeled with left and right foot tags and multiple feature points. The shoe upper image refers to the image of the shoe upper on the shoe last.

[0033] In this embodiment, 277 images of the shoe upper were collected and labeled to form a training set. The left foot was labeled with label 1, and the right foot with label 0.

[0034] To address the "limited field of view" problem of local features and improve the model's ability to model complex scenes, the YOLO model adds a context broadcasting module. This module consists of a globally average pooling layer, a linear transformation layer, a dimension expansion layer, and a feature fusion layer connected in sequence. The globally average pooling layer performs global pooling on the input features to extract global context features. The linear transformation layer performs dimension mapping and feature transformation on the global context features. The dimension expansion layer restores the global context features after the linear transformation layer to the same dimension as the input features. The feature fusion layer element-wise superimposes the output of the dimension expansion layer with the original input features, ultimately outputting an enhanced feature tensor. The average pooling layer follows the formula... To achieve this, the linear layer is based on the formula. Implementation, the extension layer is based on the formula To achieve this, the output layer follows the formula. In this implementation, G represents the global context feature output by the average pooling layer, X represents the input feature, specifically the feature tensor extracted from the training set images through convolution, or the feature tensor extracted from the shoe upper image to be detected through convolution, dim is the dimension parameter of the feature tensor, used to define the number of channels or spatial dimension of the feature, and N is the total number of pixels in the feature tensor. G is the set of feature vectors for all pixels in the i-th row of the input features, W is the linear layer weight matrix, b' is the linear layer bias parameter, Broadcast is to expand G' to the target size size, B is the batch size (i.e., the number of samples in each batch), and C is the number of channels.

[0035] Step S3: Obtain the image of the shoe upper to be detected. Use a pre-trained YOLO model to obtain the left and right foot information and feature points of the shoe upper to be detected, and segment the outermost contour of the shoe upper to be detected. Also, segment the corresponding local feature regions based on each feature point, and calculate the actual central axis of the shoe upper to be detected based on the outermost contour. The segmentation of the outermost contour and the segmentation of local feature regions using the YOLO model are existing technologies. More specifically, the trained YOLO model includes YOLO-seg and YOLO-pose. YOLO-seg is trained using training data labeled with segmentation masks and left / right foot classifications to segment and obtain the outermost contour and left / right foot information. YOLO-pose, labeled with bounding boxes and keypoints (corresponding feature points), is used to segment the local feature regions.

[0036] Step S4: Calculate the shoe size of the shoe upper to be detected based on the outermost contour obtained in step S3, and obtain the matching template based on the shoe size and left and right foot information;

[0037] According to the formula Predict shoe size ,in, The parameters a and b are calculated using the least squares method, x is the shoe length of the shoe upper to be tested obtained from the outermost contour, and n is the number of samples used for the least squares calculation. Let k be the actual shoe length of the k-th sample. For the kth sample, the actual shoe size is represented by a standard image of the shoe upper with an accurate last.

[0038] Step S5: Obtain the template feature region and local feature region corresponding to the same feature point, calculate the convex hull of the template feature region and the local feature region respectively, and calculate the normalized matching degree of the two convex hulls. When [0, If the normalized matching degree between the two features exceeds 60% of the total number of feature points, and the deviation between the actual centerline and the standard centerline is within the set threshold, the last is considered to be accurate; otherwise, the last is considered to be skewed. The set threshold is 3mm~5mm.

[0039] According to the formula Calculate template feature region The convex hull, according to the formula Calculate the local adjustment area The convex hull is the convex hull, and this process is existing technology.

[0040] Each feature point corresponds to a normalized matching degree, and the i-th normalized matching degree... According to the formula The calculation is performed, where m is the number of pixels in the standard shoe upper image that form the convex hull of the template feature region, and m' is the number of pixels in the shoe upper image to be detected that form the convex hull of the local feature region. Let be the coordinates of the i-th pixel within the convex hull of the template feature region. The convex hull of the local feature region is the first The position coordinates of each pixel for and The matching function returns an angle value, where... This refers to existing technology, specifically as follows: , The pixels obtained using the neighbor difference approximation method Tangent direction angle at that point The pixels obtained using the neighbor difference approximation method The tangent direction angle at that point.

[0041] In this invention, the terms "first," "second," and "third," etc., are used only to distinguish similar objects and are not necessarily used to describe a specific order or sequence, nor should they be construed as indicating or implying relative importance. The use of terms such as "upper," "lower," "left," "right," "front," and "rear" to indicate orientation or positional relationships is based on the orientation or positional relationships shown in the accompanying drawings and is only for the convenience of describing the invention, not to indicate or imply that the device referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation on the scope of protection of this invention. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances.

[0042] Furthermore, in the description of this application, unless otherwise stated, "multiple" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.

[0043] The above are merely specific embodiments of the present invention, but the design concept of the present invention is not limited thereto. Any non-substantial modifications made to the present invention using this concept shall be considered as infringing upon the protection scope of the present invention.

Claims

1. A method for testing shoe upper lasts, characterized in that: Includes the following steps: Step S1: Construct a template library using standard shoe upper images with accurate lasting. The template library includes standard shoe upper images for the left and right feet corresponding to each shoe size under each shoe style. The standard shoe upper images are set with multiple feature points and template feature regions containing each feature point. Step S2: Obtain a training set at the processing site and train the YOLO model so that the YOLO model can recognize the left and right foot information of the shoe upper and extract feature points. The training set includes shoe upper images of all shoe sizes under all shoe styles. The shoe upper images are labeled with left and right foot tags and multiple feature points. The shoe upper image refers to the image of the shoe upper that has been put on the shoe last. Step S3: Obtain the image of the shoe upper to be detected. Use the trained YOLO model to obtain the left and right foot information and feature points of the shoe upper to be detected, and segment the outermost contour of the shoe upper to be detected, and segment the corresponding local feature regions according to each feature point. Step S4: Calculate the shoe size of the shoe upper to be detected based on the outermost contour obtained in step S3, and obtain the matching template based on the shoe size and left and right foot information; Step S5: Obtain the template feature region and local feature region corresponding to the same feature point, calculate the convex hull of the template feature region and the local feature region respectively, and calculate the normalized matching degree of the two convex hulls. When [0, When the normalized matching degree between the features exceeds 60% of the total number of feature points, the last is considered to be accurate; otherwise, the last is considered to be skewed. The YOLO model includes a context broadcasting module, which comprises a global average pooling layer, a linear transformation layer, a dimension expansion layer, and a feature fusion layer connected in sequence. The global average pooling layer performs global pooling on the input features to extract global context features. The linear transformation layer performs dimension mapping and feature transformation on the global context features. The dimension expansion layer restores the global context features after the linear transformation layer to the same dimension as the input features. The feature fusion layer performs element-wise superposition of the output of the dimension expansion layer with the original input features, and finally outputs an enhanced feature tensor. In step S5, the i-th normalized matching degree According to the formula The calculation is performed, where m is the number of pixels in the standard shoe upper image that form the convex hull of the template feature region, and m' is the number of pixels in the shoe upper image to be detected that form the convex hull of the local feature region. Let be the coordinates of the i-th pixel within the convex hull of the template feature region. The convex hull of the local feature region is the first The position coordinates of each pixel for and The matching function returns the angle value.

2. The method for detecting shoe upper lasts according to claim 1, characterized in that: In step S1, the feature points marked on the standard shoe upper image are from areas with high contrast and different shapes and colors from other areas of the shoe upper. These areas include shoelace holes, shoe upper logos, and / or seams of different fabrics on the shoe upper.

3. The method for detecting shoe upper lasts according to claim 2, characterized in that: In step S1, obtaining the template feature region specifically includes: a rectangular or circular frame centered on the feature point of the standard shoe upper image, with a distance of 5-10 pixels from the feature point.

4. A method for detecting shoe upper lasts according to claim 1, 2, or 3, characterized in that: In step S4, according to the formula Predict shoe size ,in, The parameters a and b are calculated using the least squares method, x is the shoe length of the shoe upper to be tested obtained from the outermost contour, and n is the number of samples used for the least squares calculation. Let k be the actual shoe length of the k-th sample. For the kth sample, the actual shoe size is represented by a standard image of the shoe upper with an accurate last.

5. A method for testing shoe upper lasts according to claim 1, 2, or 3, characterized in that: In step S1, the template library also includes a standard centerline of a standard shoe upper image. In step S3, the actual centerline of the shoe upper image to be detected is calculated based on the outermost contour. In step S5, the deviation between the actual centerline and the standard centerline is compared. If the deviation is within a set threshold, the last is determined to be accurate; otherwise, the last is determined to be skewed.

6. A method for detecting shoe upper lasts according to claim 1, 2, or 3, characterized in that: In step S2, the left foot is labeled with label 1 and the right foot is labeled with label 0.

Citation Information

Patent Citations

  • Machine vision-based sock defect detection method and system

    AU2025259943B1

  • Foot shoe tree model construction system based on image reconstruction and parameterization

    CN109032073A