A multi-modal AI-based footwear multi-faceted appearance defect detection system

By combining a multimodal AI system with a multispectral camera and spectral light source, along with convolutional neural networks and soft K-means clustering algorithms, the problems of difficulty in identifying small defects and poor robustness in existing technologies have been solved, achieving efficient and accurate detection of appearance defects in footwear products.

CN120703107BActive Publication Date: 2025-11-11DONGGUAN CHUANGSHI AUTOMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511215394.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2025-11-11
Estimated Expiration
2045-08-28

AI Technical Summary

Technical Problem

Existing footwear product appearance quality inspection equipment is unable to identify minor defects, and existing visual neural network technology has poor robustness and cannot effectively identify a variety of appearance defects.

Method used

Employing a multimodal AI system that combines ordinary vision and spectral vision, image data under different frequencies of light is acquired through a multispectral camera and multispectral light source. By combining convolutional neural networks, region proposal networks, and soft K-means clustering algorithms, the system can efficiently identify appearance defects in footwear products.

Benefits of technology

It improves the ability to identify minor defects, enhances the system's robustness and self-learning ability, and can optimize the defect identification model in real time, thereby improving detection accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120703107B_ABST
    Figure CN120703107B_ABST
Patent Text Reader

Abstract

This invention relates to the field of visual inspection technology, and more particularly to a multi-modal AI-based system for detecting defects in the appearance of footwear. The problem this invention addresses is that ordinary cameras struggle to identify minute defects, and existing visual neural network technologies can only identify a fixed number of appearance defects in footwear products. The technical solution includes a defect detection hardware system, located at the end of the footwear production line, used to acquire images of each footwear product to determine whether each pair meets production standards. This invention utilizes an image acquisition mechanism with multiple multispectral cameras and multiple multispectral light sources. The cameras acquire RGB images as well as spectral images; different spectral images possess distinct geometric and color features, enabling the multispectral cameras to capture minute defects in the appearance of footwear products, thus improving their defect recognition capabilities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of visual inspection technology, and in particular to a multi-modal AI-based system for detecting defects in the appearance of footwear. Background Technology

[0002] Appearance quality is an important component of the overall quality of shoes, directly affecting consumers' purchasing decisions and brand image. As residents' income levels rise, consumers' demand for high-quality footwear products is increasing, prompting manufacturers to pay more attention to appearance quality testing in order to enhance the market competitiveness of their footwear products.

[0003] Existing footwear appearance quality inspection equipment is often located at the end of the footwear production line and relies on manual assistance. It uses ordinary cameras to collect information about footwear products and visual neural networks to determine whether there are defects on the surface of the footwear products. However, existing visual neural network technology can only identify a fixed number of footwear appearance defects, has poor robustness, and is difficult to identify some small defects using ordinary cameras. To address this, we propose a multi-modal AI-based multi-faceted appearance defect detection system for footwear. Summary of the Invention

[0004] To address the limitations of ordinary cameras in identifying minute defects and the fact that existing visual neural network technologies can only identify a fixed number of appearance defects in footwear products, this invention provides a multi-modal AI-based footwear appearance defect detection system that combines ordinary vision and spectral vision.

[0005] To achieve the above objectives, the present invention provides the following technical solution:

[0006] A multi-modal AI-based system for detecting defects in the appearance of footwear includes:

[0007] The defect detection hardware system is set at the end of the footwear production line. It is used to pair footwear products in pairs and capture images of each footwear product individually in order to fully detect defects in the appearance of the footwear products and thus determine whether each pair of footwear products meets the production standards.

[0008] The computing server interacts with the defect detection hardware system via Ethernet to exchange data and control flows. It receives footwear product image data collected by the defect detection hardware system. The computing server contains a computing model for detecting image data, which can quickly process image data and determine whether footwear products meet production requirements. It can also send the judgment information back to the defect detection hardware system to provide information support for the defect detection hardware system.

[0009] The storage server interacts with the computing server via Ethernet to exchange data streams, and it stores the computing server's computation results as reference data for the computing server's subsequent autonomous learning.

[0010] As a preferred option, the defect detection hardware system includes:

[0011] The base is the main body of the defect detection hardware system, used to fix the various functional modules.

[0012] The transmission module, mounted on the base, is used to move footwear products produced on the footwear production line to the appearance defect detection point;

[0013] The blocking module, mounted on the base, is used to group footwear products by pairs and further separate each footwear product to ensure that each footwear product can undergo comprehensive quality inspection.

[0014] The clamping module is used to clamp and move a single shoe product, thereby enabling the single shoe product to be inspected for appearance defects at a designated location;

[0015] The recognition module, installed on the base, is used to collect multi-view image data of footwear products under different frequency light conditions and send the image data to the computing server;

[0016] The sorting module, installed on the base, is used to restrict the position of footwear products, allowing a pair of shoes to be regrouped.

[0017] The control circuit, installed inside the base, is used to control the operation of the transmission module, blocking module, clamping module, identification module and sorting module, and can receive the calculation results from the computing server and adjust the sorting module based on the calculation results.

[0018] Preferably, the transmission module includes:

[0019] The first conveyor belt is used to move the footwear products produced on the footwear production line to the defect detection hardware system;

[0020] The second conveyor belt, located above the first conveyor belt, is used to move footwear products to the appearance defect inspection area.

[0021] The third conveyor belt is a dual-channel conveyor belt, used to transport defective and non-defective footwear products from different channels.

[0022] Preferably, the blocking module includes:

[0023] The first barrier mechanism is used to count footwear products by pairs, ensuring that each pair of footwear products is inspected individually;

[0024] The second blocking mechanism is used to separate footwear products individually, ensuring that each footwear product is tested separately to guarantee the quality of footwear product testing and to prevent two footwear products from being tested simultaneously and causing overlap.

[0025] As a preferred embodiment, the identification module includes:

[0026] The lifting mechanism is used to adjust the image acquisition height of footwear products, providing a platform for the inspection of footwear products. The top of the lifting mechanism is equipped with a transmission part, which can send the inspected footwear products away from the lifting mechanism.

[0027] The image acquisition mechanism consists of multiple multispectral cameras and multiple multispectral light sources. Each multispectral camera corresponds to a multispectral light source. The multispectral light sources can emit light of different frequencies at the same time. By utilizing the different absorption and reflection effects of the multispectral light sources on the different surface materials of footwear products, the multispectral cameras can acquire spectral images of footwear products with obvious geometric and color features.

[0028] As an alternative, the multispectral camera also has the ability to capture ordinary RGB images. Through a multi-channel approach, it can merge the spectral image and the RGB image into a single image and send this image as image data to the computing server for processing.

[0029] As a preferred option, the classification module includes:

[0030] A baffle mechanism, located above the third conveyor belt, is used to restrict the position of footwear products on the third conveyor belt, thereby regrouping two footwear products that were separated during inspection.

[0031] The push rod mechanism, located at the baffle mechanism, is used to move the footwear products to the output end of the third conveyor belt when the inspected footwear products are unqualified.

[0032] Preferably, the computing server includes:

[0033] The training module is used to extract data stored in the storage server and train the computational model.

[0034] The computing module is used to calculate image data and extract features from the image data through a computing model, thereby obtaining defect data corresponding to the image data. The defect data is sent to the storage server and control circuit via Ethernet.

[0035] Preferably, the computing model included in the computing server includes:

[0036] Convolutional Neural Networks (CNNs) extract deep features from image data using convolutional kernels and generate feature maps.

[0037] Region Proposal Network (RPN) is used to generate candidate boxes in feature maps to determine whether footwear products have appearance defects in the feature maps;

[0038] RoI Pooling is used to normalize candidate boxes of different sizes, thereby converting them into feature maps of a fixed size.

[0039] Soft K-Means clustering is used to identify the type of defect in candidate boxes.

[0040] As a preferred approach, after training and processing a certain number of images, the number of clusters in soft K-means clustering can be adjusted based on the elbow principle to maintain the optimal number of clusters, thereby enabling soft K-means clustering to identify new defect types based on changes in the number of clusters during training.

[0041] This invention provides an intelligent robot control system and method. Compared with the prior art, it has the following advantages:

[0042] This invention sets up an image acquisition mechanism, which uses multiple multispectral cameras and multiple multispectral light sources. While acquiring ordinary RGB images, the multispectral cameras can also acquire spectral images under different frequencies of light. Different spectral images have obvious geometric and color features, thereby enabling the multispectral cameras to capture minute defects in the appearance of footwear products and improve the defect recognition capability of the multispectral cameras.

[0043] This invention sets up soft K-means clustering in the computational model. After training and computation on a certain number of images, the number of clusters in the soft K-means clustering can be adjusted based on the elbow principle to keep the number of clusters in the soft K-means clustering at the optimal level. This allows the computational model to learn and optimize in real time, thereby improving the system's ability to identify appearance defects in footwear products. Attached Figure Description

[0044] Figure 1 This is a schematic block diagram of the structure of the present invention;

[0045] Figure 2 This is a first structural diagram of the defect detection hardware system of the present invention;

[0046] Figure 3 This is a second structural diagram of the defect detection hardware system of the present invention;

[0047] Figure 4 This is a flowchart of the defect identification process for footwear products in this invention.

[0048] In the diagram: 1-Defect detection hardware system, 100-Base, 101-Transmission module, 1011-First conveyor belt, 1012-Second conveyor belt, 1013-Third conveyor belt, 102-Blocking module, 1021-First blocking mechanism, 1022-Second blocking mechanism, 103-Clamping module, 104-Identification module, 1041-Lifting mechanism, 1042-Image acquisition mechanism, 105-Classification module, 1051-Baffle mechanism, 1052-Push rod mechanism. Detailed Implementation

[0049] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0050] Example 1

[0051] A multi-modal AI-based system for detecting defects in the appearance of footwear includes:

[0052] The defect detection hardware system 1 is set at the end of the footwear production line. It is used to pair footwear products in pairs and capture images of each footwear product individually in order to fully detect defects in the appearance of the footwear products and thus determine whether each pair of footwear products meets the production standards.

[0053] The computing server interacts with the defect detection hardware system 1 via Ethernet for data and control flow. It is used to receive footwear product image data collected by the defect detection hardware system 1. The computing server contains a computing model for detecting image data, which can quickly process image data and determine whether footwear products meet production requirements. It can also send the judgment information back to the defect detection hardware system 1 to provide information support for the defect detection hardware system 1.

[0054] The storage server interacts with the computing server via Ethernet to exchange data streams, and it stores the computing server's computation results as reference data for the computing server's subsequent autonomous learning.

[0055] Defect detection hardware system 1 includes:

[0056] The base 100 is the main body of the defect detection hardware system 1, used to fix the modules of each function;

[0057] The transmission module 101 is installed on the base 100 and is used to move the footwear products produced on the footwear production line to the appearance defect detection point.

[0058] The blocking module 102, mounted on the base 100, is used to group footwear products by pairs and further separate each footwear product to ensure that each footwear product can undergo comprehensive quality inspection.

[0059] The clamping module 103 is used to clamp and move a single shoe product, so that the single shoe product can be inspected for appearance defects at a designated position.

[0060] The recognition module 104 is installed on the base 100 and is used to collect multi-view image data of footwear products under different frequency light conditions and send the image data to the computing server.

[0061] The classification module 105, installed on the base 100, is used to restrict the position of footwear products, so that a pair of footwear products can be regrouped.

[0062] The control circuit, installed inside the base 100, is used to control the operation of the transmission module 101, the blocking module 102, the clamping module 103, the identification module 104, and the sorting module 105. It can also receive the calculation results from the computing server and adjust the sorting module 105 based on the calculation results.

[0063] The transmission module 101 includes:

[0064] The first conveyor belt 1011 is used to move the footwear products produced by the footwear production line to the defect detection hardware system 1;

[0065] The second conveyor belt 1012, located above the first conveyor belt 1011, is used to move footwear products to the appearance defect inspection area.

[0066] The third conveyor belt 1013 is a dual-channel conveyor belt used to transport defective and non-defective footwear products from different channels.

[0067] The blocking module 102 includes:

[0068] The first blocking mechanism 1021 is used to count footwear products by pairs to ensure that each pair of footwear products is inspected individually;

[0069] The second blocking mechanism 1022 is used to separate footwear products individually to ensure that each footwear product is tested separately, thereby ensuring the quality of footwear product testing and preventing the situation of two footwear products being tested at the same time from overlapping.

[0070] The identification module 104 includes:

[0071] The lifting mechanism 1041 is used to adjust the image acquisition height of footwear products and provide a testing platform for footwear products. The top of the lifting mechanism 1041 is equipped with a transmission part, which can send the tested footwear products away from the lifting mechanism 1041.

[0072] The image acquisition mechanism 1042 consists of multiple multispectral cameras and multiple multispectral light sources. Each multispectral camera corresponds to a multispectral light source. The multispectral light sources can emit light of different frequencies at the same time. By utilizing the different absorption and reflection effects of the multispectral light sources on the different surface materials of footwear products, the multispectral cameras can acquire spectral images of footwear products with obvious geometric and color features.

[0073] Multispectral cameras also have the ability to capture ordinary RGB images. Through multi-channel processing, they can merge spectral images and RGB images into a single image and send this image as image data to a computing server for processing.

[0074] Classification module 105 includes:

[0075] The baffle mechanism 1051, located above the third conveyor belt 1013, is used to restrict the position of footwear products on the third conveyor belt 1013, thereby regrouping two footwear products that were separated during inspection.

[0076] The push rod mechanism 1052, located at the baffle mechanism 1051, is used to move the footwear product to the output end of the third conveyor belt 1013 when the inspected footwear product is unqualified.

[0077] The computing server includes:

[0078] The training module is used to extract data stored in the storage server and train the computational model.

[0079] The computing module is used to calculate image data and extract features from the image data through a computing model, thereby obtaining defect data corresponding to the image data. The defect data is sent to the storage server and control circuit via Ethernet.

[0080] The computing models contained within the computing server include:

[0081] Convolutional Neural Networks (CNNs) extract deep features from image data using convolutional kernels and generate feature maps.

[0082] The Region Proposal Network (RPN) is used to generate candidate boxes in the feature map to determine whether footwear products in the feature map have appearance defects.

[0083] RoI Pooling is used to normalize candidate boxes of different sizes, thereby converting them into feature maps of a fixed size.

[0084] Soft K-Means clustering is used to identify the type of defect in the candidate bounding box.

[0085] Specifically, initially, staff need to input a footwear product dataset into the storage server so that the storage server has the dataset to train the computational model in the computing server. The footwear product dataset includes images of footwear products without appearance defects and images of footwear products with appearance defects, and the minimum amount of image data in this dataset should be: in, To minimize the amount of image data in the footwear product dataset, The types of appearance defects for footwear to be entered include defects such as excess glue, rough edges, scratches, and crooked stitching, as well as no defects. This embodiment represents the minimum amount of data for any type of footwear appearance defect. Considering the large-scale production of footwear products on a footwear production line, this embodiment... The value is set to 50 to prevent the model from ignoring footwear appearance defects with insufficient data during training, which could lead to overfitting during training.

[0086] Subsequently, the training module of the computing server acquires the footwear product dataset from the storage server via Ethernet and begins training the computing model. During training, the image data used for training is first normalized to a fixed size. Then, depth features are extracted from the image data using the convolutional kernels of a Convolutional Neural Network (CNN), resulting in a feature map. A Region Proposal Network (RPN) generates a series of candidate boxes on the feature map. The classification branch of the RPN determines whether the candidate boxes contain features indicating footwear product defects, and the regression branch corrects the position of the candidate boxes. At this point, the computing model receives a large number of candidate boxes, which increases the computational burden. To address this, the size of the candidate boxes is limited, their extent within the feature map is assessed, and the probability of each candidate box identifying footwear product defects is considered. A small subset of candidate boxes that meet the size limit, are within the feature map's range, and have a high probability of identifying footwear product defects are retained. These candidate boxes are then normalized using RoI Pooling to ensure that the feature maps within each candidate box have the same size. Finally, soft K-means clustering is applied. K-Means performs probabilistic classification of feature maps within candidate boxes. For the number of clusters in soft K-means clustering, the number of clusters is adjusted using the elbow principle after every 5000 images of training, so that the algorithm can effectively identify new types of defects that have not been labeled.

[0087] In this embodiment, the Convolutional Neural Network (CNN) employs a combination of two convolutional layers and one pooling layer, resulting in a total of nine layers. This allows the CNN to effectively extract texture features from footwear products layer by layer, transforming low-dimensional image data into high-dimensional feature maps. This enables efficient extraction of image features from the image data. Each convolutional layer calculation is linearly corrected using the ReLU activation function. Furthermore, except for the last convolutional layer which uses a 1×1 kernel, all other convolutional layers use 3×3 kernels. This allows the CNN to extract image features from the image data efficiently and with low overhead using the 3×3 kernels, and to adjust the dimensionality of the image features using the final 1×1 kernel, linearly combining image features from different channels to enhance the expressive power of the image features.

[0088] Region Proposal Network (RPN) is a neural network that generates candidate boxes at each location in a feature map. The number of candidate boxes generated at each location in the candidate box feature map is: in, The number of candidate boxes generated for each location in the feature map. The number of candidate boxes based on their base dimensions. The number of categories with different aspect ratios for candidate boxes. The number of categories for scaling the candidate boxes.

[0089] By setting candidate boxes, the computational model can quickly determine the location of defects in footwear products, thus facilitating further classification of these defects.

[0090] Soft K-Means clustering is an unsupervised clustering method. Unlike K-Means clustering, which classifies products based on the Euclidean distance between the corresponding coordinates of feature values ​​and the corresponding coordinates of cluster centers, thus directly determining the type of surface defects in footwear products, Soft K-Means clustering uses the Euclidean distance between the corresponding coordinates of feature values ​​and the corresponding coordinates of cluster centers to create a probability distribution. This ambiguity makes the classification of surface defects in footwear products more suitable for the probabilistic modeling requirements of neural network classifiers. The specific formula is as follows: in, For the first The defect represented by the feature map data belongs to the first... The probability of a defect represented by a cluster center. For the degree of ambiguity, when The larger, The higher the probability that a defect represented by a certain feature map data belongs to a specific cluster center, the closer the result is to that of K-means clustering. For the first The coordinates of the feature values ​​in a feature map are positively correlated with the number of features in the feature map. For the first The coordinates of each cluster center For the first The coordinates of the corresponding feature values ​​in the feature map are related to the coordinates of the first feature map. The Euclidean distance between the corresponding coordinates of the cluster centers The total number of cluster centers represents the total number of defect types in footwear products. Equivalent to confidence level, when When the value is higher than 0.9, it is considered that the first... The defect represented by the feature map data belongs to the first... The defects represented by each cluster center.

[0091] In soft When updating cluster centers using mean clustering, the formula for calculating the corresponding coordinates of the cluster centers is: in, For the new first The coordinates of the cluster centers For the first The defect represented by the feature map data belongs to the first... The probability of a defect represented by a cluster center. For the first The coordinates of the feature values ​​in each feature map are weighted by probability to make... Higher feature map data contributes more to cluster center updates.

[0092] It is worth noting that soft After training and processing on a certain number of images, mean clustering can be used to cluster soft data based on the elbow principle. The number of clusters in mean clustering is adjusted to make the soft Mean clustering maintains an optimal number of clusters, thus enabling soft clustering to achieve optimal results. Mean clustering can identify new defect types based on changes in the number of clusters during training.

[0093] Specifically, the elbow principle is an empirical method used in... In mean-matrix clustering, determining the optimal number of clusters is based on the probability-weighted Euclidean distance between the coordinates of the corresponding eigenvalues ​​in the feature map and the corresponding coordinates of the cluster centers. This is achieved by analyzing the sum of squared errors of this Euclidean distance. The trend of value changes is analyzed to find the "inflection point" (i.e., the elbow) where the rate of error decrease slows down significantly, thus selecting an appropriate number of clusters. This example uses curve fitting to obtain the functional relationship between error and the number of clusters. By taking the second derivative of the fitted curve function and comparing the results under different numbers of clusters, a suitable number of clusters can be obtained. The number of clusters must not be less than the number of defect types in the footwear product dataset initially entered by the staff. Adjusting the number of clusters allows the computational model to quickly identify new types of unrecorded footwear product appearance defects based on the updated cluster centers, or to further refine the classification of existing footwear product appearance defects, thereby significantly improving the robustness of the computational model and demonstrating strong self-learning capabilities. Furthermore, since footwear product inspection often involves a large amount of image data, to improve computational efficiency, when updating the number of clusters using the elbow principle, only a maximum number of clusters three times the initial value is randomly selected. The image data is used for training, where... This refers to the number of image data points in the footwear product dataset entered by staff under initial conditions.

[0094] Specifically, whether the new cluster represents a new type of footwear product appearance defect or a further refinement of existing footwear product appearance defects can be determined through the [further analysis]. The defect represented by the feature map data belongs to the first... The probability of a defect represented by a cluster center The criteria are as follows: If the probability of a defect represented by the feature map data belonging to a certain cluster center is significantly higher than the probability of it belonging to other cluster centers, then the new cluster is considered a further refinement of the original footwear product appearance defects. Otherwise, if the probability of a defect represented by the feature map data belonging to a certain cluster center is similar to the probability of it belonging to other cluster centers, and this probability is relatively high, then the new cluster is considered a new type of footwear product appearance defect. A pseudo-label is then created for this footwear product appearance defect to temporarily replace the manually assigned semantic label, until a semantic label is manually assigned. In this embodiment, when the probability that the defect represented by the feature map data belongs to the defect represented by a certain cluster center is the highest, and the ratio of this probability to the probability that the defect represented by the feature map data belongs to the defect represented by other cluster centers is greater than 1.5, then the new cluster is considered to be a further refinement of the original footwear product appearance defects. When the probability that the defect represented by the feature map data belongs to the defect represented by a certain cluster center is the highest, and the difference between this probability and the probability that the defect represented by the feature map data belongs to the defect represented by one of the cluster centers does not exceed 7%, then the new cluster is considered to be a new type of footwear product appearance defect.

[0095] It is worth noting that since the distance calculation, probability calculation, and cluster center update of soft K-means clustering are all differentiable operations, soft K-means clustering can optimize convolutional neural networks and region proposal networks through backpropagation. Moreover, as an unsupervised classification algorithm, soft K-means clustering can directly start backpropagation optimization of convolutional neural networks and region proposal networks based on the defect location and defect type while calculating the location and defect type of the appearance defects of footwear products. This allows the computational model to learn and optimize autonomously in real time, thereby improving the footwear product appearance defect recognition capability of this system.

[0096] After completing the training, the staff officially started the system. Footwear products were transported from the footwear production line to the first conveyor belt 1011 of the transmission module 101. The first conveyor belt 1011 was started, moving the footwear products. Then, the first blocking mechanism 1021 was activated, blocking the footwear products on the first conveyor belt 1011 until two footwear products on the first conveyor belt 1011 came into contact. At this point, the two footwear products were considered to be grouped. The first blocking mechanism 1021 then reset, allowing the group of footwear products to continue moving on the first conveyor belt 1011. When the group of footwear products was below the clamping module 103, the first conveyor belt 1011... The process is temporarily halted to facilitate the clamping module 103's gripping of the footwear products. The clamping module 103 grips and transports the group of footwear products onto the second conveyor belt 1012. The second conveyor belt 1012 starts and moves the group of footwear products. The second blocking mechanism 1022 is activated, separating the group of footwear products. One footwear product is not blocked by the second blocking mechanism 1022 and is thus transported by the second conveyor belt 1012 to the lifting mechanism 1041. The lifting mechanism 1041 raises the footwear product, placing it at the center of multiple multispectral cameras. Subsequently, multiple multispectral light sources are illuminated, and the light frequencies of the light sources are gradually changed. As the light frequencies of the light sources change, the multispectral cameras capture multiple images of the footwear product. By observing feature images, and based on the different absorption rates of different materials on different spectra of footwear products, color differences and more three-dimensional appearance defect features of footwear products can be obtained. It is worth noting that multiple appearance feature images of footwear products include RGB visual images and multi-band spectral images. The multispectral camera integrates multiple appearance feature images of footwear products into a single multi-channel image, realizing multimodal fusion of different types of images. This image is then transmitted as image data to the computing module of the computing server via Ethernet. The computing module determines the location and type of defects in the footwear products based on the input image data, packages the defect type and defect location into defect data (if no appearance defect is detected, the defect data is empty), and transmits it via Ethernet. The data is transmitted to the control circuit and storage server. After the inspection of one shoe product is completed, the lifting mechanism moves the shoe product to the third conveyor belt 1013 via its own transmission mechanism. At this time, it is in the starting state. The baffle mechanism 1051 restricts the movement of the shoe product on the third conveyor belt 1013 by the baffle. The above-mentioned inspection of the appearance defects of the shoe products is repeated, and another shoe product is moved to the third conveyor belt 1013, so that the two shoe products are regrouped under the drive of the third conveyor belt 1013. The control circuit adjusts the push rod mechanism 1052 based on the received defect data. When the defect data is not empty, the control push rod mechanism 1052 pushes the shoe product, so that the shoe product leaves the system at the defective product exit of the third conveyor belt 1013; otherwise, it does not.The baffle mechanism 1051 resets, causing the third conveyor belt 1013 to carry the group of footwear products out of the system at the finished product outlet of the third conveyor belt 1013, thus realizing real-time detection of the appearance quality and screening of the footwear products.

[0097] It is worth noting that after each inspection of the appearance defects of footwear products is completed, the image data and defect data are transmitted via Ethernet and stored in the storage server to facilitate further training of the computational model. Furthermore, when the footwear products are in motion, the image acquisition mechanism 1042 is not in operation. At this time, the training module will extract data to continue training the computational model, enabling the computational model to learn itself in a timely manner and to have the ability to identify new defect types, thus giving the computational model higher accuracy in identifying appearance defects of footwear products.

[0098] In summary, the intelligent robot control system and method of Embodiments 1 and 2 enable robots to have self-sensing, self-adjusting, and self-learning working modes, achieving highly flexible work, and can replace industrial workers and service industry employees, with broad application prospects.

Claims

1. A multi-modal AI-based system for detecting defects in the appearance of footwear, characterized in that, Including: The defect detection hardware system (1) is set at the end of the footwear production line. It is used to pair footwear products in pairs and collect images of each footwear product individually in order to fully detect defects in the appearance of footwear products and then determine whether each pair of footwear products meets the production standards. The computing server interacts with the defect detection hardware system (1) via Ethernet to exchange data and control flows. It is used to receive footwear product image data collected by the defect detection hardware system (1). The computing server contains a computing model for detecting image data. It can quickly process image data and determine whether footwear products meet production requirements. It can also send the judgment information back to the defect detection hardware system (1) to provide information support for the defect detection hardware system (1). The storage server interacts with the computing server via Ethernet for data stream communication, and the storage server is used to store the computing server's computing results as reference data for the computing server's subsequent autonomous learning. The defect detection hardware system (1) includes: The base (100) is the main body of the defect detection hardware system (1) and is used to fix the modules of each function; A transmission module (101) is installed on the base (100) for moving footwear products produced on the footwear production line to the appearance defect detection point; A blocking module (102) is installed on the base (100) for grouping footwear products by pairs and further separating each footwear product to ensure that each footwear product can be subjected to comprehensive quality inspection. The clamping module (103) is used to clamp and move a single footwear product, thereby enabling the single footwear product to undergo appearance defect detection at a designated location; The identification module (104) is installed on the base (100) and is used to collect multi-view image data of footwear products under different frequency light and send the image data to the computing server; A classification module (105), mounted on the base (100), is used to restrict the position of footwear products so that a pair of footwear products can be regrouped; The control circuit is installed inside the base (100) and is used to control the operation of the transmission module (101), the blocking module (102), the clamping module (103), the identification module (104) and the classification module (105). It can also receive the calculation results from the computing server and adjust the classification module (105) based on the calculation results. The blocking module (102) includes: The first blocking mechanism (1021) is used to count footwear products by pairs to ensure that each pair of footwear products is inspected individually; The second blocking mechanism (1022) is used to separate footwear products individually to ensure that each footwear product is tested separately, thereby ensuring the quality of footwear product testing and avoiding the situation where two footwear products are tested at the same time and overlap occurs. The recognition module (104) includes an image acquisition mechanism (1042), which consists of multiple multispectral cameras and multiple multispectral light sources. The multispectral cameras and the multispectral light sources correspond one-to-one. The multispectral light sources can emit light of different frequencies at the same time. By utilizing the different absorption and reflection effects of the multispectral light sources on the different surface materials of footwear products, the multispectral cameras can acquire spectral images of footwear products with obvious geometric features and obvious color features.

2. The multi-modal AI-based footwear multi-faceted appearance defect detection system according to claim 1, characterized in that, The transmission module (101) includes: The first conveyor belt (1011) is used to move the footwear products produced by the footwear production line to the defect detection hardware system (1); The second conveyor belt (1012), located above the first conveyor belt (1011), is used to move footwear products to the appearance defect detection area; The third conveyor belt (1013) is a dual-channel conveyor belt used to transport defective and non-defective footwear products from different channels.

3. The multi-modal AI-based footwear multi-faceted appearance defect detection system according to claim 2, characterized in that, The identification module (104) also includes: The lifting mechanism (1041) is used to adjust the image acquisition height of footwear products and provide a testing platform for footwear products. The top of the lifting mechanism (1041) is provided with a transmission part, which can send the tested footwear products away from the lifting mechanism (1041).

4. The multi-modal AI-based footwear multi-faceted appearance defect detection system according to claim 3, characterized in that, The multispectral camera also has the ability to capture ordinary RGB images. It combines the spectral image and the RGB image into a single image using a multi-channel method, and sends the image as image data to the computing server for processing.

5. A multi-modal AI-based footwear multi-faceted appearance defect detection system according to claim 3, characterized in that, The classification module (105) includes: A baffle mechanism (1051), located above the third conveyor belt (1013), is used to restrict the position of footwear products on the third conveyor belt (1013), thereby regrouping two footwear products that were separated during inspection. The push rod mechanism (1052), located at the baffle mechanism (1051), is used to move the footwear product to the output end of the third conveyor belt (1013) when the footwear product being inspected is unqualified.

6. The multi-modal AI-based footwear multi-faceted appearance defect detection system according to claim 2, characterized in that, The computing server includes: A training module is used to extract data stored in the storage server and train a computational model. The computing module is used to calculate image data and extract features from the image data through a computing model, thereby obtaining defect data corresponding to the image data. The defect data is sent to the storage server and the control circuit via Ethernet.

7. A multi-modal AI-based footwear multi-faceted appearance defect detection system according to claim 6, characterized in that, The computing server includes the following computing models: Convolutional Neural Networks (CNNs) extract deep features from image data using convolutional kernels and generate feature maps. Region Proposal Network (RPN) is used to generate candidate boxes in feature maps to determine whether footwear products have appearance defects in the feature maps; RoI Pooling is used to normalize candidate boxes of different sizes, thereby converting them into feature maps of a fixed size. Soft K-Means clustering is used to identify the type of defect in candidate boxes.

8. A multi-modal AI-based footwear multi-faceted appearance defect detection system according to claim 7, characterized in that, After training and processing a certain number of images, the number of clusters in soft K-means clustering is adjusted based on the elbow principle to maintain the optimal number of clusters. This allows soft K-means clustering to identify new defect types based on changes in the number of clusters during training.

Citation Information

Patent Citations

  • Vamp logo multi-azimuth vision detection method and system

    CN106903075A

  • E-TPU shoe insole defect detection method and system based on background light

    CN112837289A