Footwear multi-surface appearance defect detection system based on multi-mode AI

By combining a multimodal AI system with multispectral cameras and spectral vision technology, the problem of poor robustness of existing footwear appearance quality inspection equipment has been solved, and efficient recognition of small defects and autonomous learning optimization have been achieved.

CN120703107AActive Publication Date: 2025-09-26DONGGUAN CHUANGSHI AUTOMATION TECH CO LTD

Patent Information

Application Number
CN202511215394.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2025-09-26
Estimated Expiration
2045-08-28

AI Technical Summary

Technical Problem

Existing footwear appearance quality inspection equipment has poor robustness and is unable to identify a fixed number of defects, especially small defects that are difficult for ordinary cameras to identify.

Method used

A multimodal AI system is used, combining ordinary vision and spectral vision. Image data under different frequencies of light are collected through multispectral cameras and multispectral light sources. Combined with convolutional neural networks, region proposal networks and soft K-means clustering algorithms, efficient identification of appearance defects of footwear products can be achieved.

Benefits of technology

It improves the ability to identify appearance defects of footwear products, can autonomously learn and optimize in real time, identify small defects, and enhance the robustness and accuracy of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120703107A_ABST
    Figure CN120703107A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of visual inspection, in particular to a footwear multi-surface appearance defect detection system based on multi-modal AI. The invention aims to solve the problems that a common camera is difficult to identify and small, and the existing visual neural network technology can only identify the appearance defect of a fixed number of footwear products. According to the technical scheme, the system comprises a defect detection hardware system and the like, and the defect detection hardware system is arranged at the tail of a footwear product production line and used for conducting image collection on each footwear product so as to judge whether each pair of footwear products meet the production standard or not. According to the invention, the image acquisition mechanism is arranged, the image acquisition mechanism passes through a plurality of multispectral cameras and a plurality of multispectral light sources, the spectral cameras can acquire spectral images while acquiring RGB images, and different spectral images have obvious geometric features and color features. Therefore, the multispectral camera can capture tiny defects of the appearance of the footwear product, and the defect identification capability of the multispectral camera is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of visual inspection technology, and in particular to a multi-modal AI-based footwear multi-surface appearance defect detection system. Background Art

[0002] Appearance quality is an important component of the overall quality of shoes, which directly affects consumers' purchasing decisions and brand image. As residents' income levels increase, consumers' demand for high-quality footwear products is increasing, prompting manufacturers to pay more attention to the testing of appearance quality in order to enhance the market competitiveness of their own footwear products.

[0003] Existing footwear appearance quality inspection equipment is often installed at the end of the footwear production line. It is based on manual assistance and uses ordinary cameras to collect information about footwear products. It relies on visual neural networks to determine whether there are defects on the surface of footwear products. However, existing visual neural network technology can only identify a fixed number of footwear appearance defects and has poor robustness. In addition, some small defects are difficult to identify with ordinary cameras. In this regard, a footwear multi-faceted appearance defect detection system based on multimodal AI is provided. Summary of the Invention

[0004] In view of the fact that ordinary cameras have difficulty in identifying small defects and that existing visual neural network technology can only identify a fixed number of appearance defects of footwear products, the present invention provides a multi-modal AI-based footwear multi-faceted appearance defect detection system based on a mixture of ordinary vision and spectral vision.

[0005] To achieve the above objectives, the present invention is implemented through the following technical solutions: A multi-modal AI-based footwear multi-surface appearance defect detection system, including: The defect detection hardware system, installed at the end of the footwear production line, is used to pair footwear products and capture individual images of each shoe to fully detect defects in the shoe's appearance and determine whether each pair of shoes meets production standards; A computing server, which interacts with the defect detection hardware system via Ethernet for data flow and control flow, is used to receive image data of footwear products collected by the defect detection hardware system. The computing server includes a computing model for detecting image data, can quickly process the image data and determine whether the footwear products meet production requirements, and can also send the determination information back to the defect detection hardware system to provide information support for the defect detection hardware system; The storage server interacts with the computing server through Ethernet for data flow, and the storage server is used to store the computing server's computing results as reference data for subsequent autonomous learning of the computing server.

[0006] Preferably, the defect detection hardware system includes: The base is the main body of the defect detection hardware system and is used to fix the modules of various functions; A transfer module, mounted on the base, for moving footwear produced by the footwear production line to an appearance defect inspection location; The blocking module is installed on the base and is used to group the footwear products into pairs and further separate each shoe product to ensure that each shoe product can undergo comprehensive quality inspection; The clamping module is used to clamp and move a single shoe product so that the single shoe product can be inspected for appearance defects at a specified location; The recognition module is installed on the base and is used to collect multi-view image data of footwear products under different frequency lights and send the image data to the computing server; a classification module, mounted on the base, for limiting the positions of the footwear products so that a pair of footwear products can be regrouped; The control circuit is installed inside the base and is used to control the operation of the transmission module, blocking module, clamping module, identification module and classification module, and can receive the calculation results of the calculation server and adjust the classification module based on the calculation results.

[0007] Preferably, the transmission module includes: a first conveyor belt, for moving footwear produced by the footwear production line to a defect detection hardware system; a second conveyor belt, located above the first conveyor belt, for moving the footwear products to an appearance defect inspection location; The third conveyor belt is a double-channel conveyor belt, which is used to transport defective footwear products and non-defective footwear products out from different channels.

[0008] Preferably, the blocking module includes: a first blocking mechanism for counting the footwear products in pairs to ensure that each pair of footwear products is individually inspected; The second blocking mechanism is used to separate the footwear products individually, ensuring that each footwear product is tested separately to ensure the quality of the footwear testing and avoid the situation where two footwear products are tested at the same time and overlap each other; Preferably, the identification module includes: A lifting mechanism is used to adjust the height of the image acquisition of the footwear product, providing a testing platform for the footwear product. A transmission part is provided on the top of the lifting mechanism to transport the tested footwear product away from the lifting mechanism. The image acquisition mechanism consists of multiple multispectral cameras and multiple multispectral light sources. The spectral cameras correspond one-to-one to the spectral light sources. The multispectral light sources can emit light of different frequencies at the same time. The different surface materials of footwear products have different absorption and reflection effects on the multispectral light sources, so that the multispectral camera can capture spectral images of footwear products with obvious geometric and color characteristics.

[0009] Preferably, the multispectral camera also has the ability to capture ordinary RGB images. Through a multi-channel approach, the spectral image and the RGB image can be merged into one image, and the image is sent to the computing server in the form of image data for computing.

[0010] Preferably, the classification module includes: a baffle mechanism, located above the third conveyor belt, for limiting the position of the footwear products on the third conveyor belt, thereby allowing two footwear products separated during inspection to be regrouped; The push rod mechanism is located at the baffle mechanism and is used to move the shoe products to the output end of the third conveyor belt when the shoe products being tested are unqualified.

[0011] Preferably, the computing server includes: The training module is used to extract data stored in the storage server and train the computing model; The computing module is used to calculate image data and extract features of the image data through a computing model, thereby obtaining defect data corresponding to the image data through calculation. The defect data is sent to the storage server and control circuit via Ethernet.

[0012] Preferably, the computing model included in the computing server includes: Convolutional neural network (CNN), which extracts deep features from image data through convolution kernels and generates feature maps; The Region Proposal Network (RPN) is used to generate candidate boxes in the feature map and determine whether the footwear product in the feature map has appearance defects; RoI Pooling is used to normalize candidate boxes of different sizes and convert them into feature maps of fixed size; Soft K-Means clustering (Soft K-Means) is used to identify the defect type in the candidate box.

[0013] Preferably, after training and computing a certain number of images, the number of clusters of the soft K-means clustering can be adjusted based on the elbow principle to keep the number of clusters of the soft K-means clustering at the optimal number, thereby enabling the soft K-means clustering to identify new defect types according to the change in the number of clusters during the training process.

[0014] The present invention provides an intelligent robot control system and method. Compared with the prior art, it has the following advantages: The present invention provides an image acquisition mechanism that uses multiple multispectral cameras and multiple multispectral light sources. The spectral cameras can capture spectral images under different frequencies of light while capturing ordinary RGB images. Different spectral images have distinct geometric and color features, enabling the multispectral camera to capture tiny defects in the appearance of footwear products, thereby improving the defect recognition capability of the multispectral camera. The present invention sets soft K-means clustering in the calculation model. After training and calculation on a certain number of images, the soft K-means clustering can adjust the cluster number of the soft K-means clustering based on the elbow principle, so that the cluster number of the soft K-means clustering is maintained at an optimal number, and the calculation model can be autonomously learned and optimized in real time, thereby improving the system's ability to identify appearance defects of footwear products. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 It is a structural schematic diagram of the present invention; Figure 2 is a first structural diagram of the defect detection hardware system of the present invention; Figure 3 is a second structural diagram of the defect detection hardware system of the present invention; Figure 4 This is a flow chart of footwear product defect identification in the present invention.

[0016] In the figure: 1-defect detection hardware system, 100-base, 101-transmission module, 1011-first conveyor belt, 1012-second conveyor belt, 1013-third conveyor belt, 102-blocking module, 1021-first blocking mechanism, 1022-second blocking mechanism, 103-clamping module, 104-identification module, 1041-lifting mechanism, 1042-image acquisition mechanism, 105-classification module, 1051-baffle mechanism, 1052-push rod mechanism. DETAILED DESCRIPTION

[0017] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0018] Example 1

[0019] A multi-modal AI-based footwear multi-surface appearance defect detection system, including: Defect detection hardware system 1, installed at the end of the footwear production line, is used to pair footwear products and capture individual images of each shoe product to fully detect defects in the appearance of the footwear products and further determine whether each pair of footwear products meets production standards; The computing server interacts with the defect detection hardware system 1 via Ethernet for data flow and control flow, and is used to receive the image data of the footwear product collected by the defect detection hardware system 1. The computing server includes a computing model for detecting the image data, can quickly process the image data and determine whether the footwear product meets production requirements, and can also send the determination information back to the defect detection hardware system 1 to provide information support for the defect detection hardware system 1; The storage server interacts with the computing server through Ethernet for data flow, and the storage server is used to store the computing server's computing results as reference data for subsequent autonomous learning of the computing server.

[0020] The defect detection hardware system 1 includes: The base 100 is the main body of the defect detection hardware system 1 and is used to fix the modules of various functions; The transmission module 101 is installed on the base 100 and is used to move the footwear products produced by the footwear production line to the appearance defect detection location; The blocking module 102 is mounted on the base 100 and is used to group the footwear products into pairs and further separate each footwear product to ensure that each footwear product can undergo a comprehensive quality inspection; The clamping module 103 is used to clamp and move a single shoe product so that the single shoe product can be inspected for appearance defects at a designated location; The recognition module 104 is installed on the base 100 and is used to collect multi-view image data of the footwear product under different frequency lights and send the image data to the computing server; The classification module 105 is installed on the base 100 and is used to limit the position of the footwear products so that a pair of footwear products can be regrouped; The control circuit is installed inside the base 100 and is used to control the operation of the transmission module 101, the blocking module 102, the clamping module 103, the identification module 104 and the classification module 105, and can receive the calculation results of the calculation server and adjust the classification module 105 based on the calculation results.

[0021] The transmission module 101 includes: The first conveyor belt 1011 is used to move the footwear products produced by the footwear production line to the defect detection hardware system 1; The second conveyor belt 1012 is located above the first conveyor belt 1011 and is used to move the footwear products to the appearance defect inspection area; The third conveyor belt 1013 is a dual-channel conveyor belt, which is used to transport defective footwear products and non-defective footwear products out from different channels.

[0022] The blocking module 102 includes: The first blocking mechanism 1021 is used to count the footwear products by pairs to ensure that each pair of footwear products is individually inspected; The second blocking mechanism 1022 is used to separate the footwear products individually, ensuring that each footwear product is tested separately to ensure the quality of the footwear testing and avoid the situation where two footwear products are tested at the same time and are covered; The identification module 104 includes: The lifting mechanism 1041 is used to adjust the height of the image acquisition of the footwear product and provide a testing platform for the footwear product. The top of the lifting mechanism 1041 is provided with a transmission part, which can send the tested footwear product away from the lifting mechanism 1041; The image acquisition mechanism 1042 is composed of multiple multispectral cameras and multiple multispectral light sources. The spectral cameras correspond one to one with the spectral light sources. The multispectral light sources can emit light of different frequencies at the same time. The different surface materials of the footwear products have different absorption and reflection effects on the multispectral light sources, so that the multispectral camera can capture spectral images of footwear products with obvious geometric features and obvious color features.

[0023] The multispectral camera also has the ability to capture ordinary RGB images. Through a multi-channel approach, it can merge the spectral image and the RGB image into one image, and send the image as image data to the computing server for calculation.

[0024] The classification module 105 includes: a baffle mechanism 1051, located above the third conveyor belt 1013, for limiting the position of the footwear products on the third conveyor belt 1013, thereby regrouping two footwear products that were separated during inspection; The push rod mechanism 1052 is located at the baffle mechanism 1051 and is used to move the shoe product to the output end of the third conveyor belt 1013 when the shoe product being inspected is unqualified.

[0025] The computing servers include: The training module is used to extract data stored in the storage server and train the computing model; The computing module is used to calculate image data and extract features of the image data through a computing model, thereby obtaining defect data corresponding to the image data through calculation. The defect data is sent to the storage server and control circuit via Ethernet.

[0026] The computing models contained in the computing server include: Convolutional neural network (CNN) extracts deep features from image data through convolution kernels and generates feature maps; The region proposal network (RPN) is used to generate candidate boxes in the feature map and determine whether the footwear product in the feature map has appearance defects; Interest pool RoI Pooling is used to normalize candidate boxes of different sizes and convert them into feature maps of fixed size; Soft K-Means clustering is used to identify the defect type in the candidate box.

[0027] Specifically, in the initial state, the staff needs to first enter the shoe product dataset into the storage server so that the storage server has the dataset to train the computing model in the computing server. The staff first enters the shoe product dataset into the storage server so that the computing server can call the shoe product dataset for training. The shoe product dataset includes image data of shoe products without appearance defects and image data of shoe products with appearance defects, and the image data volume of the dataset should be at least: in, is the minimum image data size of the footwear product dataset, The types of shoe appearance defects entered include those with glue overflow, burrs, scratches, crooked seams, and no defects. is the minimum amount of data for any type of shoe appearance defect. Considering the mass production of footwear products on a footwear production line, in this embodiment The value is set to 50 to prevent the model from ignoring the appearance defects of shoes with too little data during training, which may cause the model to overfit during training.

[0028] Subsequently, the training module of the computing server obtains the footwear product dataset of the storage server through Ethernet and starts to train the computing model. During the training process, the image data used for training is first normalized so that the size of the image data is scaled to a fixed size. Then, the deep features of the image data are extracted by the convolution kernel of the convolutional neural network (CNN) to obtain a feature map. A series of candidate boxes are generated on the feature map through the region proposal network (RPN). The classification branch of the region proposal network (RPN) is used to determine whether the candidate box contains the features of the appearance defects of the footwear product, and the regression branch of the region proposal network (RPN) is used to correct the position of the candidate box. At this time, the computing model obtains a large number of candidate boxes. Too many candidate boxes will increase the computing burden of the computing model. By limiting the size of the candidate box, judging whether the candidate box exceeds the range of the feature map, and the probability of the candidate box judging the appearance defects of the footwear product, a small number of candidate boxes that meet the size limit, are within the range of the feature map, and have a high judgment rate of the appearance defects of the footwear product are retained. Then, the candidate box is normalized by the interest pool (RoI Pooling) so that the feature maps within the candidate box have the same size. Finally, the soft K-means clustering (Soft K-means clustering) is used to cluster the candidate boxes. K-Means) is used to perform probabilistic classification on the feature maps within the candidate boxes. For soft K-means clustering, the number of clusters is adjusted using the elbow principle after every 5,000 training images, enabling the algorithm to effectively identify new types of defects that have not been labeled.

[0029] In this embodiment, the convolutional neural network (CNN) adopts a method of combining every two convolution layers with a pooling layer, with a total of 9 layers of neural networks, so that the convolutional neural network can effectively extract the texture features of footwear products layer by layer, and convert low-dimensional image data into high-dimensional feature maps, so as to effectively extract image features in the image data. Among them, each convolution layer calculation must be linearly corrected by the activation function ReLU. In addition, except for the convolution kernel of the last convolution layer, which is a 1×1 convolution kernel, the convolution kernels of other convolution layers are all 3×3 convolution kernels, so that the convolutional neural network can extract image features from image data with low overhead and high efficiency through the 3×3 convolution kernel, and enhance the expression ability of image features by adjusting the dimension of the image features through the last 1×1 convolution kernel and linearly combining the image features of different channels.

[0030] The Region Proposal Network (RPN) is a neural network that generates candidate boxes at each position in the feature map. The number of candidate boxes generated at each position in the candidate box feature map is: in, The number of candidate boxes generated for each position of the feature map, is the number of basic sizes of the candidate box, is the number of types of candidate box aspect ratios, The number of types of candidate box scaling.

[0031] By setting the candidate box, the computational model can quickly determine the defect location of the footwear product, thereby facilitating further classification of the footwear product defects.

[0032] Soft K-Means clustering is an unsupervised clustering method. Unlike K-Means clustering, which classifies based on the Euclidean distance between the corresponding coordinate points of the eigenvalues ​​in the feature map and the corresponding coordinate points of the cluster center, it directly determines the type of defect that the surface defects of the footwear products in the feature map belong to. Soft K-Means clustering, on the other hand, uses the Euclidean distance between the corresponding coordinate points of the eigenvalues ​​in the feature map and the corresponding coordinate points of the cluster center to perform probability distribution. This fuzzifies the type of surface defects in the footwear products in the feature map, making it more suitable for the probabilistic modeling requirements of neural network classifiers. The specific formula is: in, For the The defect represented by the feature map data belongs to The probability of the defect represented by the cluster center, is the degree of blur, when The bigger, The higher the probability that the defect represented by a feature graph data belongs only to a specific cluster center point, the closer the result is to the result of K-means clustering. For the The coordinate points corresponding to the eigenvalues ​​in the feature map, the dimension of the coordinate points is positively correlated with the number of features in the feature map, For the The coordinate points corresponding to the cluster centers, For the The corresponding coordinate points of the eigenvalues ​​in the feature map are The Euclidean distance between the corresponding coordinate points of the cluster centers, is the total number of cluster centers, representing the total number of defect types of footwear products, Equivalent to confidence, when When the value is higher than 0.9, it is considered that The defect represented by the feature map data belongs to The defects represented by the cluster centers.

[0033] In the soft When the mean clustering method updates the cluster center, the calculation formula for the corresponding coordinate points of the cluster center is: in, For the new The coordinate values ​​of the cluster centers, For the The defect represented by the feature map data belongs to The probability of the defect represented by the cluster center, For the The coordinate points corresponding to the eigenvalues ​​in the feature map are weighted by probability. Higher feature map data contributes more to the cluster center update.

[0034] It is worth noting that soft After a certain number of images have been trained and operated, mean clustering can be used to classify soft The number of clusters of mean clustering is adjusted to make the soft The number of clusters in mean clustering is kept at the optimal number, so that the soft Mean clustering can identify new defect types according to the changes in the number of clusters during training.

[0035] Specifically, the elbow principle is an empirical method used to The core idea of ​​determining the optimal number of clusters in mean clustering is to calculate the probability-weighted Euclidean distance between the corresponding coordinate points of the eigenvalues ​​in the feature map and the corresponding coordinate points of the cluster centers. The trend of changes in the value is analyzed to find the "inflection point" (i.e., the elbow) where the error decrease rate slows down significantly, so as to select the appropriate number of clusters. The analysis method adopted in this example is the curve fitting method. The functional relationship between the error and the number of clusters is obtained through curve fitting, and the quadratic derivative of the fitted curve function is taken. By comparing the results of the quadratic derivative of the curve function under different cluster numbers, a more appropriate number of clusters can be obtained. Among them, the number of clusters shall not be less than the number of defect types in the footwear product data set entered by the staff under the initial conditions. By adjusting the number of clusters, the operation model can quickly identify new types of footwear product appearance defects that have not been entered according to the update of the cluster center, or further refine the classification of the original footwear product appearance defects, thereby greatly improving the robustness of the operation model and making the operation model show strong self-learning ability. In addition, since footwear product detection often has a large amount of image data, in order to improve computational efficiency, when updating the number of clusters through the elbow principle, only a maximum number of 3 times is randomly selected. The image data is trained, where The number of image data of the footwear product dataset entered by the staff under initial conditions.

[0036] Specifically, whether the new cluster is a new type of appearance defect of footwear products or a further refinement of the original appearance defect of footwear products, the first The defect represented by the feature map data belongs to The probability of defects represented by cluster centers When the probability that the defect represented by the feature map data belongs to the defect represented by a certain cluster center is significantly higher than the probability that it belongs to the defect represented by other cluster centers, the new cluster is considered to be a further refinement of the original footwear product appearance defect. Otherwise, when the probability that the defect represented by the feature map data belongs to the defect represented by a certain cluster center is close to the probability that it belongs to the defect represented by other cluster centers, and the probability is high, the new cluster is considered to be a new type of footwear product appearance defect, and a pseudo-label is made for the footwear product appearance defect to temporarily replace the manually assigned semantic label until the semantic label is manually assigned. In this embodiment, when the probability that the defect represented by the feature graph data belongs to the defect represented by a certain cluster center is the highest, and the ratio of this probability to the probability that the defect represented by the feature graph data belongs to the defect represented by other cluster centers is greater than 1.5, the new cluster is considered to be a further refined classification of the original footwear product appearance defect; when the probability that the defect represented by the feature graph data belongs to the defect represented by a certain cluster center is the highest, and the difference between this probability and the probability that the defect represented by the feature graph data belongs to the defect represented by one of the cluster centers does not exceed 7%, the new cluster is considered to be a new type of footwear product appearance defect.

[0037] It is worth noting that since the distance calculation, probability calculation and cluster center update of soft K-means clustering are all differentiable operations, soft K-means clustering can optimize the convolutional neural network and region proposal network through back propagation. Moreover, as an unsupervised classification algorithm, soft K-means clustering can calculate the location and defect type of the appearance defects of the footwear products, and directly start the back propagation optimization of the convolutional neural network and region proposal network based on the defect location and defect type, so that the operation model can be autonomously learned and optimized in real time, thereby improving the system's ability to identify appearance defects of footwear products.

[0038] After completing the training, the staff officially started the system to work, wherein the shoe products will be transported to the first conveyor belt 1011 of the transmission module 101 through the shoe product production line, and the first conveyor belt 1011 will be started, and the first conveyor belt 1011 will drive the shoe products to move, and then the first blocking mechanism 1021 will be started to block the shoe products on the first conveyor belt 1011 until the two shoe products on the first conveyor belt 1011 touch each other. At this time, it is determined that the two shoe products are grouped together, and then the first blocking mechanism 1021 is reset to allow the grouped shoe products to continue to move on the first conveyor belt 1011. When the grouped shoe products are located below the clamping module 103, the first conveyor belt 1011 will The conveyor belt 1010 stops temporarily to facilitate the clamping module 103 to clamp the footwear products. The clamping module 103 clamps the grouped footwear products and moves them to the second conveyor belt 1012. The second conveyor belt 1012 starts and drives the grouped footwear products to move. The second blocking mechanism 1022 starts to separate the grouped footwear products. One of the footwear products is not blocked by the second blocking mechanism 1022, so that the footwear product is transferred to the lifting mechanism 1041 by the second conveyor belt 1012. The lifting mechanism 1041 lifts the footwear product so that the footwear product is located in the center of multiple multispectral cameras, and then lights up multiple multispectral light sources and gradually changes the light frequency of the light source. When the light frequency of the light source changes, the multispectral camera captures multiple external images of the footwear products. The multispectral camera can integrate multiple appearance feature images of footwear products into a single multi-channel image based on the different absorption rates of different spectra of different materials of footwear products, thereby obtaining the color difference and more three-dimensional appearance defect features of footwear products. It is worth noting that multiple appearance feature images of footwear products include RGB visual images and multi-band spectral images. The multispectral camera integrates multiple appearance feature images of footwear products into a single multi-channel image to achieve multimodal fusion of different types of images, and transmits the image to the computing module of the computing server in the form of image data through Ethernet. The computing module determines the defect location and defect type of the footwear product through the input image data, and packages the defect type and defect location into defect data (if no appearance defect of the footwear product is detected, the defect data is empty), and transmits the image to the computing server through Ethernet. The data is transmitted to the control circuit and the storage server. After completing the inspection of a shoe product, the lifting mechanism will move the shoe product to the third conveyor belt 1013 through its own transmission part. At this time, it is in the starting state. The baffle mechanism 1051 limits the movement of the shoe product on the third conveyor belt 1013 through the baffle, repeats the above-mentioned inspection of the appearance defects of the shoe product, and moves the other shoe product to the third conveyor belt 1013, so that the two shoe products are re-grouped under the drive of the third conveyor belt 1013. The control circuit adjusts the push rod mechanism 1052 based on the received defect data. When the defect data is not empty, the push rod mechanism 1052 is controlled to push the shoe product so that the shoe product leaves the system at the defective product exit of the third conveyor belt 1013. Otherwise,The baffle mechanism 1051 is reset to allow the third conveyor belt 1013 to carry the grouped footwear products out of the system at the finished product outlet of the third conveyor belt 1013, thereby achieving real-time detection of the appearance quality of the footwear products and screening of the footwear product quality.

[0039] It is worth noting that after each appearance defect inspection of a footwear product is completed, the image data and defect data will be transmitted via Ethernet and saved in a storage server to facilitate subsequent further training of the computational model. When the footwear product is moving, the image acquisition mechanism 1042 is not in working condition. At this time, the training module will extract data to continue training the computational model, so that the computational model can self-learn in a timely manner and has the ability to identify new defect types, so that the computational model has a higher recognition accuracy for the appearance defects of footwear products.

[0040] In summary, the intelligent robot control systems and methods of Examples 1 and 2 can enable robots to have self-sensing, self-adjusting, and self-learning working modes, achieve highly flexible work, and can replace industrial workers and service industry practitioners, with broad application prospects.

Claims

1. A multi-faceted shoe appearance defect detection system based on multimodal AI, characterized by: Includes: A defect detection hardware system (1) is provided at the end of a footwear production line, and is used to pair footwear products and capture individual images of each footwear product, so as to fully detect defects in the appearance of the footwear products and further determine whether each pair of footwear products meets the production standards; An operation server is configured to interact with the defect detection hardware system (1) via Ethernet for data flow and control flow, and is configured to receive image data of the footwear product collected by the defect detection hardware system (1). The operation server includes an operation model for detecting the image data, is capable of quickly processing the image data and determining whether the footwear product meets production requirements, and is capable of sending the determination information back to the defect detection hardware system (1), thereby providing information support for the defect detection hardware system (1); The storage server interacts with the operation server via Ethernet to perform data flow, and the storage server is used to store the operation results of the operation server as reference data for subsequent autonomous learning of the operation server.

2. The multi-faceted appearance defect detection system for footwear based on multimodal AI according to claim 1, characterized in that: The defect detection hardware system (1) includes: The base (100) is the main body of the defect detection hardware system (1) and is used to fix the modules of various functions; A transmission module (101), mounted on the base (100), is used to move footwear products produced by the footwear production line to an appearance defect detection location; a blocking module (102), mounted on the base (100), for grouping the footwear products into pairs and further separating each footwear product to ensure that each footwear product can undergo a comprehensive quality inspection; A clamping module (103) is used to clamp and move a single shoe product, so that the single shoe product can be inspected for appearance defects at a designated location; an identification module (104), mounted on the base (100), for collecting multi-view image data of footwear products under light of different frequencies, and sending the image data to the computing server; a classification module (105), mounted on the base (100), for limiting the position of footwear products so that a pair of footwear products can be regrouped; The control circuit is installed inside the base (100) and is used to control the operation of the transmission module (101), the blocking module (102), the clamping module (103), the identification module (104) and the classification module (105), and is capable of receiving a calculation result from the calculation server and adjusting the classification module (105) based on the calculation result.

3. The multi-faceted appearance defect detection system for footwear based on multimodal AI according to claim 2, characterized in that: The transmission module (101) includes: A first conveyor belt (1011) is used to move the footwear products produced by the footwear production line to the defect detection hardware system (1); a second conveyor belt (1012), located above the first conveyor belt (1011), and used to move the footwear products to an appearance defect inspection location; The third conveyor belt (1013) is a dual-channel conveyor belt, used for conveying defective footwear products and non-defective footwear products from different channels.

4. The multi-faceted appearance defect detection system for footwear based on multimodal AI according to claim 2, characterized in that: The blocking module (102) includes: A first blocking mechanism (1021) is used to count the footwear products in pairs, ensuring that each pair of footwear products is individually inspected; The second blocking mechanism (1022) is used to separate the footwear products individually, ensuring that each footwear product is separated and tested individually, thereby ensuring the quality of the footwear testing and avoiding the situation where two footwear products are covered when tested at the same time.

5. The multi-faceted appearance defect detection system for footwear based on multimodal AI according to claim 2, characterized in that: The identification module (104) includes: A lifting mechanism (1041) is used to adjust the image acquisition height of the footwear product and provide a testing platform for the footwear product. A transmission part is provided on the top of the lifting mechanism (1041) and can send the tested footwear product away from the lifting mechanism (1041); The image acquisition mechanism (1042) is composed of a plurality of multispectral cameras and a plurality of multispectral light sources. The spectral cameras correspond to the spectral light sources one by one. The multispectral light sources can emit light of different frequencies at the same time. The multispectral light sources are able to absorb and reflect light of different surface materials of footwear products in different ways, so that the multispectral camera can acquire spectral images of footwear products with obvious geometric features and obvious color features.

6. The multi-faceted appearance defect detection system for footwear based on multimodal AI according to claim 5, characterized in that: The multispectral camera also has the ability to capture ordinary RGB images. Through a multi-channel approach, it can merge the spectral image and the RGB image into one image, and send the image as image data to the computing server for computing.

7. The multi-faceted appearance defect detection system for footwear based on multimodal AI according to claim 3, characterized in that: The classification module (105) includes: a baffle mechanism (1051), located above the third conveyor belt (1013), for limiting the position of the footwear products on the third conveyor belt (1013), thereby regrouping two footwear products that were separated during detection; The push rod mechanism (1052) is located at the baffle mechanism (1051) and is used to move the shoe product to the output end of the third conveyor belt (1013) when the shoe product being tested is unqualified.

8. The multi-faceted appearance defect detection system for footwear based on multimodal AI according to claim 2, characterized in that: The computing server includes: A training module, used to extract data stored in the storage server and train a computing model; The computing module is used to calculate the image data and extract the features of the image data through the computing model, thereby obtaining defect data corresponding to the image data through calculation, and the defect data is sent to the storage server and the control circuit via Ethernet.

9. The multi-faceted appearance defect detection system for footwear based on multimodal AI according to claim 8, characterized in that: The computing models contained in the computing server include: Convolutional neural network (CNN), which extracts deep features from image data through convolution kernels and generates feature maps; The Region Proposal Network (RPN) is used to generate candidate boxes in the feature map and determine whether the footwear products in the feature map have appearance defects; RoI Pooling is used to normalize candidate boxes of different sizes and convert them into feature maps of fixed size; Soft K-Means clustering (Soft K-Means) is used to identify the defect type in the candidate box.

10. The multi-faceted appearance defect detection system for footwear based on multimodal AI according to claim 9, characterized in that: After training and computing a certain number of images, the number of clusters of the soft K-means clustering can be adjusted based on the elbow principle to keep the number of clusters of the soft K-means clustering at the optimal number, so that the soft K-means clustering can identify new defect types according to the change of the number of clusters during the training process.

Citation Information

Patent Citations

  • Vamp logo multi-azimuth vision detection method and system

    CN106903075A

  • E-TPU shoe insole defect detection method and system based on background light

    CN112837289A

  • Automatic photographing system for shoes

    CN113038030A

  • Multispectral-thermal imaging composite defect detection device for leather production line and self-optimization method

    CN120253869A

  • Inspection Method for Shoes Sole

    KR101972805B1

Cited By

  • Safety shoe quality detection method based on artificial intelligence visual inspection

    CN120894639A