A Vision-Based Unsupervised Defect Recognition and Segmentation Method for Intelligent Manufacturing
Through the unsupervised identification and segmentation method, combined with the SAM model and context prompts, the problem of scarcity of labeled data in intelligent manufacturing is solved, and high-precision and real-time defect segmentation is achieved.
Patent Information
- Application Number
- CN202510401975.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-01
AI Technical Summary
In the field of intelligent manufacturing, it is difficult for the existing technology to achieve high-precision defect segmentation when labeled data is scarce, and traditional supervised learning methods rely on labeled data, have low detection efficiency and are not strong in real time.
The unsupervised recognition segmentation method based on vision is adopted, and defect segmentation is achieved by using SAM model and context prompts, combining image encoding, shadow detection and unsupervised clustering.
In the scenario of scarce labeled data, high-precision industrial defect segmentation can be realized, and real-time detection of product assembly lines can be carried out.
Smart Images

Figure CN119919430B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and specifically to a vision-based unsupervised defect recognition and segmentation method for intelligent manufacturing. Background Art
[0002] In the field of intelligent manufacturing, accurate and real-time defect segmentation is very important for ensuring product quality. Especially in manufacturing fields such as aerospace and automotive, quality control is crucial for safety. Precise identification and positioning of deviations through defect segmentation are an indispensable part of maintenance and error correction in the production process and a key step to ensure high quality.
[0003] Different from the medical field where computer vision technology is widely used, in industrial environments, the forms of defects are diverse, and most scenarios lack sufficient labeled samples. This complexity requires detection methods to be more adaptable.
[0004] With the rapid development of artificial intelligence technology, early solutions based on support vector product (SVM) and random forest (RandomForest) are increasingly being replaced by deep learning-based solutions. Recently, models based on computer vision and U-Net are currently a widely used solution for defect segmentation in the industrial field. However, in real-world scenarios, the labeled data in the industrial manufacturing field is often insufficient or even missing, which results in the inability to apply traditional supervised learning training methods for U-Net and its variants in many cases.
[0005] The machine learning-based solutions have the following disadvantages:
[0006] First, the generalization ability is not strong, and it is difficult for the model to be universal among different products;
[0007] Second, many methods require manual setting of anomaly thresholds;
[0008] Third, the accuracy is lower than that of deep learning-based solutions.
[0009] The supervised solutions based on U-Net and others have the following disadvantages:
[0010] First, they rely heavily on labeled data, and a large number of labeled samples are often required, otherwise it is easy to overfit;
[0011] Second, some models are very complex, with low detection efficiency and poor real-time performance;
[0012] Therefore, to address the above problems, a vision-based unsupervised defect recognition and segmentation method and system for intelligent manufacturing are needed to achieve high-precision detection under the premise of insufficient labeled data. Summary of the Invention
[0013] The object of the present invention is to provide a vision-based unsupervised defect recognition and segmentation method for intelligent manufacturing. Compared with most current supervised defect segmentation schemes, the present invention realizes high-precision defect segmentation in the industrial field based on the basic model SAM and context cues in the scenario of scarce labeled data. And with the help of the lightweight decoder of SAM, real-time detection of the product pipeline can be carried out.
[0014] The present invention is implemented as follows:
[0015] The present invention provides a vision-based unsupervised defect recognition and segmentation method for intelligent manufacturing, which is specifically executed according to the following steps:
[0016] S1: First, in the pipeline, collect the cut images after segmentation; S2: Then encode the collected images, specifically encoded in the form of embedding vectors;
[0017] In step S2, it is specifically executed according to the following steps:
[0018] S2.1: First, pass through a conv convolution to compress its original feature dimension to obtain patch embedding, as shown in Equation (1):
[0019] Equation (1);
[0020] Among them, I is the input image or patch, K is the weight of the convolution kernel that is gradually updated with gradient descent, I mn is the pixel value of the image at position m , n of the pixel value, K ijmn is the convolution kernel at position i , j multiplied by the image position m , n value, ( I ∗ K ) ij is the result at position i , j after the convolution operation;
[0021] S2.2: Flatten the features to obtain a one-dimensional vector, as shown in Equation (2):
[0022] Equation (2);
[0023] S2.3: Perform patch embedding to splice position information to obtain position embedding, as shown in Equation (3):
[0024] Equation (3);
[0025] Wherein, PE is a learnable embedding vector that self-learns the encoding of each position during the training process, PE The dimension of E is the same as
[0026] S2.4: The position embedding passes through 12 transformer blocks based on window partition and 4 global attention blocks to learn the image features; as follows:
[0027] Wherein, Transformer Blocker represents the output after being processed by K local attention blocks, and LayerNorm is a layer normalization operation;
[0028] S2.5: Finally, it undergoes two-layer conv convolution for dimensionality reduction to obtain the final image embbeding, as shown in Equation (4):
[0029] Equation (4);
[0030] Wherein, represents dimensionality reduction through two-layer conv convolution.
[0031] S3: Then perform shadow detection, and then use CBENet to calculate the color of the processed image and the original image; specifically, it is executed according to the following steps:
[0032] S3.1: Use BGShadowNet to generate a rough shadow removal map for the original image, and then use CBENet to calculate the color of the processed image and the original image, as shown in Equation (5):
[0033] Equation (5);
[0034] Wherein, A and B are two embedding vectors, and ||A|| and ||B|| are the L2 norms of the vectors respectively;
[0035] S3.2: Pass the data in Step S3.1 through Equations (6) - (7), calculate their similarity, and then delete the images with obvious deviations in the training image dataset;
[0036] Equation (6);
[0037] Equation (7);
[0038] Wherein, is the image to be deleted, , is the average feature of the training set images, is a preset threshold.
[0039] S4: Perform unsupervised clustering model training on the encoded training set images;
[0040] Specifically, it is executed according to the following steps:
[0041] S4.1: First, initialize the centroid. Select one from the encodings of the random training samples as the centroid, calculate the distance from each sample encoding vector to the centroid, denoted by D(x). Subsequently, calculate the probability of each sample being selected as the next centroid. After calculating the probabilities of all samples, divide the probabilities according to intervals, and randomly generate a number between 0 and 1 as the new centroid, as shown in Equation (8):
[0042] Equation (8);
[0043] S4.2: Assign clustering clusters and update the centroid. Calculate the Euclidean distance from each sample encoding vector to all centroids, and assign it to the nearest centroid, as shown in Equation (9):
[0044] Equation (9);
[0045] Wherein, x represents the current centroid, y represents the current sample, and d(x, y) represents the distance from x to y;
[0046] S4.3: Recalculate the centroid of each cluster again, as shown in Equation (10):
[0047] Equation (10);
[0048] Wherein, u j is the centroid of the j-th cluster, and C j is the set of all samples belonging to the j-th cluster;
[0049] S4.4: Repeat the above steps until the centroid encoding vector no longer changes.
[0050] S5: During the training process, after aggregation by the image processing module, the centroid image is obtained. Subsequently, using the binary threshold of OpenCV, the centroid is divided into the foreground containing defects and the background containing the remaining image, and use these foreground centroids containing defects as the prompt of the basic model SAM;
[0051] S6: Segment the defect; the mask of the test set in the training image will be calculated for similarity with the mask generated by SAM to evaluate the accuracy of the segmentation, as shown in Equation (5):
[0052] Equation (5);
[0053] Where A and B are two embedding vectors, and ||A|| and ||B|| are the L2 norms of the vectors respectively.
[0054] Furthermore, the present invention provides a vision-based unsupervised defect recognition and segmentation system for intelligent manufacturing, including an image processing module, which is used to encode images; and during model training, remove images containing shadow noise; during model training, perform unsupervised clustering on the images in the training set;
[0055] A context prompt information generation module, which is used to generate a prompt for tuning the base model. During training, after aggregation by the image processing module, a centroid image is obtained. Subsequently, using the binary threshold of OpenCV, the centroid is divided into a foreground containing defects and a background containing the remaining image. These foreground centroids containing defects are used as prompts for the base model SAM;
[0056] A defect segmentation module that generates effective segmentation masks for any given input and prompt. These masks are usually subsets, parts, or the entire image with different segmentation accuracies.
[0057] Furthermore, the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the steps of a vision-based unsupervised defect recognition and segmentation method for intelligent manufacturing as described in any one of the above.
[0058] Furthermore, the present invention provides a computer-readable storage medium, which includes an embedded processing system and a stored program. When the program runs under the control of the embedded system, it controls a vision-based unsupervised defect recognition and segmentation method as described in any one of the above.
[0059] Compared with the prior art, the beneficial effects of the present invention are:
[0060] Compared with most current supervised defect segmentation schemes, the present invention realizes high-precision defect segmentation in the industrial field based on the basic model SAM and context cues in scenarios with scarce labeled data. And with the help of the lightweight decoder of SAM, it can perform real-time detection on the product assembly line. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, so they should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other relevant drawings can also be obtained based on these drawings.
[0062] Figure 1 is the flowchart of the method of the present invention;
[0063] Figure 2 is the SAM framework diagram based on unsupervised cues of the present invention;
[0064] Figure 3 is the structural diagram of data processing of the intelligent manufacturing defect segmentation pipeline of the present invention;
[0065] Figure 4 is the system structure diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0066] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the present invention to be protected, but is only for the selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0067] Please refer to Figures 1 - 4 , the present invention provides a vision-based unsupervised defect recognition and segmentation method for intelligent manufacturing, which is specifically executed according to the following steps:
[0068] S1: First, in the assembly line, collect the incision images of the segmented incisions; specifically, collect images through a camera, and when necessary, perform supplementary shooting with lights to obtain high-definition images of the collected incisions;
[0069] S2: Then encode the collected image, specifically encoded in the form of an embedding vector; the specific steps are as follows:
[0070] S2.1: First, pass through a conv convolution to compress its original feature dimension to obtain patch embedding, as shown in Equation (1):
[0071] Equation (1);
[0072] Among them, I is the input image or patch, K is the weight of the convolution kernel that is gradually updated with gradient descent, I mn is the image at position m , n of the pixel value, K ijmn is the convolution kernel at position i , j multiplied by the image position m , n value, ( I ∗ K ) ij is the result after the convolution operation at position i , j ;
[0073] S2.2: Flatten the features to obtain a one-dimensional vector, as shown in Equation (2):
[0074] Equation (2);
[0075] S2.3: Concatenate the patch embedding with the position information to obtain the position embedding, as shown in Equation (3):
[0076] Equation (3);
[0077] Among them, PE is a learnable embedding vector that learns the encoding of each position during the training process, PE has the same dimension as E ;
[0078] S2.4: The position embedding passes through 12 transformer blocks based on window partition and 4 global attention blocks to learn the image features; as follows:
[0079] ;
[0080] Among them, Transformer Blocker represents the output after being processed by K local attention blocks, and LayerNorm is a layer normalization operation;
[0081] S2.5: Finally, dimensionality reduction is performed through two layers of conv convolutions to obtain the final image embbeding, as shown in Equation (4):
[0082] Equation (4);
[0083] Among them, represents dimensionality reduction through two conv convolutions.
[0084] S3: Then perform shadow detection, and then calculate the color of the processed image and the original image using CBENet; specifically, it is executed according to the following steps:
[0085] S3.1: Use BGShadowNet to generate a rough shadow removal map for the original image, and then calculate the color of the processed image and the original image using CBENet, as shown in Equation (5):
[0086] Equation (5);
[0087] Among them, A and B are two embedding vectors, and ||A|| and ||B|| are the L2 norms of the vectors respectively;
[0088] S3.2: Pass the data in step S3.1 through Equations (6) - (7), calculate the similarity between the two, and then delete the images with obvious deviations in the training image dataset;
[0089] Equation (6);
[0090] Equation (7);
[0091] Among them, is the image to be deleted, , is the average feature of the training set images, is the preset threshold.
[0092] S4: Train an unsupervised clustering model on the encoded training set images; specifically, it is executed according to the following steps:
[0093] S4.1: First, initialize the centroids. Select one from the encodings of the random training samples as the centroid, calculate the distance from each sample encoding vector to the centroid, denoted as D(x). Subsequently, calculate the probability of each sample being selected as the next centroid. After calculating the probabilities of all samples, divide the probabilities by intervals, and randomly generate a number between 0 and 1 as the new centroid, as shown in Equation (8):
[0094] Equation (8);
[0095] S4.2: Assign clustering clusters and update the centroids. Calculate the Euclidean distance from each sample encoding vector to all centroids, and assign it to the nearest centroid, as shown in Equation (9):
[0096] Equation (9);
[0097] Among them, x represents the current centroid, y represents the current sample, and d(x, y) represents the distance from x to y;
[0098] S4.3: Recalculate the centroid of each cluster again, as shown in Equation (10):
[0099] Equation (10);
[0100] Among them, u j is the centroid of the j-th cluster, and C j is the set of all samples belonging to the j-th cluster;
[0101] S4.4: Repeat the above steps until the centroid encoding vector no longer changes.
[0102] S5: During the training process, after aggregation by the image processing module, the centroid image is obtained. Subsequently, using the binary threshold of OpenCV, the centroid is divided into the foreground containing defects and the background containing the remaining image. Use these foreground centroids containing defects as the prompt for the basic model SAM;
[0103] ;
[0104] Using the centroid as the prompt can reduce the additional overhead caused by selecting the prompt area from each image and does not require any labeled data;
[0105] S6: Segment the defects; the mask of the test set in the training image will be calculated for similarity with the mask generated by SAM to evaluate the accuracy of the segmentation, as shown in Equation (5):
[0106] Equation (5);
[0107] Among them, A and B are two embedding vectors, and ||A|| and ||B|| are the L2 norms of the vectors respectively.
[0108] In this embodiment, the present invention provides a vision-based unsupervised defect recognition and segmentation system for intelligent manufacturing, including an image processing module, which is used to encode images; and during model training, remove images containing shadow noise; during model training, perform unsupervised clustering on the images in the training set.
[0109] A context prompt information generation module, which is used to generate a prompt for tuning the base model. During training, after aggregation by the image processing module, a centroid image is obtained. Subsequently, using the binary threshold of OpenCV, the centroid is divided into a foreground containing defects and a background containing the remaining image, and these foreground centroids containing defects are used as the prompt for the base model SAM.
[0110] A defect segmentation module that generates effective segmentation masks for any given input and prompt, and these masks are usually subsets, parts, or the entire image with different segmentation accuracies.
[0111] In this embodiment, the present invention provides a second embodiment. The present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the steps of a vision-based unsupervised defect recognition and segmentation method for intelligent manufacturing as described in any one of the above.
[0112] In this embodiment, the present invention provides a third embodiment. The present invention provides a computer-readable storage medium, which includes an embedded processing system and a stored program. When the program runs under the control of the embedded system, it controls a vision-based unsupervised defect recognition and segmentation method as described in any one of the above.
[0113] The above is only the preferred embodiment of the present invention and is not used to limit the present invention. For those skilled in the art, there are various changes and modifications to the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A vision-based unsupervised defect recognition and segmentation method for intelligent manufacturing, characterized in that: The specific implementation steps are as follows: S1: First, in the pipeline, collect the incision images of the segmented incisions; S2: Then encode the collected images, specifically encoded in the form of embedding vectors; The specific implementation steps are as follows: S2.1: First, pass through a conv convolution to compress its original feature dimension to obtain patch embedding; S2.2: Flatten the features to obtain a one-dimensional vector; S2.3: Concatenate the position information to obtain position embedding; S2.4: The position embedding passes through 12 transformer blocks based on window partition and 4 global attention blocks to learn the image features; S2.5: Finally, pass through two layers of conv convolution for dimensionality reduction to obtain the final image embbeding; S3: Then perform shadow detection, and then use CBENet to calculate the color of the processed image and the original image; The specific implementation steps are as follows: S3.1: Use BGShadowNet to generate a rough shadow removal map for the original image, and then use CBENet to calculate the color of the processed image and the original image; S3.2: Pass the data in step S3.1 through the following formula. After calculating the similarity between the two, delete the images with obvious deviations in the training image dataset; ; Among them, is the image to be deleted, , is the average feature of the training set images, is the preset threshold value; S4: Train an unsupervised clustering model for the encoded training set images; S5: During the training process, after aggregation by the image processing module, obtain the centroid image. Subsequently, use the binary threshold of OpenCV to divide the centroid into the foreground containing defects and the background containing the remaining image. Use these foreground centroids containing defects as the prompt for the basic model SAM; S6: Segment the defects.
2. The unsupervised recognition and segmentation method for defects based on vision for intelligent manufacturing according to claim 1, characterized in that: In step S4, the specific implementation steps are as follows: S4.1: First, initialize the centroid. Select one from the encodings of random training samples as the centroid. Calculate the distance from each sample encoding vector to the centroid, denoted by D(x). Subsequently, calculate the probability of each sample being selected as the next centroid. After calculating the probabilities of all samples, divide the probabilities according to intervals and randomly generate a number between 0 and 1 as the new centroid; S4.2: Assign the clustering clusters and update the centroid. Calculate the Euclidean distance from each sample encoding vector to all centroids and assign it to the nearest centroid; S4.3: Recalculate the centroid of each cluster; S4.4: Repeat the above steps until the centroid encoding vector no longer changes.
3. A vision-based unsupervised defect recognition and segmentation method for intelligent manufacturing according to claim 1, characterized in that: In step S6, the mask of the test set in the training image will be calculated for similarity with the mask generated by SAM to evaluate the segmentation accuracy.
Citation Information
Patent Citations
SAM-based unsupervised surgical instrument image segmentation method and system
CN117218340A
Image segmentation model training method, handheld object identification method, equipment and medium
CN117893756A