Workpiece crack visual detection method based on feature learning
Patent Information
- Application Number
- CN202410576082.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-10
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2044-05-10
AI Technical Summary
[0004]这些方法在公开数据集取得了良好的效果,但是面对工件裂纹图像并不能取得较好的效果
本发明提出的一种基于特征学习的工件裂纹视觉检测方法,由于实际工件图像的数据裂纹的细小和与正常区域的相似度极高,在现有技术上效果不好,促使我们改变了原有的基于图像分割方式的缺陷检测,将工件图像的裂纹检测问题看作一个基于图像的分类问题,最终通过分割块回归到整体图像的检测结果,采用这种新方式,可以有效解决复杂状况(背景噪声干扰,弱裂纹)的裂纹检测问题。
Smart Images

Figure CN118351381B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer vision and graphics processing technology, and relates to a visual detection method for workpiece cracks based on feature learning. It mainly addresses the problem of crack detection in DR images of industrial products, specifically involving a crack detection method based on a sliding window and image features. Background Technology
[0002] Crack detection technology is an important technique widely used in various engineering fields. It aims to detect and assess cracks in structures early to ensure the safety and reliability of facilities, buildings, or mechanical components. Cracks can arise from material fatigue, stress concentration, environmental factors, or manufacturing defects, and can exist in various complex materials and components such as steel structures, concrete members, pipes, and aerospace parts.
[0003] Manually inspecting cracks directly has significant limitations, requiring substantial manpower and time. Therefore, image-based crack detection technology is of paramount importance. Currently, image-based crack detection technology, with the rapid development of convolutional neural networks, mainly consists of two methods: The first is a supervised method, which trains the network by inputting images with crack defects labeled (including categories, bounding boxes, or pixel-wise representations). A "crack" here refers to a labeled region or image. The method proposed by Tao et al. ("X. Tao, D. Zhang, W. Ma, X. Liu, and D. Xu. Automatic metallic surface defect detection and recognition with convolutional neural networks. Applied Sciences, 8, 1575, 2018") uses an autoencoder for supervised feature learning. The second method is an unsupervised crack detection method, which typically only requires normal, crack-free samples for network training; this is also known as one-class learning. This method focuses more on crack-free (i.e., normal sample) features. When a previously unseen feature (abnormal feature) is found during crack detection, a crack is considered detected. In this case, the crack signifies an anomaly, hence the method is also called anomaly detection. The method proposed by Li et al. ("Z. Li, N. Li, K. Jiang, Ma. Zhang, W. Xing,...") H. Xiao, and Y. Gong, "Superpixel Masking and Inpainting for Self-Supervised Anomaly Detection. In Proceedings of British Machine Vision Conference, 2020", describes a method for superpixel segmentation and reconstruction of an input image. Defects are determined by comparing the differences between the reconstructed blocks and the input blocks.
[0004] These methods have achieved good results on public datasets, but they do not perform well on workpiece crack images. Summary of the Invention
[0005] Technical problems to be solved To overcome the shortcomings of existing technologies, this invention proposes a visual crack detection method for workpieces based on feature learning. This method is a crack detection method based on a sliding window and image feature learning. The method divides the image into blocks using a sliding window and then uses a deep learning method based on image features to determine whether cracks exist in the image blocks. This approach improves the accuracy of crack detection on workpiece image sets.
[0006] Technical solution A visual detection method for workpiece cracks based on feature learning, characterized by the following steps: Step 1: Enhance the contrast of images of workpieces suspected of containing cracks and reduce noise interference in the images; Step 2: Use a sliding window to view the image Data is segmented to obtain multiple image patches, forming an image set. To obtain the denoised training samples ; The image set The composition is: images In region Q, using The window is obtained by sliding and dividing the region Q with a step size s. Image patches constitute an image set ,in: This represents the number of image patches along the y-axis of the image. This represents the number of image patches along the x-axis of the image. Image width, Image length; Step 3: Transfer the training samples Divide the dataset into training and testing datasets. Input the training dataset into the classification and detection network for training. Once the classification and detection network is trained, it will classify and detect the data from step 1 after processing in step 2, and determine whether the image patch is abnormal. The classification and detection network includes a visual attention mechanism and non-local feature enhancement network backbone, a sampling feature pyramid model, and a classifier. The visual attention mechanism and non-local feature enhancement network backbone includes three cascaded attention modules. The sampling feature pyramid model includes two upsampling modules and a fusion module. The output of the third attention module is connected to the second upsampling module, the output of the second attention module is connected to the first upsampling module, the output of the second upsampling module is connected to the first upsampling module, and the output of the first upsampling module is connected to the fusion module. The output of the fusion module is connected to an MLP+Softmax classifier. The specific process is as follows: The training dataset is used with a visual attention mechanism and a non-local feature enhancement network backbone to extract and learn global feature information, resulting in feature information of different sizes. The network backbone has three self-attention layers, generating three different feature sizes. Specifically, the first attention module outputs feature map A, the second attention module outputs feature map B, and the third attention module outputs feature map C. After upsampling, the feature map is fused with the feature map output from the second attention module, and then subjected to further upsampling. The sampling feature pyramid model fuses feature information of different sizes generated by the backbone to obtain multi-scale feature maps; wherein: the C feature map is upsampled to output the D feature map, the D feature map is fused with the B feature map and then upsampled again to output the E feature map, and the E feature map is fused with the A feature map to obtain the F feature map. For the fused F feature map, MLP+softmax and confidence scores are used for discrimination. Specifically, the fused F feature map is stretched into a one-dimensional sequence, and then MLP+softmax is used to process the one-dimensional sequence, outputting binary classification confidence scores p1 and p2. If the scores are greater than the set confidence threshold, the classification is determined. Image blocks are identified as abnormal, i.e., image blocks with cracks; otherwise, they are image blocks without cracks.
[0007] The image contrast enhancement employs a histogram equalization image processing algorithm, which transforms the histogram distribution of the image into an approximately uniform distribution, thereby enhancing the image contrast.
[0008] The histogram equalization image processing algorithm is as follows: First, the image is segmented into many overlapping or non-overlapping regions. Then, the histogram of each region is calculated, and histogram equalization is applied to the pixels within each small region to enhance local contrast. Next, a contrast-limiting formula is used for contrast cropping, and the cropped regions are re-interpolated and combined to form the final equalized image. The contrast-limiting formula is:
[0009] in: To compare the pixel values of the image after the restriction, This is a weighted blending value for the image channels. It is the set of pixel values corresponding to the pixels of an image block. The value of the red channel in the image. The value of the green channel of the image. This represents the value of the blue channel in the image.
[0010] The ratio of the training to the test dataset is 8:2.
[0011] The method of using a visual attention mechanism for global feature learning is as follows: After positional encoding, global feature information is learned through a visual attention mechanism. The image, after positional encoding, becomes Q, K, and V, and then passes through a self-attention layer to learn global feature information. The final feature output is: .
[0012] The fusion formula for the feature pyramid model is: , Features at different scales.
[0013] The confidence threshold .
[0014] An electronic device is characterized by comprising a processor and a memory, the processor being configured to execute a computer program stored in the memory to implement the steps of the feature learning-based visual inspection method for workpiece cracks.
[0015] A readable storage medium, characterized in that a computer program is stored on the readable storage medium, and when the computer program is executed by a processor, it implements the steps of the workpiece crack visual detection method based on feature learning.
[0016] A computer program product characterized by comprising computer-executable instructions, which, when executed, are used to implement the method described.
[0017] Beneficial effects This invention proposes a visual crack detection method for workpieces based on feature learning. Due to the small size of cracks in actual workpiece images and their high similarity to normal areas, existing technologies have not been effective. This prompted us to change the original defect detection method based on image segmentation. Instead, we treat the crack detection problem in workpiece images as an image-based classification problem. Finally, we regress the segmented blocks to the detection result of the overall image. This new approach can effectively solve the crack detection problem under complex conditions (background noise interference, weak cracks).
[0018] In this invention, we simultaneously employ histogram equalization in the network structure to eliminate background interference and amplify the color difference between workpiece cracks and normal areas. The final detection result is then obtained through a classifier. Compared to other existing technologies, this method can accurately identify fine cracks in workpieces, and it demonstrates fast training speed, high recognition rate (accuracy reaches 50%), and recall rate exceeding 90% on real-world workpiece datasets. These performance metrics indicate that this method provides a more beneficial effect for workpiece image crack detection. The results are shown in Table 2. Attached Figure Description
[0019] Figure 1: Classification and detection network structure diagram of the present invention Figure 2 : Detailed flowchart of the implementation of this invention Detailed Implementation The present invention will now be further described in conjunction with the embodiments and accompanying drawings: This invention proposes a crack detection method based on sliding window and image features to overcome the limitations of general scene classification methods that struggle to detect cracks in workpiece image data for this specific problem. The technical solution comprises two main parts: sliding window segmentation and deep learning classification.
[0020] Sliding window segmentation: 1. Data augmentation is performed on data suspected of having cracks. Histogram equalization is used to process the data, making the histogram distribution of the image approximately uniform, thereby enhancing the image contrast and reducing noise interference, resulting in denoised training samples.
[0021] 2. Use a sliding window for data segmentation, for a given image. In region Q, using The window is obtained by sliding and dividing the region Q with a step size s. A set of image patches that transforms a large image with background interference into a low-resolution image set that focuses solely on the object. .
[0022] Deep learning classification: 1. For the image set P, randomly select and divide the training and test datasets, with a training-to-test ratio of 8:2.
[0023] 2. A visual attention mechanism and a non-local feature enhancement method are used as the network backbone to extract and learn the global feature information of the training samples. The training images are processed into image features through the feature network, which further amplifies the differences between abnormal and normal samples. The network model can learn the feature information of positive and negative samples.
[0024] 3. The sampling feature pyramid model fuses feature information of different sizes generated by the backbone to obtain a multi-scale feature map.
[0025] 4. Using a classification structure, the image features are normalized and then Softmax is used for discrimination. The abnormal confidence of the output image patch is then used to filter the confidence (set to 0.6) according to the designed threshold. The confidence is judged as abnormal and that cracks exist if the confidence is greater than the threshold; otherwise, cracks do not exist.
[0026] The network structure used in this invention is shown in [reference needed]. Figure 1 As shown.
[0027] In this embodiment of the invention, reference is made to Figure 2 The implementation of the sliding window handling module on the left is as follows: Step 1: Histogram equalization of the image The initial dataset is subjected to histogram equalization to reduce background and noise interference and enhance image contrast. Specifically, the image is first segmented into many overlapping or non-overlapping regions. Then, the histogram for each region is calculated, and histogram equalization is applied to the pixels within each small region to enhance local contrast. Next, contrast cropping is performed to limit the effect of histogram equalization and avoid over-enhancing contrast. The formula for limiting the degree of contrast is:
[0028] in: To compare the pixel values of the image after the restriction, This is a weighted blending value for the image channels. It is the set of pixel values corresponding to the pixels of an image block. The value of the red channel in the image. This refers to the green channel value of the image. This represents the value of the blue channel in the image.
[0029] Step 2, slide window to split data For an already equalized image, a sliding window is used for data partitioning. For a given image... In region Q, using The window is obtained by sliding and dividing the region Q with a step size s. A set of image patches that transforms a large image with background interference into a low-resolution image set that focuses solely on the object. .
[0030] m and n are respectively:
[0031] in, This represents the number of image patches along the y-axis of the image. This represents the number of image patches along the x-axis of the image. Image width, Image length; Reference Figure 2 The implementation steps of the sliding window processing module of the present invention are as follows: Step 3, Attention Feature Extraction: The segmented training samples are then processed using a visual attention mechanism for global feature learning. Specifically, after positional encoding, global feature information is learned through a visual attention mechanism. The image, after positional encoding, becomes Q, K, and V, and then passes through a self-attention layer to learn global feature information. The final feature output is as follows:
[0032] The entire feature learning and extraction module has three self-attention layers, which generate three features of different sizes.
[0033] Step 4, Feature Pyramid Fusion: For Different scale features are obtained by layer-by-layer upsampling and interpolation to fuse image features at different scales. The feature fusion formula is as follows:
[0034] Features at different scales.
[0035] Step 5, Result Classification: For the fused features, MLP+Softmax and confidence scores are used for discrimination. First, the feature image is stretched into a one-dimensional sequence, then MLP+Softmax is used to process the sequence, outputting binary classification confidence scores p1 and p2. If a certain confidence score is greater than p1, then the classification is performed. If the value is 1, it is classified into the corresponding category (1 indicates the presence of a crack, 0 indicates the absence of a crack). The discrimination formula is as follows:
[0036] After training is completed, the network structure of this invention is used to determine the category of the input image (i.e. whether there is a crack) during image test inference.
[0037] Steps 3 to 5 above employ a classification detection network including a visual attention mechanism and a non-local feature enhancement network backbone, a sampling feature pyramid model, and a classifier; training samples Divide the dataset into training and testing datasets. Input the training dataset into the classification and detection network for training. Once the classification and detection network is trained, it will classify and detect the data from step 1 after processing in step 2, and determine whether the image patch is abnormal. The classification and detection network includes a visual attention mechanism and non-local feature enhancement network backbone, a sampling feature pyramid model, and a classifier. The visual attention mechanism and non-local feature enhancement network backbone includes three cascaded attention modules. The sampling feature pyramid model includes two upsampling modules and a fusion module. The output of the third attention module is connected to the second upsampling module, the output of the second attention module is connected to the first upsampling module, the output of the second upsampling module is connected to the first upsampling module, and the output of the first upsampling module is connected to the fusion module. The output of the fusion module is connected to an MLP+Softmax classifier. The specific process is as follows: The training dataset is used with a visual attention mechanism and a non-local feature enhancement network backbone to extract and learn global feature information, resulting in feature information of different sizes. The network backbone has three self-attention layers, generating three different feature sizes. Specifically, the first attention module outputs feature map A, the second attention module outputs feature map B, and the third attention module outputs feature map C. After upsampling, the feature map is fused with the feature map output from the second attention module, and then subjected to further upsampling. The sampling feature pyramid model fuses feature information of different sizes generated by the backbone to obtain multi-scale feature maps; wherein: the C feature map is upsampled to output the D feature map, the D feature map is fused with the B feature map and then upsampled again to output the E feature map, and the E feature map is fused with the A feature map to obtain the F feature map. For the fused F feature map, MLP+softmax and confidence scores are used for discrimination. Specifically, the fused F feature map is stretched into a one-dimensional sequence, and then MLP+softmax is used to process the one-dimensional sequence, outputting binary classification confidence scores p1 and p2. If the scores are greater than the set confidence threshold, the classification is determined. Image blocks are identified as abnormal, i.e., image blocks with cracks; otherwise, they are image blocks without cracks.
[0038] The effects of this invention can be further illustrated by the following simulation experiments.
[0039] 1. Simulation conditions: This invention is a simulation performed using Python software on a system with an Intel® i5-3470 3.2GHz CPU, 4GB of memory, and a Windows 7 operating system. The data used in the experiment were self-acquired workpiece images.
[0040] 2. Simulation Content First, the images are segmented using the training set according to the method in the specific implementation, and then trained using the network model. After training is completed, the images in the segmented test set are classified, and the classification accuracy and recall are calculated by combining the results with the real labels.
[0041] To demonstrate the effectiveness of the algorithm, we selected two algorithms for comparison. One is the one proposed by Tao et al., “X. Tao, D. Zhang, W. Ma, X. Liu, and D. Xu. Automatic metallic surface defect detection and recognition with convolutional neural networks. Applied Sciences, 8, 1575, 2018”. The other is the one proposed by Li et al., “Z. Li, N. Li, K. Jiang, Ma. Zhang, W. Xing, H. Xiao, and Y. Gong, Superpixel Masking and Inpainting for Self-Supervised Anomaly Detection. In Proceedings of British Machine Vision Conference, 2020”, with parameter tuning, and the average accuracy, recall, and F-score were calculated. The comparison results are shown in Table 2.
[0042] Table 1 Experimental Data Setup
[0043] Table 1 shows the training data used for learning. The comparative experiment followed the image count settings in Table 1, and the results are shown in Table 2. The average precision of this invention is approximately 52.13%, and the average recall is approximately 93.12%. This indicates that the overall detection performance of this method is good. It demonstrates that this method has a good effect on crack detection in industrial data.
[0044] Table 2 Comparison of Experimental Results
[0045] In summary, this invention combines feature learning with crack detection, exploring how to learn more effective feature information through labeled training sets, enabling the method to achieve high accuracy and strong robustness in crack detection. At the same time, the sliding window algorithm is used to process the data, thereby enhancing and denoising the dataset and improving the accuracy of the data during detection.
Claims
1. A visual detection method for workpiece cracks based on feature learning, characterized in that... The steps are as follows: Step 1: Enhance the contrast of images of workpieces suspected of containing cracks and reduce noise interference in the images; Step 2: Use a sliding window to view the image Data segmentation yields multiple image patches, forming an image set. The training samples after denoising are obtained. ; The image set The composition is: images In region Q, using The window is obtained by sliding and dividing the region Q with a step size s. Image patches constitute an image set ,in: ; This represents the number of image patches along the y-axis of the image. This represents the number of image patches along the x-axis of the image. Image width, Image length; Step 3: Transfer the training samples Divide the dataset into training and testing datasets. Input the training dataset into the classification and detection network for training. Once the classification and detection network is trained, it will classify and detect the data from step 1 after processing in step 2, and determine whether the image patch is abnormal. The classification and detection network includes a visual attention mechanism and non-local feature enhancement network backbone, a sampling feature pyramid model, and a classifier. The visual attention mechanism and non-local feature enhancement network backbone includes three cascaded attention modules. The sampling feature pyramid model includes two upsampling modules and a fusion module. The output of the third attention module is connected to the second upsampling module, the output of the second attention module is connected to the first upsampling module, the output of the second upsampling module is connected to the first upsampling module, and the output of the first upsampling module is connected to the fusion module. The output of the fusion module is connected to the MLP+Sofmax classifier. The specific process is as follows: The training dataset is used with a visual attention mechanism and a non-local feature enhancement network backbone to extract and learn global feature information, resulting in feature information of different sizes. The network backbone has three self-attention layers, generating three different feature sizes. Specifically, the first attention module outputs feature map A, the second attention module outputs feature map B, and the third attention module outputs feature map C. After upsampling, the feature map is fused with the feature map output from the second attention module, and then subjected to further upsampling. The sampling feature pyramid model fuses feature information of different sizes generated by the backbone to obtain multi-scale feature maps; wherein: the C feature map is upsampled to output the D feature map, the D feature map is fused with the B feature map and then upsampled again to output the E feature map, and the E feature map is fused with the A feature map to obtain the F feature map. For the fused F feature map, MLP+softmax and confidence scores are used for discrimination. Specifically, the fused F feature map is stretched into a one-dimensional sequence, and then MLP+softmax is used to process the one-dimensional sequence, outputting binary classification confidence scores p1 and p2. If the scores are greater than the set confidence threshold, the classification is determined. Image blocks are identified as abnormal, i.e., image blocks with cracks; otherwise, they are image blocks without cracks.
2. The workpiece crack visual detection method based on feature learning according to claim 1, characterized in that: The image contrast enhancement employs a histogram equalization image processing algorithm, which transforms the histogram distribution of the image into an approximately uniform distribution, thereby enhancing the image contrast.
3. The workpiece crack visual detection method based on feature learning according to claim 2, characterized in that: The histogram equalization image processing algorithm is as follows: First, the image is segmented into many overlapping or non-overlapping regions. Then, the histogram of each region is calculated, and histogram equalization is applied to the pixels within each small region to enhance local contrast. Next, a contrast-limiting formula is used for contrast cropping, and the cropped regions are re-interpolated and combined to form the final equalized image. The contrast-limiting formula is: in: To compare the pixel values of the image after the restriction, This is a weighted blending value for the image channels. It is the set of pixel values corresponding to the pixels of an image block. The value of the red channel in the image. The value of the green channel of the image. This represents the value of the blue channel in the image.
4. The workpiece crack visual detection method based on feature learning according to claim 1, characterized in that: The ratio of the training to the test dataset is 8:
2.
5. The workpiece crack visual detection method based on feature learning according to claim 1, characterized in that: The method of using a visual attention mechanism for global feature learning is as follows: After positional encoding, global feature information is learned through a visual attention mechanism. The image, after positional encoding, becomes Q, K, and V, and then passes through a self-attention layer to learn global feature information. The final feature output is: .
6. The workpiece crack visual detection method based on feature learning according to claim 1, characterized in that: The fusion formula for the feature pyramid model is: , Features at different scales.
7. The workpiece crack visual detection method based on feature learning according to claim 1, characterized in that: The confidence threshold .
8. An electronic device, characterized in that, It includes a processor and a memory, wherein the processor is used to execute a computer program stored in the memory to implement the steps of the feature learning-based visual inspection method for workpiece cracks as described in any one of claims 1 to 7.
9. A readable storage medium, characterized in that, The readable storage medium stores a computer program that, when executed by a processor, implements the steps of the feature learning-based visual inspection method for workpiece cracks as described in any one of claims 1 to 7.
10. A computer program product, characterized in that... It includes computer-executable instructions, which, when executed, are used to implement the method of any one of claims 1 to 7.
Citation Information
Patent Citations
Metal part crack detection method based on deep learning and ultrasonic infrared image
CN115456938A
Machine vision-based spacecraft workpiece CT image defect detection network and method
CN117934390A