Weld defect detection system based on improved U-Net network

By improving the U-Net network and combining it with depthwise separable convolutional residual blocks, multi-attention feature fusion modules, and CBCNM modules, the problem of insufficient feature extraction in weld defect detection is solved, and efficient and accurate weld detection is achieved in complex environments.

CN120852262APending Publication Date: 2025-10-28TIANJIN POLYTECHNIC UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410502147.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-25
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing deep learning methods fail to effectively integrate spatial information, global features, and local features during the welding process, resulting in insufficient feature extraction capabilities for weld defect detection in complex environments, especially in low-contrast weld images where end-to-end accurate detection is difficult.

Method used

An improved U-Net network is adopted, which combines the depthwise separable convolutional residual block (DSCR), the multi-attention feature fusion module (AFFM), and the CBCNM module. The feature extraction capability is improved through depthwise separable convolution and attention mechanism, and the ONNX Runtime library is used for model inference detection.

Benefits of technology

It improves the accuracy and efficiency of weld defect detection, can accurately extract weld features in complex environments, reduces model complexity, and adapts to resource-constrained environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120852262A_ABST
    Figure CN120852262A_ABST
Patent Text Reader

Abstract

The invention provides a weld defect detection system based on an improved U-Net network, which uses an improved U-Net semantic segmentation network for training, and uses a trained network model to segment a weld image for detecting the weld quality. The deep separable convolution residual block DSCR, the multi-attention feature fusion module AFFM, the CBCNM module and the U-Net are combined to enable the network to segment the weld joint more accurately, then the improved U-Net network model is packaged into a deep learning detection algorithm of the system, weld joint detection intelligence is achieved, and therefore weld joint quality is accurately detected, and weld joint defects are judged.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of machine vision technology, specifically a weld defect detection system based on an improved U-Net network. Background Art

[0002] With the development of deep learning technology and the improvement of computer equipment performance, the application of artificial intelligence-related technologies has developed rapidly. Semantic segmentation has always been one of the important tasks of artificial intelligence in the field of computer vision. By assigning category labels to each pixel in an image, it helps humans understand images more accurately and plays a key role in many practical applications. Compared with traditional machine learning methods, deep convolutional neural networks have stronger feature extraction and object representation capabilities. For the automotive manufacturing industry, welding technology is a crucial technology that directly affects the quality and performance of the entire vehicle. During the welding process, external environmental factors and other non-human factors can interfere with the robotic welding process, leading to various unpredictable welding defects such as weld breaks, weld leaks, cracks, and slag inclusions. Timely and accurate detection of weld defects during the welding process is essential for improving efficiency and saving costs.

[0003] With the development of deep learning in the field of image segmentation, deep learning-based weld seam image segmentation algorithms have made breakthrough progress. Among them, "Li Pengfei. Industrial Defect Segmentation and Automatic Detection Based on Semantic Context [D]. Tianjin University, 2021" proposed a context refinement network called CRNet. This network has been applied to the field of automated detection of welding defects. By analyzing the visual attributes of the target and its surrounding environment, it improves the model's adaptability to features and recognition accuracy. The paper "Bai Lianfa, Wang Yeyu, Zhao Zhuang, et al. A Weld Seam Tracking Method Based on Feature Segmentation [P]. Jiangsu Province: CN202110277763.0, 2022-11-11" uses a multi-scale feature fusion approach to improve the network's ability to extract welding feature points. It uses feature maps from the encoder and decoder to stitch together, thereby obtaining more spatial and semantic information. The paper "Luo Renzhe, Li Huadu, Tang Xiang, et al. A Pipeline Weld Seam Image Segmentation Method Based on Improved UNet [P]. Sichuan Province: CN202310392019.4, 2023-07-11" uses a parallel attention mechanism to improve the U-Net network encoder's perception and feature extraction capabilities for weld seam areas. However, existing deep learning methods do not comprehensively consider spatial information, global features, and local features, and their ability to extract weld seam features in complex environments needs further improvement. In order to combine spatial information, global features and local features, and extract valuable feature information from low-contrast weld images in a short time, so as to achieve accurate end-to-end extraction of weld images, this invention proposes a weld defect detection system based on an improved U-Net network. Summary of the Invention

[0004] In view of this, the present invention aims to propose a weld defect detection system based on an improved U-Net network, which is used to complete the semantic segmentation task of weld images and achieves high-precision feature information results.

[0005] To achieve the above objectives, this invention proposes a weld defect detection system based on an improved U-Net network, comprising the following steps:

[0006] S1: Develop weld inspection software based on an improved U-Net network;

[0007] S2: Use the labelme tool to label and augment the images captured by the camera;

[0008] S3: Train the improved U-Net semantic segmentation model using weld seam images, and update the improved U-Net semantic segmentation network during the training process using Depthically Separable Convolutional Residual Block (DSCR), Multiple Attention Feature Fusion (AFFM), and CBCNM modules to obtain the trained model.

[0009] S4: The weld seam detection system uses ONNX Runtime library functions in the C++ program to call the trained model file to perform inference detection on the weld seam image.

[0010] Furthermore, step S2, which involves labeling and data augmentation of the images captured by the camera, specifically includes:

[0011] First, a noise reduction algorithm is used to denoise the weld area to reduce noise interference. Then, the denoised image is enhanced to improve detection tolerance. Finally, LabelMe is used to annotate the camera image and add corresponding labels.

[0012] Furthermore, the depth-separable convolutional residual block (DSCR) mentioned in step S3 specifically refers to:

[0013] The Depthwise Separable Convolutional Residual Block (DSCR) uses depthwise separable convolution to reduce the number of parameters. Each DSCR Block first uses a 1x1 convolution to increase the dimensionality, then uses a 3x3 depthwise separable convolution (DWConv) to extract features, and then uses a 1x1 convolution for dimensionality reduction. To maintain spatial similarity in the extracted features, a combination of depthwise separable convolution and point convolution is used to refine each feature while maintaining the feature dimensionality in line with the network structure requirements. Finally, a 1x1 convolution is used at the residual edge to make the number of output channels the same as the backbone, and then the channels are summed. The Depthwise Separable Convolutional Residual Block (DSCR) is used for encoder feature extraction.

[0014] Furthermore, the multi-attention feature fusion module AFFM mentioned in step S3 is specifically as follows:

[0015] In the decoder, the deep feature maps acquire category-related semantic information through the channel attention module (CAM), while the shallower feature maps in the encoder acquire location-related spatial information through the spatial attention module (SAM). These two types of feature maps are superimposed and further input into the channel attention module, ultimately producing a comprehensive feature map that incorporates both spatial and semantic information. This module has the ability to extract image features from multiple dimensions and can effectively improve the category determination and location localization of weld seam areas.

[0016] The channel attention module processes the input feature map through a global average pooling layer, then uses two multilayer perceptron layers to improve generalization performance, followed by non-linear feature transformation using an MLP. After pixel-level summation, the resulting attention weights are activated by an activation function. Finally, a scaling factor is used to weight the attention weights onto the features of each channel. The calculation formula can be defined as:

[0017] Mc(X)=W1(M(M(δ1[AvgPool(X)])))

[0018] In the formula, X is the input feature map, W1 represents the scaling factor used to adjust the importance of features, M represents the multilayer perceptron, ReLU function, δ1 represents 1×1 convolution, and AvgPool represents the average pooling operation.

[0019] The spatial attention module first extracts weight information, then multiplies this weight information with the original feature map to obtain the attention enhancement effect. It learns the correlation between pixels in space using the weights, and then generates spatial weight coefficients using the Sigmoid function. These weight coefficients are multiplied with the input feature map to obtain the feature map infused with spatial attention. The calculation formula can be defined as:

[0020] Ms(X)=σ(δ7[AvgPool(X)])

[0021] In the formula, X is the input feature map, σ represents the Sigmoid function, and δ7 represents a 7×7 convolution.

[0022] Furthermore, the CBCNM module mentioned in step S3 specifically includes:

[0023] The CBCNM module replaces batch channel normalization in the original ConvMixer structure with batch normalization, placing the ConvMixer detection network at the end of the U-Net concatenation. The ConvMixer module fuses global features from the Transformer with local features from the CNN, combining the global feature extraction capabilities of the Transformer with the local feature extraction capabilities of the CNN during feature extraction. The CBCNM module better adapts to the correlation between feature channels in convolutional neural networks, improving its adaptability to small batches of samples. The calculation formula can be defined as:

[0024] z0=BCN(GELU{δ cin→h (X, stride=p, kernel size=p)})

[0025] Where BCN represents batch channel normalization, Conv cin→h This represents a convolution with input channel cin, output channel h, kernel size p, and stride p.

[0026] Z1=BCN(GELU{δ DW (z 1-1 )})+z 1-1

[0027] Z 1+1 =BCN(GELU{δ PW (z1)})

[0028] Where, δ DW Denotes channel convolution, δ PW This represents pointwise convolution.

[0029] The CBCNM layer mainly consists of depthwise convolution and pointwise convolution. Each convolution is followed by a GELU activation and batch channel normalization, finally yielding a feature map of the same size as the input.

[0030] Furthermore, the weld detection system described in step S4 utilizes ONNX Runtime library functions to call the trained model file to perform inference detection on the weld image, specifically as follows:

[0031] The network training model converts its weight files to ONNX file format using the ONNX Runtime model library. Then, the semantic segmentation inference is accelerated by CPU calls to ONNX Runtime library functions in the C++ program. ONNXRuntime is an open-source engine for efficient inference of ONNX models. ONNX is an open deep learning model exchange format that can be used to convert deep learning models from one framework to another, thereby enabling cross-platform and cross-framework model deployment and inference.

[0032] Compared with existing technologies, the weld defect detection system estimation method based on the improved U-Net network described in this invention has the following advantages:

[0033] (1) This invention employs Depthwise Separable Convolutional Residual Blocks (DSCR). The DSCR module can effectively balance the computational speed and classification accuracy of the network. Furthermore, it can effectively reduce the complexity of the model and improve its performance in resource-constrained environments.

[0034] (2) The present invention applies the attention feature fusion module AFFM to the U-Net network model. This module can extract features of multiple dimensions of the image and enhance the learning of defect region category and location information.

[0035] (3) This invention integrates the CBCNM module with the U-Net network model and replaces the batch channel normalization in the original ConvMixer structure with the batch normalization, thereby better adapting to the correlation between feature channels in the convolutional neural network and improving the adaptability to small batch samples. Attached Figure Description

[0036] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an undue limitation of the invention. In the drawings:

[0037] Figure 1 This is a flowchart of a weld defect detection system based on an improved U-Net network according to the present invention;

[0038] Figure 2 This is a structural diagram of the Depth Separable Convolutional Residual Block (DSCR) of the present invention;

[0039] Figure 3 This is a structural diagram of the multi-attention feature fusion module AFFM of the present invention;

[0040] Figure 4 This is a network diagram of the CBCNM module of the present invention;

[0041] Figure 5 The weld image input for this invention;

[0042] Figure 6 This is the semantic segmentation map predicted by the present invention; Detailed Implementation

[0043] This invention proposes a weld defect detection system based on an improved U-Net network. The invention will be described in more detail below with reference to the accompanying drawings and specific embodiments.

[0044] In this embodiment, the following steps are included:

[0045] Step 1: After selecting the image samples to be detected, preprocess the images first, and then input the images into the U-Net improved semantic segmentation network. Figure 5 The processed image shown.

[0046] Step 2: After inputting the image, it first undergoes a 3×3 depthwise separable convolution, ReLU activation, batch normalization (BN) operation, and one max pooling operation for downsampling. Then it undergoes further processing... Figure 2 After the depthwise separable convolutional residual block is shown, five preliminary effective feature layers are obtained. The first four layers are upsampled using convolution. Then, the upsampled feature maps are fused with the feature maps after the depthwise separable convolutional residual block and further processed. Figure 3 The multi-attention feature fusion module shown is followed by a standard convolutional feature extraction module to further extract features. In layer 5, the feature map and its attachments... Figure 4 The feature maps output by the CBCNM module are concatenated, and finally a 1x1 convolution is used for channel adjustment to adjust the number of channels in the final feature layer to match the number of classes, thus obtaining the prediction result. (The last sentence appears to be incomplete and possibly refers to a separate process.) Figure 6 The predicted semantic segmentation graph.

[0047] Step 3: Convert the weight file of the network training model to ONNX file format using the ONNX runtime model library. Then, in the C++ program, call the ONNX runtime library functions to accelerate semantic segmentation inference using the CPU. After running the communication thread, monitor in real time whether the PLC server is sending information. Upon receiving information, save the record and simultaneously feed it back to the main thread. After the image processing thread receives the weld seam image into the global variable area, process the image using the detection algorithm, send the processed image and detection results to the main thread for further processing, and then clear the global variable area to await the next detection. When the main thread receives information from the communication thread and the image sent by the image processing thread, it displays the image on the main interface and simultaneously notifies the communication thread to provide feedback.

[0048] Step 4: Determine if the image meets the detection standards. If it does not, staff will conduct a re-inspection and verification. If it meets the standards, the data will be stored and recorded. The detection results will be displayed on the main interface and sent to the PLC server.

[0049] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A weld defect detection system based on an improved U-Net network, characterized in that, Includes the following steps: S1: Develop weld inspection software based on an improved U-Net network; S2: Use the labelme tool to label and augment the images captured by the camera; S3: Train the improved U-Net semantic segmentation model using weld seam images, and update the improved U-Net semantic segmentation network during the training process using Depthically Separable Convolutional Residual Block (DSCR), Multiple Attention Feature Fusion (AFFM), and CBCNM modules to obtain the trained model. S4: The weld seam detection system uses ONNX Runtime library functions in the C++ program to call the trained model file to perform inference detection on the weld seam image.

2. The method for annotating and data augmenting images captured by a camera according to claim 1, characterized in that: The image processing and labeling described above involves using a noise reduction algorithm to denoise the weld area to reduce noise interference, followed by image enhancement to improve detection tolerance. Then, LabelMe is used to label the camera images and add appropriate tags.

3. The Depth-Separable Convolutional Residual Block (DSCR) according to claim 1, characterized in that: The Depthly Separable Convolutional Residual Block (DSCR) uses depthly separable convolution to reduce the number of parameters. Each DSCR Block first uses a 1x1 convolution to increase the dimensionality, then uses a 3x3 depthly separable convolution (DWConv) to extract features, and then uses a 1x1 convolution for dimensionality reduction. To maintain spatial similarity in the extracted features, a combination of depthly separable convolution and point convolution is used to refine each feature while maintaining the feature dimensionality in line with the network structure requirements. Finally, a 1x1 convolution is used at the residual edge to make the number of output channels the same as the backbone, and then the channels are summed. The Depthly Separable Convolutional Residual Block (DSCR) is used for encoder feature extraction.

4. The multi-attention feature fusion module AFFM according to claim 1, characterized in that: In the multi-attention feature fusion module, the deep feature map of the decoder obtains category-related semantic information through the channel attention module (CAM), while the shallower feature map of the encoder obtains location-related spatial information through the spatial attention module (SAM). These two types of feature maps are superimposed and further input into the channel attention module, ultimately producing a comprehensive feature map that includes both spatial and semantic information. This module has the ability to extract image features from multiple dimensions and can effectively improve the category determination and location positioning of the weld area.

5. The channel attention module according to claim 4, characterized in that: The aforementioned channels pass the input feature map through a global average pooling layer, then use two multilayer perceptron layers to improve generalization performance, followed by non-linear feature transformation through an MLP. After pixel-level summation, the corresponding attention weights are obtained by activation function. Finally, a scaling factor is used to weight the previously obtained attention weights onto the features of each channel. The calculation formula can be defined as follows: Mc(X)=W1(M(M(δ1[AvgPool(X)]))) In the formula, X is the input feature map, W1 represents the scaling factor used to adjust the importance of features, M represents the multilayer perceptron, ReLU function, δ1 represents 1×1 convolution, and AvgPool represents the average pooling operation.

6. The spatial attention module according to claim 4, characterized in that: The spatial attention module first extracts weight information, then multiplies the weight information with the original feature map to obtain the attention enhancement effect. It learns the correlation between pixels in space using the weights, and then generates spatial weight coefficients using the Sigmoid function. The weight coefficients are multiplied with the input feature map to obtain the feature map injected with spatial attention. The calculation formula can be defined as: Ms(X)=σ(δ7[AvgPool(X)]) In the formula, X is the input feature map, σ represents the Sigmoid function, and δ7 represents a 7×7 convolution.

7. The CBCNM module according to claim 1, characterized in that: The CBCNM Block module replaces batch channel normalization in the original ConvMixer structure with batch normalization. The ConvMixer detection network is placed at the end of the U-Net concatenation. The ConvMixer module fuses global features from the Transformer with local features from the CNN. During feature extraction, it combines the global feature extraction capabilities of the Transformer with the local feature extraction capabilities of the CNN. The CBCNM module can better adapt to the correlation between feature channels in convolutional neural networks, improving its adaptability to small batches of samples. The calculation formula can be defined as: z0=BCN(GELU{δ cin→h (X,stride=p,kernel size=p)}) Where BCN represents batch channel normalization, Conv cin→h This represents a convolution with input channel cin, output channel h, kernel size p, and stride p. Z1=BCN(GELU{δ DW (With 1-1 )})+z 1-1 FROM 1+1 =BCN(GELU{δ PW (z1)}) Where, δ DW Denotes channel convolution, δ PW This represents pointwise convolution; The CBCNM layer mainly consists of depthwise convolution and pointwise convolution. Each convolution is followed by a GELU activation and batch channel normalization, finally yielding a feature map of the same size as the input.

8. The weld inspection system according to claim 1 calls the trained model file, characterized in that: The weld seam inspection system calls the model file by converting the weight file of the network training model into the ONNX file format through the ONNX Runtime model library. After calling the ONNX Runtime library function on the software C++ program side, the semantic segmentation inference is accelerated by the CPU. ONNX Runtime is an open source engine for efficient inference of ONNX models. ONNX is an open deep learning model exchange format that can be used to convert deep learning models from one framework to another, thereby realizing cross-platform and cross-framework model deployment and inference.